Transformer
Based on the attention mechanism, the structure of the nervous network, which constitutes the architectural basis of today’s large language models.
A Transformer block is typically composed of increased kneeling of advanced-feeded nerve net layers with self-solid layers; no longer links (residual connections) and layer normalization helps very deep networks to be stable trained. Original architectural encoder-coder (encoder-decoder) was in its structure; models such as GPT only use coder, BERT only the encoder part. The scalability of the transformer - i.e. showing continuous performance increase with more data and parameters - the large language model has been directly triggered by the age, the image beyond the language processing (Vision Transformer) and audio processing are also spread to areas.
