Self-Attention
The type of attention that calculates the relationship of each element in a number with all other items in the same directory.
The biggest technical advantage of the essence-dickatin is that it can be paralleled: Unlike RNNs, the account for all locations in the array can be done at the same time, with matrix palpitations; which provides a lot of acceleration in the GPUs, and the very long-distance dependency of the model - for example, the relationship between the verb at the end of a paragraph - allows direct capture. In the technique called "Multi-head" (multi-head) attention, the model simultaneously learns in parallel with the different aspects of the language (specific, meaningal). Öz-dikkat is the direct source of the processing power and scalability of the transformer architecture and therefore large language models of today.
