Attention

Self-attention

Every token builds queries, keys and values, then gathers information from the other tokens that match its query.

First page of Attention Is All You NeedIntroduced inJun 2017Attention Is All You NeedVaswani et al. · arXiv 1706.03762 ↗
Self-attention: every token asks, all others answerthe query compares with each earlier token · weights (bars) set the blendquery

Self-attention is the core transformer operation: each position emits a query that is compared against the keys of the other positions, and the resulting weights average their values. It is what lets the model relate any token to any other, at a cost that grows with the square of sequence length. Which is why so many of the surrounding tricks (GQA, sliding windows, sparse and linear attention) exist to tame it.

Adoption over time

Share of new models that have included self-attention over time.

34%2022202320242025202634%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

sentence-transformers/all-MiniLM-L6-v2 architecture graphsentence-transformers/all-MiniLM-L6-v2sentence-similarity · ↓ 252.2M · ♡ 6kOpen in visualizer cross-encoder/ms-marco-MiniLM-L6-v2 architecture graphcross-encoder/ms-marco-MiniLM-L6-v2text-ranking · ↓ 88.7M · ♡ 347Open in visualizer BAAI/bge-small-en-v1.5 architecture graphBAAI/bge-small-en-v1.5feature-extraction · ↓ 64.4M · ♡ 588Open in visualizer google/electra-base-discriminator architecture graphgoogle/electra-base-discriminator↓ 50.1M · ♡ 183Open in visualizer google-bert/bert-base-uncased architecture graphgoogle-bert/bert-base-uncasedfill-mask · ↓ 46.6M · ♡ 3kOpen in visualizer sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 architecture graphsentence-transformers/paraphrase-multilingual-MiniLM-L12-v2sentence-similarity · ↓ 45.7M · ♡ 1kOpen in visualizer BAAI/bge-m3 architecture graphBAAI/bge-m3sentence-similarity · ↓ 37.6M · ♡ 4kOpen in visualizer google-t5/t5-small architecture graphgoogle-t5/t5-smalltranslation · ↓ 24.8M · ♡ 634Open in visualizer Comfy-Org/MiniMax-H3 architecture graphComfy-Org/MiniMax-H3↓ 23.3M · ♡ 2kOpen in visualizer openai/clip-vit-base-patch32 architecture graphopenai/clip-vit-base-patch32zero-shot-image-classification · ↓ 21.9M · ♡ 2kOpen in visualizer FacebookAI/xlm-roberta-base architecture graphFacebookAI/xlm-roberta-basefill-mask · ↓ 21.0M · ♡ 923Open in visualizer jonatasgrosman/wav2vec2-large-xlsr-53-japanese architecture graphjonatasgrosman/wav2vec2-large-xlsr-53-japaneseautomatic-speech-recognition · ↓ 17.7M · ♡ 87Open in visualizer

View 2,289 catalog matches for Self-attention →

Related concepts

hfviewer renders the full architecture of 6,800+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →