Attention

Sliding-window attention

Each token only attends to a fixed window of recent tokens, keeping cost and cache flat as context grows.

First page of Longformer: The Long-Document TransformerIntroduced inApr 2020Longformer: The Long-Document TransformerBeltagy et al. · arXiv 2004.05150 ↗
Full attentionSliding windowevery token sees all before iteach token sees the last 4hybrid stacks alternate:SSSFSSSF

Rather than looking at the entire history, a sliding-window layer attends only to the last N tokens. Models like Gemma and gpt-oss interleave many sliding layers with occasional full-attention layers: the sliding layers handle local structure cheaply while the periodic full layers carry long-range information. This hybrid keeps the KV cache small without giving up long-context ability.

Adoption over time

Share of new models that have included sliding-window attention over time.

0%202220232024202520260%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

google/gemma-4-26B-A4B-it architecture graphgoogle/gemma-4-26B-A4B-itimage-text-to-text · ↓ 13.1M · ♡ 2kOpen in visualizer google/gemma-4-31B-it architecture graphgoogle/gemma-4-31B-itimage-text-to-text · ↓ 9.0M · ♡ 4kOpen in visualizer openai/gpt-oss-20b architecture graphopenai/gpt-oss-20btext-generation · ↓ 6.7M · ♡ 5kOpen in visualizer ibm-granite/granite-embedding-small-english-r2 architecture graphibm-granite/granite-embedding-small-english-r2feature-extraction · ↓ 6.3M · ♡ 83Open in visualizer openai/gpt-oss-120b architecture graphopenai/gpt-oss-120btext-generation · ↓ 5.0M · ♡ 5kOpen in visualizer answerdotai/ModernBERT-base architecture graphanswerdotai/ModernBERT-basefill-mask · ↓ 4.4M · ♡ 1kOpen in visualizer google/gemma-4-E2B-it architecture graphgoogle/gemma-4-E2B-itany-to-any · ↓ 3.4M · ♡ 971Open in visualizer google/embeddinggemma-300m architecture graphgoogle/embeddinggemma-300msentence-similarity · ↓ 2.7M · ♡ 2kOpen in visualizer Alibaba-NLP/gte-reranker-modernbert-base architecture graphAlibaba-NLP/gte-reranker-modernbert-basetext-ranking · ↓ 2.6M · ♡ 98Open in visualizer google/gemma-4-12B-it architecture graphgoogle/gemma-4-12B-itany-to-any · ↓ 2.5M · ♡ 2kOpen in visualizer google/gemma-3-4b-it architecture graphgoogle/gemma-3-4b-itimage-text-to-text · ↓ 1.8M · ♡ 2kOpen in visualizer nvidia/Gemma-4-26B-A4B-NVFP4 architecture graphnvidia/Gemma-4-26B-A4B-NVFP4text-generation · ↓ 1.7M · ♡ 149Open in visualizer

View 535 catalog matches for Sliding-window attention →

Related concepts

hfviewer renders the full architecture of 6,700+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →