Attention

Mamba / state-space block (SSM)

A recurrent state-space scan replaces attention: constant memory per token, linear time in sequence length.

First page of Mamba: Linear-Time Sequence Modeling with Selective State SpacesIntroduced inDec 2023Mamba: Linear-Time Sequence Modeling with Selective State SpacesGu & Dao · arXiv 2312.00752 ↗
Attentionthe KV cache grows with every tokenSelective state (Mamba)conv, then a gated scan rewrites SScost per token growscost stays constant

State-space blocks (Mamba-2’s SSD scan here) carry information through a fixed-size hidden state that is updated token by token, instead of comparing every token against every other one. That makes both compute and memory linear in sequence length. Hybrid models such as Nemotron-H use mostly Mamba blocks with a few attention layers mixed in for tasks that need exact token recall.

Adoption over time

Share of new models that have included Mamba/SSM blocks over time.

0%202220232024202520260%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 architecture graphnvidia/NVIDIA-Nemotron-3-Nano-4B-BF16text-generation · ↓ 3.5M · ♡ 121Open in visualizer nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 architecture graphnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16text-generation · ↓ 1.3M · ♡ 426Open in visualizer nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 architecture graphnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4text-generation · ↓ 1.0M · ♡ 431Open in visualizer nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8 architecture graphnvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8any-to-any · ↓ 959k · ♡ 63Open in visualizer nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 architecture graphnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4text-generation · ↓ 827k · ♡ 180Open in visualizer nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 architecture graphnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4text-generation · ↓ 744k · ♡ 436Open in visualizer nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 architecture graphnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16text-generation · ↓ 666k · ♡ 823Open in visualizer nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 architecture graphnvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4any-to-any · ↓ 640k · ♡ 190Open in visualizer nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 architecture graphnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16text-generation · ↓ 570k · ♡ 218Open in visualizer nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 architecture graphnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8text-generation · ↓ 409k · ♡ 360Open in visualizer nvidia/NVIDIA-Nemotron-Nano-9B-v2 architecture graphnvidia/NVIDIA-Nemotron-Nano-9B-v2text-generation · ↓ 383k · ♡ 519Open in visualizer RedHatAI/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-FP8 architecture graphRedHatAI/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-FP8text-generation · ↓ 357k · ♡ 2Open in visualizer

View 87 catalog matches for Mamba / state-space block (SSM) →

Related concepts

hfviewer renders the full architecture of 6,700+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →