Mixture of experts

Grouped expert routing

Experts are organized into groups; the router first picks groups, then experts inside them. Friendlier to multi-GPU serving.

First page of DeepSeek-V3 Technical ReportIntroduced inDec 2024DeepSeek-V3 Technical ReportDeepSeek-AI · arXiv 2412.19437 ↗
Grouped routing: pick a group, then expertsgroups map to devices · each token’s experts stay on a couple of GPUstokenGPU 0GPU 11. pick the group   2. pick experts inside it

With hundreds of experts spread across devices, letting a token pick any k experts scatters traffic everywhere. Grouped routing first selects the best expert groups (which map to devices), then the best experts within those groups. DeepSeek introduced this ‘node-limited’ routing to cap cross-device communication; you will see it in graphs as a two-stage top-k.

Adoption over time

Share of new models that have included grouped expert routing over time.

1%202220232024202520261%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

autogluon/chronos-2 architecture graphautogluon/chronos-2time-series-forecasting · ↓ 8.7M · ♡ 72Open in visualizer autogluon/chronos-2-small architecture graphautogluon/chronos-2-smalltime-series-forecasting · ↓ 6.7M · ♡ 7Open in visualizer nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 architecture graphnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16text-generation · ↓ 1.3M · ♡ 426Open in visualizer nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 architecture graphnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4text-generation · ↓ 1.0M · ♡ 431Open in visualizer zai-org/GLM-5.2-FP8 architecture graphzai-org/GLM-5.2-FP8text-generation · ↓ 978k · ♡ 266Open in visualizer zai-org/GLM-5.3 architecture graphzai-org/GLM-5.3text-generation · ↓ 974k · ♡ 2kOpen in visualizer zai-org/GLM-5.2 architecture graphzai-org/GLM-5.2text-generation · ↓ 962k · ♡ 5kOpen in visualizer nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 architecture graphnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4text-generation · ↓ 827k · ♡ 180Open in visualizer nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 architecture graphnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4text-generation · ↓ 744k · ♡ 436Open in visualizer nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 architecture graphnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16text-generation · ↓ 666k · ♡ 823Open in visualizer nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 architecture graphnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16text-generation · ↓ 570k · ♡ 218Open in visualizer nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 architecture graphnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8text-generation · ↓ 409k · ♡ 360Open in visualizer

View 116 catalog matches for Grouped expert routing →

Related concepts

hfviewer renders the full architecture of 6,700+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →