Diffusion & generation

Timestep embedding

Tells the network how noisy the input currently is, so one network can handle every denoising stage.

First page of Denoising Diffusion Probabilistic ModelsPopularized inJun 2020Denoising Diffusion Probabilistic ModelsHo et al. · arXiv 2006.11239 ↗
Timestep embedding: telling the net how noisyt samples a stack of sinusoids · the pattern becomes a conditioning vectortinjected into every block

The same network weights must denoise both nearly-pure noise and nearly-finished images. The current timestep is therefore encoded (usually sinusoidally, then through a small MLP) and injected into every block, letting the network condition its behavior on how far along the denoising process is.

Adoption over time

Share of new models that have included timestep embeddings over time.

13%2022202320242025202613%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

Comfy-Org/MiniMax-H3 architecture graphComfy-Org/MiniMax-H3↓ 23.3M · ♡ 2kOpen in visualizer Comfy-Org/Krea-2 architecture graphComfy-Org/Krea-2↓ 9.8M · ♡ 573Open in visualizer Comfy-Org/Qwen-Image-2.1 architecture graphComfy-Org/Qwen-Image-2.1↓ 6.0MOpen in visualizer MiniMaxAI/MiniMax-H3 architecture graphMiniMaxAI/MiniMax-H3image-text-to-video · ↓ 4.1M · ♡ 6kOpen in visualizer mistralai/Voxtral-Mini-4B-Realtime-2602 architecture graphmistralai/Voxtral-Mini-4B-Realtime-2602automatic-speech-recognition · ↓ 1.9M · ♡ 988Open in visualizer ResembleAI/chatterbox architecture graphResembleAI/chatterboxtext-to-speech · ↓ 1.8M · ♡ 2kOpen in visualizer nvidia/Cosmos3-Edge architecture graphnvidia/Cosmos3-Edge↓ 1.5M · ♡ 208Open in visualizer circlestone-labs/Anima architecture graphcirclestone-labs/Anima↓ 1.2M · ♡ 2kOpen in visualizer unsloth/MiniMax-H3-GGUF architecture graphunsloth/MiniMax-H3-GGUFimage-text-to-video · ↓ 979k · ♡ 302Open in visualizer Lightricks/LTX-Video architecture graphLightricks/LTX-Videoimage-to-video · ↓ 769k · ♡ 2kOpen in visualizer Tongyi-MAI/Z-Image-Turbo architecture graphTongyi-MAI/Z-Image-Turbotext-to-image · ↓ 653k · ♡ 5kOpen in visualizer google/diffusiongemma-26B-A4B-it architecture graphgoogle/diffusiongemma-26B-A4B-itimage-text-to-text · ↓ 596k · ♡ 1kOpen in visualizer

View 208 catalog matches for Timestep embedding →

Related concepts

hfviewer renders the full architecture of 6,700+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →