Audio

Conformer block

A transformer block with a convolution module inside. Attention for global context, convs for local acoustic detail.

First page of Conformer: Convolution-augmented Transformer for Speech RecognitionIntroduced inMay 2020Conformer: Convolution-augmented Transformer for Speech RecognitionGulati et al. · arXiv 2005.08100 ↗
Conformer: attention and convolution, togetherattention hears the whole utterance · the conv module keeps the local detailfeed-forward (½)self-attention · globalconvolution · local detailfeed-forward (½)

The Conformer sandwiches a depthwise-convolution module between attention and feed-forward layers (in a macaron FFN-attention-conv-FFN pattern), usually with relative position encoding in the attention. Convolutions capture local spectral patterns that pure attention handles poorly, which made Conformers the dominant speech-encoder architecture (NVIDIA’s FastConformer here).

Adoption over time

Share of new models that have included Conformer blocks over time.

3%202220232024202520263%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

google/gemma-4-E4B-it architecture graphgoogle/gemma-4-E4B-itany-to-any · ↓ 4.4M · ♡ 2kOpen in visualizer ResembleAI/chatterbox architecture graphResembleAI/chatterboxtext-to-speech · ↓ 1.8M · ♡ 2kOpen in visualizer nvidia/nemotron-3.5-asr-streaming-0.6b architecture graphnvidia/nemotron-3.5-asr-streaming-0.6bautomatic-speech-recognition · ↓ 752k · ♡ 1kOpen in visualizer nvidia/parakeet-tdt-0.6b-v3 architecture graphnvidia/parakeet-tdt-0.6b-v3automatic-speech-recognition · ↓ 577k · ♡ 1kOpen in visualizer nvidia/nemotron-speech-streaming-en-0.6b architecture graphnvidia/nemotron-speech-streaming-en-0.6bautomatic-speech-recognition · ↓ 419k · ♡ 624Open in visualizer ai-sage/GigaAM-v3 architecture graphai-sage/GigaAM-v3automatic-speech-recognition · ↓ 414k · ♡ 163Open in visualizer facebook/seamless-m4t-v2-large architecture graphfacebook/seamless-m4t-v2-largeautomatic-speech-recognition · ↓ 306k · ♡ 1kOpen in visualizer ibm-granite/granite-speech-4.1-2b architecture graphibm-granite/granite-speech-4.1-2bautomatic-speech-recognition · ↓ 236k · ♡ 162Open in visualizer CohereLabs/cohere-transcribe-03-2026 architecture graphCohereLabs/cohere-transcribe-03-2026automatic-speech-recognition · ↓ 207k · ♡ 1kOpen in visualizer google/gemma-3n-E2B-it architecture graphgoogle/gemma-3n-E2B-itimage-text-to-text · ↓ 171k · ♡ 327Open in visualizer ibm-granite/granite-speech-4.1-2b-plus architecture graphibm-granite/granite-speech-4.1-2b-plusautomatic-speech-recognition · ↓ 132k · ♡ 92Open in visualizer nvidia/canary-1b-v2 architecture graphnvidia/canary-1b-v2automatic-speech-recognition · ↓ 47k · ♡ 423Open in visualizer

View 58 catalog matches for Conformer block →

Related concepts

hfviewer renders the full architecture of 6,700+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →