Heads & prediction

LM head (output projection)

Projects the final hidden state onto vocabulary logits. Often reusing the input embedding matrix (‘tied weights’).

First page of Using the Output Embedding to Improve Language ModelsIntroduced inAug 2016Using the Output Embedding to Improve Language ModelsPress & Wolf · arXiv 1608.05859 ↗
LM head: one matrix, read twicetied weights: tokens→vectors at the input, hidden→vocab logits at the outputembedding matrix (shared)'cat'vectorhidden statelogits over the vocab

The LM head is a single linear map from the model’s hidden width to one logit per vocabulary token; softmax over those logits gives next-token probabilities. Many models tie it to the input embedding matrix, saving hidden×vocab parameters (often several hundred million) at negligible quality cost. Common in small and mid-size models, less so at the largest scales.

Adoption over time

Share of new models that have included an LM head over time.

59%2022202320242025202659%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

google-t5/t5-small architecture graphgoogle-t5/t5-smalltranslation · ↓ 24.8M · ♡ 634Open in visualizer Qwen/Qwen3-0.6B architecture graphQwen/Qwen3-0.6Btext-generation · ↓ 23.7M · ♡ 2kOpen in visualizer amazon/chronos-2 architecture graphamazon/chronos-2time-series-forecasting · ↓ 22.5M · ♡ 484Open in visualizer FacebookAI/xlm-roberta-base architecture graphFacebookAI/xlm-roberta-basefill-mask · ↓ 21.0M · ♡ 923Open in visualizer Qwen/Qwen3-VL-8B-Instruct architecture graphQwen/Qwen3-VL-8B-Instructimage-text-to-text · ↓ 19.7M · ♡ 1kOpen in visualizer jonatasgrosman/wav2vec2-large-xlsr-53-japanese architecture graphjonatasgrosman/wav2vec2-large-xlsr-53-japaneseautomatic-speech-recognition · ↓ 17.7M · ♡ 87Open in visualizer openai-community/gpt2 architecture graphopenai-community/gpt2text-generation · ↓ 15.1M · ♡ 4kOpen in visualizer google/gemma-4-26B-A4B-it architecture graphgoogle/gemma-4-26B-A4B-itimage-text-to-text · ↓ 13.1M · ♡ 2kOpen in visualizer Qwen/Qwen3-8B architecture graphQwen/Qwen3-8Btext-generation · ↓ 12.9M · ♡ 2kOpen in visualizer Qwen/Qwen2.5-7B-Instruct architecture graphQwen/Qwen2.5-7B-Instructtext-generation · ↓ 9.7M · ♡ 2kOpen in visualizer Qwen/Qwen3.5-9B architecture graphQwen/Qwen3.5-9Bimage-text-to-text · ↓ 9.2M · ♡ 2kOpen in visualizer google/gemma-4-31B-it architecture graphgoogle/gemma-4-31B-itimage-text-to-text · ↓ 9.0M · ♡ 4kOpen in visualizer

View 4,058 catalog matches for LM head (output projection) →

Related concepts

hfviewer renders the full architecture of 6,800+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →