Multimodal

Multimodal projector

A small MLP that translates vision-encoder features into the language model’s embedding space.

First page of Visual Instruction TuningPopularized inApr 2023Visual Instruction TuningLiu et al. · arXiv 2304.08485 ↗
Multimodal projector: the small bridgeone or two linear layers map vision features into the LLM’s embedding spaceMLPvision featuresLLM-space tokens

The projector (often just one or two linear layers) maps visual features into vectors that look like text-token embeddings to the LLM, which then attends over them like any other tokens. It is tiny compared to the towers it connects but is where ‘alignment’ between modalities is learned. Many multimodal training recipes tune the projector first, alone.

Adoption over time

Share of new models that have included a multimodal projector over time.

1%202220232024202520261%20222023202420252026

See it in real models

Open any of these on hfviewer to find this block in the interactive architecture graph.

microsoft/Florence-2-base architecture graphmicrosoft/Florence-2-baseimage-text-to-text · ↓ 3.0M · ♡ 400Open in visualizer moonshotai/Kimi-K3 architecture graphmoonshotai/Kimi-K3image-text-to-text · ↓ 2.0M · ♡ 11kOpen in visualizer mistralai/Voxtral-Mini-4B-Realtime-2602 architecture graphmistralai/Voxtral-Mini-4B-Realtime-2602automatic-speech-recognition · ↓ 1.9M · ♡ 988Open in visualizer google/gemma-3-4b-it architecture graphgoogle/gemma-3-4b-itimage-text-to-text · ↓ 1.8M · ♡ 2kOpen in visualizer llava-hf/llava-1.5-7b-hf architecture graphllava-hf/llava-1.5-7b-hfimage-text-to-text · ↓ 1.7M · ♡ 375Open in visualizer IDEA-Research/grounding-dino-base architecture graphIDEA-Research/grounding-dino-basezero-shot-object-detection · ↓ 1.5M · ♡ 206Open in visualizer google/medgemma-4b-it architecture graphgoogle/medgemma-4b-itimage-text-to-text · ↓ 1.1M · ♡ 1kOpen in visualizer IDEA-Research/grounding-dino-tiny architecture graphIDEA-Research/grounding-dino-tinyzero-shot-object-detection · ↓ 765k · ♡ 116Open in visualizer Infomaniak-AI/vllm-translategemma-4b-it architecture graphInfomaniak-AI/vllm-translategemma-4b-itimage-text-to-text · ↓ 746k · ♡ 14Open in visualizer ISTA-DASLab/gemma-3-27b-it-GPTQ-4b-128g architecture graphISTA-DASLab/gemma-3-27b-it-GPTQ-4b-128gimage-text-to-text · ↓ 714k · ♡ 45Open in visualizer google/gemma-3-12b-it architecture graphgoogle/gemma-3-12b-itimage-text-to-text · ↓ 603k · ♡ 837Open in visualizer microsoft/Florence-2-large architecture graphmicrosoft/Florence-2-largeimage-text-to-text · ↓ 555k · ♡ 2kOpen in visualizer

View 230 catalog matches for Multimodal projector →

Related concepts

hfviewer renders the full architecture of 6,700+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.

Browse all model graphs →