cerebras/GLM-4.7-Flash-REAP-23B-A3B neural network architecture graph

hfviewer renders an interactive architecture graph for the Hugging Face model cerebras/GLM-4.7-Flash-REAP-23B-A3B. The graph is built from the model's real structure: nodes carry source-faithful module, class, and operation names (embeddings, attention blocks, feed-forward layers, normalization, task heads), edges follow the actual forward dataflow, and repeated blocks are grouped with their true repeat counts. Where the model could be executed, the graph is trace-backed; otherwise it is derived from the reviewed configuration and source code.

cerebras/GLM-4.7-Flash-REAP-23B-A3B architecture at a glance

cerebras/GLM-4.7-Flash-REAP-23B-A3B is a text-generation model (glm4_moe_lite architecture). It uses 47 transformer layers, a hidden size of 2,048, and multi-head latent attention (MLA). The feed-forward layers use a mixture-of-experts design with 48 experts (4 active per token, 1 always-on). It has a vocabulary of 154,880 tokens and a context window of up to 202,752 tokens, in bfloat16 precision.

Model type
glm4_moe_lite
Layers
47
Hidden size
2,048
Attention
multi-head latent attention (MLA)
Feed-forward
Mixture-of-experts design with 48 experts (4 active per token, 1 always-on)
Context length
202,752 tokens
Vocabulary
154,880 tokens
Precision
bfloat16

Task: text-generation.

Architecture components in this graph, each explained in the hfviewer glossary: Gated MLP (SwiGLU) · LM head (output projection) · MoE experts · MoE router · Multi-head latent attention (MLA) · Residual (skip) connection · RMSNorm · Self-attention · Shared expert.

Related architecture graphs on hfviewer: zai-org/GLM-4.7-Flash · cyankiwi/GLM-4.7-Flash-AWQ-4bit · unsloth/GLM-4.7-Flash · lmstudio-community/GLM-4.7-Flash-MLX-8bit · lmstudio-community/GLM-4.7-Flash-MLX-6bit · GadflyII/GLM-4.7-Flash-NVFP4.

Browse more graphs from cerebras or explore other models on the hfviewer home page.

Interactive model architecture

Architecture graph for cerebras/GLM-4.7-Flash-REAP-23B-A3B.

Interactive architecture graph for cerebras/GLM-4.7-Flash-REAP-23B-A3B, visualized from Hugging Face model metadata.

Paste a Hugging Face link to visualize it
No export step, no config hunt, no model surgery. Paste the link and inspect the graph.
Graph structure Understand the high-level graph structure of different transformer models. Quickstart guide
URL magic You can replace huggingface.co with hfviewer.com in the url to view it.
Chrome Extension! View each model directly on Hugging Face! Install extension
Embed in model card Embed the architecture graph directly in your Hugging Face model card. Add to your model card!
Granularity Block
Zoom into
Community showcase

Featureyour model

Embed the visualization in your model card (README.md) and get it featured in the Community showcase.

Article moderation

Review reports.

Triage reported model articles and comments, hide abusive content, and resolve cases.

Sign in with the HannesVonEssen Hugging Face account to review reports.

No moderation reports match this filter.

Interactive article

Choose a model for your article.

Search ready hfviewer graphs or pick one of your public Hugging Face model repos. Ready models open the editor immediately; owned models that are not viewable yet can be generated first.

Owner article If the model belongs to your Hugging Face user or organization namespace, it is labeled as an owner article.
Community article If you write about another public model, it is labeled as a community article next to the graph.

Loading models...

Release watchlist.

Watch your favorite Hugging Face orgs and get an email the moment a new release's architecture graph is ready in hfviewer.

Editor - interactive article

Loading editor...

MODEL PAGES WITH HFVIEWER

Community showcase.

Add to your model card!

Hugging Face authors are adding the hfviewer model card directly to their READMEs.

If you are interested in deploying these models to edge devices, check out our other products: