A compact 126M language model built with MEGA rather than a standard transformer stack.
Give your model card an architecture graph in 10 seconds.
Paste your model ID → copy HTML
→ paste into README.md.
A larger, modern Constellation-One variant fine-tuned from ModernBERT-large.
A Gemma 4 E2B based conversational model tuned toward a more direct, personal voice.
A DistilRoBERTa-based OME v5.2 emotion classifier for English text, fine-tuned over 26 categories and reporting 98.03 percent eval accuracy.
HTML
Preview
Copied! Click to Edit README.md
Community showcase.
Hugging Face authors are adding hfviewer architecture cards directly to their READMEs.
A compact 155M-parameter latent MMDiT image model pairing frozen FLAN-T5 and CLIP conditioning with 16 dual-stream joint-attention blocks, then sampling through 50-step rectified flow before SDXL VAE decode.
A tiny text-to-3D diffusion model that turns prompts into 32x32x32 voxel occupancy grids with rectified-flow sampling. The PixelModel line stepping out of 2D: 40M parameters, a 12-layer DiT over 3D patches, trained on 28k meshes in under ten hours on a single A100.
A 97.8M-parameter causal language model built on the custom Rose X1 architecture and trained on roughly 80 billion tokens with a native 2048-token context window. The scaled-up sequel to Rose-Mini from one of our most prolific community members.
A 56M-parameter, from-scratch recurrent language model combining a per-channel leaky integrator (FWKV) with the RWKV-8 ROSA copy-signal mechanism across 14 stacked blocks. No attention anywhere; chat-tuned on ultrachat_200k.
A compact text-to-image diffusion transformer that replaces standard self-attention with bidirectional FWKV token mixing over 256 latent patches. It generates 256x256 images with frozen CLIP text conditioning, adaLN-zero modulation, and a rectified-flow VAE latent decoder.
An experimental 169.5M-parameter diffusion language model built by converting Supra-1.5-50M-Base-exp from autoregressive generation, then expanding it through layer duplication. Its 16 timestep-conditioned, bidirectional GQA blocks iteratively denoise masked tokens across a 5,120-token context.
A 24M-parameter English base model trained from scratch, the second release in IvmeLabsโ Ivme family (Turkish for โaccelerationโ) of tiny models built to punch above their weight, this time on a much heavier, topic-focused data diet.
A ~50M-parameter Python code model trained from scratch on filtered StarCoder data, part of IvmeLabsโ Ivme family (Turkish for โaccelerationโ) of tiny models built to punch above their weight.
An experimental 130M-parameter masked discrete-diffusion language model derived from IvmeLabs' bidirectional transformer family. Instead of next-token prediction, it denoises masked tokens with full-context attention to explore small dSLM generation.
A 40M-parameter text-to-image diffusion transformer that generates 256x256 images with rectified-flow sampling. Version 5 keeps the v4 architecture but trains on a much larger recaptioned CC12M image-caption set.
A 1.4B nanochat-style chat model in the Karpathy lineage, trained end to end on a single RTX 5090.
A Qwen3.5-family multimodal text-generation model from the XORTRON Criminal Computing project, fine-tuned from a DavidAU Qwen3.6 27B Fable Fusion base as an AI-safety and alignment research experiment.
A GGUF release of darkc0de/XORTRON-NXTXPRT9PRO-27B, a Qwen3.5-family multimodal model from the XORTRON Criminal Computing project. The model card frames the project as an AI-safety and alignment research experiment.
An experimental SpikeWhale base language model for compact text generation, using custom Transformers code with MLA, JEPA-inspired design, and two-expert routing.
A Jet-Long edition of Escarda-86M-Base, an ~86M-parameter from-scratch SpikeWhaleLM decoder-only language model and JEPA-distilled base checkpoint for continued pretraining or fine-tuning. It extends the usable context from 4K to 10K tokens with dynamic bifocal RoPE while leaving short-context behavior unchanged.
An action-conditioned world model for Counter-Strike 2 that takes keyboard and mouse actions and imagines the next frames. It adapts MIRA around a DINOv3-backed video codec and a small latent diffusion transformer trained for action-responsive rollouts.
A companion SpikeWhale base model for small text-generation experiments.
A Jet-Long edition of Byrne-86M-Base, an ~86M-parameter SpikeWhaleLM decoder for compact text generation and continued fine-tuning. It extends the usable context from 4K to 10K tokens with dynamic bifocal RoPE while preserving short-context behavior.
A from-scratch spiking language model: four stacked leaky integrate-and-fire neuron layers trained with surrogate gradients.
A 144M-parameter causal LM whose channel mixing runs on Kuramoto oscillators instead of a plain MLP, on a conventional transformer backbone.
A 79M-parameter research model whose channel mixing is a differentiable neighbour-sensing layer inspired by fungal colonies.
A 64M Quazimoto language model with MLA attention, Chimera Kuramoto/Growth/Wave cores, and top-2 routed MoE experts.
RobinsonLabs includes architecture links to the hfviewer graph for each of their model families.
A BrtGPT conversational text-generation checkpoint trained on LaMini-instruction, with code and math evaluations highlighted in the model card.
A physics-guided ultrasound segmentation suite covering UNet, UNet++, SegFormer, SwinUnet, TransUNet, and style-adaptation checkpoints.
A compact CALI causal language model with custom blocks and evaluation-tracker coverage across ARC, HellaSwag, MMLU, TruthfulQA, and WinoGrande.
An 80M MiniGPT-style causal language model using the Qwen2.5 tokenizer, 12 decoder blocks, causal attention, and a tied language-model head.
A 125M custom GPT-X2 language model with RoPE, SwiGLU, grouped-query attention, and curriculum-trained code/math normalization.
A compact 5M GPT-S language model using RoPE, SwiGLU, grouped-query attention, and exclusive shared attention for small-model reasoning experiments.
A MIT-licensed Ettin encoder fine-tuned for PII token classification on Nemotron-PII, reporting 96.27 F1 on the test split.
A multilingual Echo-DSRN intent classifier fine-tuned on Amazon MASSIVE, built from the recurrent Echo-DSRN embedding model for 60 text-intent labels.
A fine-tuned SAMUS checkpoint for ultrasound image segmentation, adapting Segment Anything with a ViT-B image encoder, prompt encoder, and mask decoder.
A 74M decoder-only Llama-style language model pretrained from scratch on 1.23B tokens, with eight layers and grouped-query attention.
A 49.4M-parameter text-generation model built on Rose X1, GODELEV's custom lightweight decoder-only architecture for efficient small language models.
A compact 126M language model built with MEGA rather than a standard transformer stack, with a 4096-token context length.
A larger, more modern Constellation-One variant for Cockatoo, fine-tuned from answerdotai/ModernBERT-large.
A Pegasus-X based long-document summarization model fine-tuned on synthetic summaries for general summarization with long contexts.
A Gemma 4 E2B based conversational model tuned toward a more direct, personal voice rather than a generic assistant cadence.
A BERT based OME v5 classifier for English emotion examples, fine-tuned over 26 categories and reporting 98.63 percent accuracy on eval.
A DistilRoBERTa-based OME v5.2 emotion classifier for English text, fine-tuned over 26 categories and reporting 98.03 percent eval accuracy.
A LoRA fine-tune of ModernBERT-base for binary classification, designed for high recall with controlled false positives.