LOADING DATASHEET
LOADING DATASHEET
by Inclusionai
RANKED #94 OF 228 ASSISTANTS · OVERALL #162 OF 6,618 · VIBE SCORE 6.0 · 82 VOICES
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enablin
16 mentions
4 mentions
26 weeks · 153 voices
AGING WELL
The crowd is warmer now than it was at the start.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Ling 3.0 flash once and it's yours. No subscription, no rate limits, works offline.
124B parameters, a Q4 file is about 75 GB
PC with RTX 3090 (24GB) and 64GB RAM
24 GB VRAM + 64 GB RAM
~$850 USEDIQ2_M (~49 GB) with experts offloaded to CPU, ~5-10 tok/s, 8k context
Mac Studio M2 Ultra (128GB)
128 GB UNIFIED MEMORY
~$2,500 USEDAD-Q5_K_M (~89 GB), ~15-20 tok/s, 32k context
Mac Studio M4 Max (192GB)
192 GB UNIFIED MEMORY
~$6,600 REFURBISHEDQ8_0 (~136 GB) or Q6_K (~105 GB), full precision / longest context (up to 256K context)
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Ling 3.0 flash takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMPC with RTX 3090 (24GB) and 64GB RAM | SWEET SPOTMac Studio M2 Ultra (128GB) | FULL POWERMac Studio M4 Max (192GB) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~67 S | ~29 S | NO DATA |
| Summarize a documenta long report boiled down to the points that matter | ~3.3 MIN★★★★★ | ~86 S★★★★★ | —★★★★★ |
| Build a websitea small landing page, markup and styles together | ~18 MIN★★★★★ | ~7.6 MIN★★★★★ | —★★★★★ |
| Build a backendan API with routes, storage and tests | ~56 MIN★★★★★ | ~24 MIN★★★★★ | —★★★★★ |
| Build a gamea playable browser game in one file | ~33 MIN★★★★★ | ~14 MIN★★★★★ | —★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. A DASH MEANS NO SPEED WAS RESEARCHED FOR THAT RIG. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Ling 3.0 flash yet to call it. We found 82 posts but people did not say much either way.
74 POSTS · 8 COMMENTS · REDDIT · X · GITHUB · LEMMY · HACKER NEWS · BLUESKY
Checked picks first: tags say why you would switch, stars say how fully each one stands in for Ling 3.0 flash.
inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face
AntLing-3.0-flash is now live on OpenRouter, and free to use through August 3, 2026
inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8
AntLing’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages.
🧵 We’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, an
inclusionAI/Ling-3.0-flash · Hugging Face
Everyone posts day-one impressions. What's still in your stack a month later?
The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them.
ling 3.0 flash/tiny base models
124B total but only ~5B active-this is exactly the shape I want for my box. Are low-active MoEs just the local sweet spot now?
Ling-3.0 (BailingMoE3) lands in llama.cpp mainline - Quick benchmarks on Intel Arc B580
The executor slot doesn't need a smart model, it needs one that fails loudly