LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #25 OF 266 ASSISTANTS · OVERALL #38 OF 6,544 · VIBE SCORE 6.9 · 649 VOICES
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
93 mentions
37 mentions
26 weeks · 6,280 voices
AGING WELL
The crowd is warmer now than it was at the start.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Qwen3.8 Flash once and it's yours. No subscription, no rate limits, works offline.
125B MoE with 6B active parameters plus 51B n-gram embeddings, a Q4 quant takes about 94 GB
Apple Mac Studio M2 Max (96GB Unified Memory)
96 GB UNIFIED MEMORY
~$2,200 USEDQ2 quant or Q3 with n-gram offloading, ~10 tok/s, 8k context
Apple Mac Studio M2 Ultra (192GB Unified Memory)
192 GB UNIFIED MEMORY
~$4,500 USEDQ4_K_M quant, ~22 tok/s, 32k context
Workstation with 4x NVIDIA GeForce RTX 3090 24GB
96 GB VRAM AND 128 GB SYSTEM RAM
~$3,800 USEDFP8 / NVFP4 quant with host embedding offloading, ~42 tok/s, 64k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Qwen3.8 Flash takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMApple Mac Studio M2 Max (96GB Unified Memory) | SWEET SPOTApple Mac Studio M2 Ultra (192GB Unified Memory) | FULL POWERWorkstation with 4x NVIDIA GeForce RTX 3090 24GB |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~50 S★★★★★ | ~23 S★★★★★ | ~12 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~2.5 MIN★★★★★ | ~68 S★★★★★ | ~36 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~13 MIN★★★★★ | ~6.1 MIN★★★★★ | ~3.2 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~42 MIN★★★★★ | ~19 MIN★★★★★ | ~9.9 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~25 MIN★★★★★ | ~11 MIN★★★★★ | ~6 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly frustrated with Qwen3.8 Flash right now. Biggest praise: help with code. Biggest gripe: working when you need it.
513 POSTS · 136 COMMENTS · REDDIT · BLUESKY · LEMMY · GITHUB · BLOG · HACKER NEWS · DEV.TO
2 thumbs up · 8 thumbs down
4 thumbs up · 1 thumbs down
1 thumbs up · 2 thumbs down
“Qwen 3.8 Flash Next (qwen4exp): fatal broadcastshapes crash at step 2 of any tool run — QSA indexer budget vs KV length.”
Checked picks first: tags say why you would switch, stars say how fully each one stands in for Qwen3.8 Flash.
Qwen3.8-Flash-Next announced
Run Qwen3.8-Flash-Next locally! 🔥
Qwen releases Qwen3.8-Flash-Next!
All Qwen3.8-Flash-Next quants are up!
Run Qwen3.8-Flash-Next GGUF 1.7x Faster!
Qwen3.8-Flash-Next now available in Unsloth!
Qwen 3.8 Flash Next: Beating DS V4 Flash at half the parameters, stronger than Opus 4.6
Qwen3.8-Flash-Next 176B on a 16GB card: yes it works, here's how
New Unsloth Release - much faster performance + new model support
MIND BLOWING - QWEN3.8-FLASH-NEXT - RTX 3090 + 128DDR5 +/-20tok/s
Qwen3.8-Next-Flash MTP Models Uploaded??
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s