LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #34 OF 267 ASSISTANTS · OVERALL #51 OF 6,566 · VIBE SCORE 6.8 · 550 VOICES
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
96 mentions
22 mentions
26 weeks · 789 voices
HONEYMOON FADING
Early praise is cooling off in recent weeks.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Qwen3.6 35B A3B once and it's yours. No subscription, no rate limits, works offline.
35B parameters (3B active per token), a Q4 file is about 18 GB
NVIDIA RTX 4060 Ti 16GB
16 GB VRAM
~$390 NEWIQ3_S or Q3_K_S quant, ~15-20 tok/s, 16k context (requires partial CPU offload)
NVIDIA RTX 3090 24GB
24 GB VRAM
~$800 USEDIQ4_XS or Q4_K_M quant, ~30-45 tok/s, 32k context (fully in VRAM)
Apple Mac Studio M2 Ultra (128GB)
128 GB UNIFIED MEMORY
~$2,500 USEDBF16 (unquantized) or Q8_0, ~20-30 tok/s, up to 262k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Qwen3.6 35B A3B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA RTX 4060 Ti 16GB | SWEET SPOTNVIDIA RTX 3090 24GB | FULL POWERApple Mac Studio M2 Ultra (128GB) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~29 S★★★★★ | ~13 S★★★★★ | ~20 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~86 S★★★★★ | ~40 S★★★★★ | ~60 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~7.6 MIN★★★★★ | ~3.6 MIN★★★★★ | ~5.3 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~24 MIN★★★★★ | ~11 MIN★★★★★ | ~17 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~14 MIN★★★★★ | ~6.7 MIN★★★★★ | ~10 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
The internet is split on Qwen3.6 35B A3B. Biggest praise: help with code. Biggest gripe: saying no too much.
424 POSTS · 126 COMMENTS · REDDIT · HACKER NEWS · GITHUB · X · LEMMY · DEV FORUMS · BLOG
0 thumbs up · 10 thumbs down
8 thumbs up · 0 thumbs down
1 thumbs up · 6 thumbs down
5 thumbs up · 1 thumbs down
2 thumbs up · 3 thumbs down
2 thumbs up · 1 thumbs down
0 thumbs up · 3 thumbs down
Checked picks first: tags say why you would switch, stars say how fully each one stands in for Qwen3.6 35B A3B.
Qwen3.6 35B-A3B (Q8_0, no KV quant) single prompt in opencode: "Create a beautiful, relaxing flight simulator in a single html file with mountains, clouds, and endless procedural terrain"
Qwen3.6-35B-A3B: Agentic coding power, now open to all
Qwen3.6 is out now!
2.5x faster Qwen3.6 NVFP4 Unsloth quants
Run Qwen3.6 MTP GGUFs locally!
Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7
Local AI News You Missed - April 2026
2-bit Qwen3.6-35B-A3B GGUF is amazing! Made 30+ successful tool calls
Qwen3.6 MTP Unsloth Experimental GGUFs
Last week in Generative Image & Video
A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM
I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM