LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #112 OF 228 ASSISTANTS · OVERALL #185 OF 6,618 · VIBE SCORE 5.9 · 62 VOICES
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
14 mentions
4 mentions
26 weeks · 86 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Qwen3.5 122B A10B once and it's yours. No subscription, no rate limits, works offline.
122B total parameters (10B active MoE), a Q4_K_M quant file is about 72 GB
Apple Mac Studio M2 Max (96GB unified memory)
96 GB UNIFIED MEMORY
~$2100 USEDQ3_K_M quant, ~14 tok/s, 16k context with unified memory
Dual NVIDIA GeForce RTX 3090 (2x24GB) desktop with 64GB RAM
48 GB VRAM PLUS 64 GB RAM
~$1750 USEDQ4_K_M quant with KTransformers offloading, ~22 tok/s, 32k context
Workstation with 4x NVIDIA GeForce RTX 4090 24GB
96 GB VRAM
~$7800 NEWFP8 / Q8 quant fully in VRAM, ~45 tok/s, 128k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Qwen3.5 122B A10B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMApple Mac Studio M2 Max (96GB unified memory) | SWEET SPOTDual NVIDIA GeForce RTX 3090 (2x24GB) desktop with 64GB RAM | FULL POWERWorkstation with 4x NVIDIA GeForce RTX 4090 24GB |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~36 S | ~23 S | ~11 S |
| Summarize a documenta long report boiled down to the points that matter | ~1.8 MIN | ~68 S | ~33 S |
| Build a websitea small landing page, markup and styles together | ~9.5 MIN★★★★★ | ~6.1 MIN★★★★★ | ~3 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~30 MIN★★★★★ | ~19 MIN★★★★★ | ~9.3 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~18 MIN★★★★★ | ~11 MIN★★★★★ | ~5.6 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly frustrated with Qwen3.5 122B A10B right now. Biggest gripe: saying no too much.
39 POSTS · 23 COMMENTS · REDDIT · HACKER NEWS · LEMMY · GITHUB · DEV FORUMS · X · BLUESKY
0 thumbs up · 6 thumbs down
0 thumbs up · 3 thumbs down
Other models the crowd has fully reviewed, starting with multimodal models like this one.
AI News You Missed - March 2026
Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM
Qwen just released Qwen 3.5 medium model Series: Qwen 3.5 Flash plus 3 more models
Ideogram 4 compared with Krea 2 in natural language prompting
Fixed three bugs that made Qwen3.5-122B a daily driver on Mac Studio
Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark
NVFP4 still isn't faster than FP8 on Blackwell (SM120) - some numbers from Qwen3.6-27B
All Qwen model oneshots: 1109 outputs to look at and compare!
[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarms
Real-world reality check on Qwen for autonomous coding agents
Professional-grade local AI on consumer hardware — 80B stable on 44GB mixed VRAM (RTX 5060 Ti ×2 + RTX 3060) for under €800 total. Full compatibility matrix included.