LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #97 OF 228 ASSISTANTS · OVERALL #167 OF 6,618 · VIBE SCORE 6.0 · 72 VOICES
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
22 mentions
4 mentions
26 weeks · 94 voices
AGING WELL
The crowd is warmer now than it was at the start.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Qwen3.5 35B A3B once and it's yours. No subscription, no rate limits, works offline.
35B total parameters with 3B active per token (MoE), a Q4 file is about 20 GB
Nvidia RTX 4060 Ti 16GB
16 GB VRAM + 32 GB SYSTEM RAM
~$450Q3_K_M or Q4 with partial RAM offload, ~15 tok/s, 8k context
Nvidia RTX 3090 24GB
24 GB VRAM
~$750 USEDQ4_K_M fully in VRAM, ~45 tok/s, 32k context
Apple Mac Studio M2 Ultra
128 GB UNIFIED MEMORY
~$3500Q8 or BF16, ~65 tok/s, 128k+ context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Qwen3.5 35B A3B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNvidia RTX 4060 Ti 16GB | SWEET SPOTNvidia RTX 3090 24GB | FULL POWERApple Mac Studio M2 Ultra |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~33 S | ~11 S | ~8 S |
| Summarize a documenta long report boiled down to the points that matter | ~1.7 MIN★★★★★ | ~33 S★★★★★ | ~23 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~8.9 MIN★★★★★ | ~3 MIN★★★★★ | ~2.1 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~28 MIN★★★★★ | ~9.3 MIN★★★★★ | ~6.4 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~17 MIN★★★★★ | ~5.6 MIN★★★★★ | ~3.8 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly frustrated with Qwen3.5 35B A3B right now.
53 POSTS · 19 COMMENTS · DEV FORUMS · X · HACKER NEWS · LEMMY · REDDIT · GITHUB
0 thumbs up · 4 thumbs down
Other models the crowd has fully reviewed, starting with multimodal models like this one.
Has anyone created a "Local LLM Survival Kit"?
Qwen just released Qwen 3.5 medium model Series: Qwen 3.5 Flash plus 3 more models
Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark
Open WebUI “terminal-aware” skills are scary powerful. I made a skill-building workflow that seems to work well for developing them.
Qwen3.5 27B Uncensored Heretic Native MTP Preserved is Out Now With the Full 15 MTPs Preserved and Retained, Available in Safetensors, GGUFs, NVFP4, NVFP4 GGUFs and GPTQ-Int4 Formats!
Qwen3.5 27B Uncensored Heretic Native MTP Preserved is Out Now With the Full 15 MTPs Preserved and Retained, Available in Safetensors, GGUFs, NVFP4, NVFP4 GGUFs and GPTQ-Int4 Formats!
All Qwen model oneshots: 1109 outputs to look at and compare!
Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti
Built an open-source one-prompt-to-cinematic-reel pipeline on a single GPU — FLUX.2 [klein] for character keyframes, Wan2.2-I2V for animation, vision critic with auto-retry, music + 9-language narration in the same pipeline
Built an open-source one-prompt-to-cinematic-reel pipeline on a single GPU — FLUX.2 [klein] for character keyframes, Wan2.2-I2V for animation, vision critic with auto-retry, music + 9-language narration in the same pipeline
Professional-grade local AI on consumer hardware — 80B stable on 44GB mixed VRAM (RTX 5060 Ti ×2 + RTX 3060) for under €800 total. Full compatibility matrix included.
Has anyone got qwen3.5 to work with ollama?