LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #81 OF 228 ASSISTANTS · OVERALL #136 OF 6,618 · VIBE SCORE 6.1 · 195 VOICES
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
31 mentions
21 mentions
26 weeks · 325 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Qwen3.5 9B once and it's yours. No subscription, no rate limits, works offline.
9.65B parameters, a Q4_K_M GGUF is about 6.6 GB (FP16 full precision is about 19 GB)
NVIDIA GeForce RTX 3060 12GB (Desktop)
12 GB VRAM
~$230 USEDQ4_K_M quant, ~45 to 55 tok/s, 8k to 16k context entirely offloaded to GPU
Apple Mac mini M4 (24GB Unified Memory)
24 GB UNIFIED RAM
~$799 NEWQ8_0 quant or Q4 with large 64k+ context, ~35 to 45 tok/s, silent operation
NVIDIA GeForce RTX 4090 24GB (Desktop)
24 GB VRAM
~$1650 USEDFP16 unquantized or Q8_0 with full 262k context window offloaded to VRAM, 90+ tok/s
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Qwen3.5 9B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce RTX 3060 12GB (Desktop) | SWEET SPOTApple Mac mini M4 (24GB Unified Memory) | FULL POWERNVIDIA GeForce RTX 4090 24GB (Desktop) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~10 S★★★★★ | ~13 S★★★★★ | —★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~30 S★★★★★ | ~38 S★★★★★ | —★★★★★ |
| Build a websitea small landing page, markup and styles together | ~2.7 MIN★★★★★ | ~3.3 MIN★★★★★ | —★★★★★ |
| Build a backendan API with routes, storage and tests | ~8.3 MIN★★★★★ | ~10 MIN★★★★★ | —★★★★★ |
| Build a gamea playable browser game in one file | ~5 MIN★★★★★ | ~6.3 MIN★★★★★ | —★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. A DASH MEANS NO SPEED WAS RESEARCHED FOR THAT RIG. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
The internet is split on Qwen3.5 9B. Biggest praise: help with code. Biggest gripe: saying no too much.
162 POSTS · 33 COMMENTS · DEV FORUMS · LEMMY · GITHUB · REDDIT · DEV.TO · X · BLOG
4 thumbs up · 0 thumbs down
0 thumbs up · 4 thumbs down
Other models the crowd has fully reviewed, starting with multimodal models like this one.
Petition to add a rule for people to add their DAMN quant levels to their posts
I created a super harmful model ! :D (by tweaking it's J-Space!!!)
I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM
A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone
Use Qwen3.5 as an AI Assistant, Captioner or Image Analyzer inside of Comfyui!
I got tired of manually prompting every single clip for my AI music videos, so I built a 100% local open-source (LTX Video desktop + Gradio) app to automate it, meet - Synesthesia
Is it real qwen3.5 9B beat oss:120b?
My Local Setup for Agentic Sessions with Ollama + Qwen 3.5 9B
I built a free local video captioner specifically tuned for LTX-2.3 training —
Nanbeige4.2-3B drops: 3B params claiming to beat 9B/12B models on agentic tasks (atleast according to them)
Why have 8B-12B models been dropped?
Is Ling 3 tiny underrated for its size?