LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #106 OF 228 ASSISTANTS · OVERALL #178 OF 6,618 · VIBE SCORE 5.9 · 119 VOICES
Qwen3.5 4B is a compact open-weight multimodal model from Alibaba for reasoning, coding, visual understanding, tool use, and structured output.
12 mentions
9 mentions
26 weeks · 184 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Qwen3.5 4B once and it's yours. No subscription, no rate limits, works offline.
4B parameters, a Q4 file is about 2.7 GB and full BF16 is around 8.5 GB
Nvidia GTX 1660 Super
6 GB VRAM
~$110 USEDQ4_K_M, ~35 tok/s, 8k context
Nvidia RTX 4060
8 GB VRAM
~$299 NEWQ8_0 or full BF16, ~75 tok/s, 32k context
Nvidia RTX 4070 Super
12 GB VRAM
~$590 NEWBF16 unquantized, ~110 tok/s, full extended context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Qwen3.5 4B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNvidia GTX 1660 Super | SWEET SPOTNvidia RTX 4060 | FULL POWERNvidia RTX 4070 Super |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~14 S★★★★★ | ~7 S★★★★★ | ~5 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~43 S★★★★★ | ~20 S★★★★★ | ~14 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~3.8 MIN★★★★★ | ~1.8 MIN★★★★★ | ~73 S★★★★★ |
| Build a backendan API with routes, storage and tests | ~12 MIN★★★★★ | ~5.6 MIN★★★★★ | ~3.8 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~7.1 MIN★★★★★ | ~3.3 MIN★★★★★ | ~2.3 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly happy with Qwen3.5 4B. Biggest praise: help with code.
101 POSTS · 18 COMMENTS · REDDIT · GITHUB · BLUESKY · X · LEMMY · DEV.TO · DEV FORUMS · BLOG
2 thumbs up · 2 thumbs down
4 thumbs up · 0 thumbs down
3 thumbs up · 0 thumbs down
Other models the crowd has fully reviewed, starting with multimodal models like this one.
I built an open-source RAG system that actually understands images, tables, and document structure — not just text chunks
Qwen3.5-4B-Base-ZitGen-V1
[P] TurboQuant for weights: near‑optimal 4‑bit LLM quantization with lossless 8‑bit residual – 3.2× memory savings
NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction
tencent/EVIE-Preview-4.5B · Hugging Face
NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction (self-hostable) [P]
Many people reached out to me in the past asking about my local agent stack as well as how I set up my local agent stack. So, I thought it might be useful to put together a little tutorial on how to set up a local (coding) agent using open-source tools and…
would you laugh at me if I ran gemma-4-26b on a 4 core Xeon, with 32GB RAM, no GPU?
you can just watch a language model think now. i built a way to visualize the words AI doesn’t say
[Study/Models] Flint: Compressing Reasoning Without Breaking It
Agents-A1-4B (Qwen3.7-4B ???) : Scaling the Horizon, Not the Parameters
Smallest coding model possible