LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #136 OF 228 ASSISTANTS · OVERALL #220 OF 6,618 · VIBE SCORE 5.7 · 211 VOICES
qwen3-4b-heretic on Hugging Face (text generation). 20,121 downloads. Open weights for local or hosted use.
32 mentions
17 mentions
26 weeks · 204 voices
AGING WELL
The crowd is warmer now than it was at the start.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download qwen3 4B once and it's yours. No subscription, no rate limits, works offline.
4B parameters, a Q4 file is about 2.5 GB, full FP16 is about 8 GB
NVIDIA GeForce GTX 1660 Super
6 GB VRAM
~$110 USEDQ4_K_M, ~35 tok/s, 8k context
NVIDIA GeForce RTX 4060
8 GB VRAM
~$290 NEWQ8_0 or FP16, ~75 tok/s, 32k context
NVIDIA GeForce RTX 4070
12 GB VRAM
~$530 NEWFP16, ~110 tok/s, full 128k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long qwen3 4B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce GTX 1660 Super | SWEET SPOTNVIDIA GeForce RTX 4060 | FULL POWERNVIDIA GeForce RTX 4070 |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~14 S★★★★★ | ~7 S★★★★★ | ~5 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~43 S★★★★★ | ~20 S★★★★★ | ~14 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~3.8 MIN★★★★★ | ~1.8 MIN★★★★★ | ~73 S★★★★★ |
| Build a backendan API with routes, storage and tests | ~12 MIN★★★★★ | ~5.6 MIN★★★★★ | ~3.8 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~7.1 MIN★★★★★ | ~3.3 MIN★★★★★ | ~2.3 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about qwen3 4B yet to call it. We found 211 posts but people did not say much either way.
170 POSTS · 41 COMMENTS · GITHUB · REDDIT · HACKER NEWS · DEV FORUMS · X · DEV.TO · STACK OVERFLOW · LEMMY · BLOG
Checked picks first: tags say why you would switch, stars say how fully each one stands in for qwen3 4B.
Z-Image styles: 70 examples of how much can be done with just prompting.
Playing with z image turbo on an idle day
I successfully replaced CLIP with an LLM for SDXL
I tested 10 LLMs locally on my MacBook Air M1 (8GB RAM!) – Here's what actually works-
If You Already Pay for an LLM Service, Running Local Embeddings and Rerankers Feels More Useful Than Running Local LLMs
Z image turbo (Low vram workflow) GGUF
You can now run Qwen's new Qwen3 model on your own local device! (10GB RAM min.)
Z-Image + Qwen3 4b: The abliterated text encoder debate is pure vibes. I measured it. Here are the numbers - Abliterlitics
💻 I optimized Qwen3:30B MoE to run on my RTX 3070 laptop at ~24 tok/s - full breakdown inside
[WIP] Still experimenting, but the next Z-Image Power Nodes will have no limits!!
Qwen3:4b Too Many Model thoughts to respond to a simple "hi"
Conditioning Enhancer (Qwen/Z-Image): Post-Encode MLP & Self-Attention Refiner