LOADING DATASHEET
LOADING DATASHEET
by Alibaba
RANKED #120 OF 271 ASSISTANTS · OVERALL #202 OF 6,567 · VIBE SCORE 5.9 · 74 VOICES
Qwen3.5 0.8B is a lightweight open-weight multimodal model from Alibaba for fast reasoning, visual understanding, tool use, and JSON output.
8 mentions
3 mentions
26 weeks · 109 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Qwen3.5 0.8B once and it's yours. No subscription, no rate limits, works offline.
0.8B parameters, a Q4 file is about 0.6 GB (BF16 is ~1.8 GB)
Raspberry Pi 5 (4GB RAM)
4 GB RAM
~$60 NEWQ4_K_M on CPU, ~25 tok/s, 8k context
Nvidia GeForce GTX 1650 (4GB VRAM)
4 GB VRAM
~$80 USEDQ8 or FP16 fully offloaded to VRAM, ~140 tok/s, 32k context
Nvidia GeForce RTX 3060 (12GB VRAM)
12 GB VRAM
~$280 NEWBF16 / FP16, max native 262k context window, ~260 tok/s
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Qwen3.5 0.8B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMRaspberry Pi 5 (4GB RAM) | SWEET SPOTNvidia GeForce GTX 1650 (4GB VRAM) | FULL POWERNvidia GeForce RTX 3060 (12GB VRAM) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~20 S | ~4 S | ~2 S |
| Summarize a documenta long report boiled down to the points that matter | ~60 S★★★★★ | ~11 S★★★★★ | ~6 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~5.3 MIN★★★★★ | ~57 S★★★★★ | ~31 S★★★★★ |
| Build a backendan API with routes, storage and tests | ~17 MIN★★★★★ | ~3 MIN★★★★★ | ~1.6 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~10 MIN★★★★★ | ~1.8 MIN★★★★★ | ~58 S★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Qwen3.5 0.8B yet to call it. We found 74 posts but people did not say much either way.
72 POSTS · 2 COMMENTS · REDDIT · GITHUB · LEMMY · HACKER NEWS
Other models the crowd has fully reviewed, starting with multimodal models like this one.
Qwen3.5 Small models out now!
You can now Fine-tune Qwen3.5 locally! (5GB VRAM)
Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0
[P] TurboQuant for weights: near‑optimal 4‑bit LLM quantization with lossless 8‑bit residual – 3.2× memory savings
Finetuned Qwen3.5 0.8b and I must say it is very good
OvisOCR2: a promising 0.8B local document parser
Where are small Models like Qwen3 0.6B and Qwen3.5 0.8B used ? Huggingface shows 2.88 million downloads this month.[D]
MUST use this to make the text more readable!
Qwen3.5 0.8B Finetuned for Steroids and Peptides
OvisOCR2 (0.8B): first end-to-end model to top OmniDocBench - I threw 827 real scanned medical docs at it, here's everything I learned
A benchmark for tiny LLMs based on a real world problem: natural language file search (using monkeSearch)
eTPS — Effective Tokens Per Second: A Better Way to Measure Local LLM Performance