LOADING DATASHEET
LOADING DATASHEET
by Meta
RANKED #209 OF 228 ASSISTANTS · OVERALL #320 OF 6,618 · VIBE SCORE 5.2 · 165 VOICES
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
17 mentions
22 mentions
26 weeks · 82 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Llama 4 Scout once and it's yours. No subscription, no rate limits, works offline.
109B parameters (MoE, 17B active), a Q4 file is about 61 GB
2x NVIDIA RTX 3090 (Used)
48 GB VRAM
~$1,600 USEDQ3_K_M, ~15 tok/s, 8k context (requires tensor parallelism/split weights)
Mac Studio M2 Ultra (128GB Unified Memory)
128 GB UNIFIED MEMORY
~$3,000 USEDQ4_K_M or Q5_K_M, ~12 tok/s, 32k context comfortably
Mac Studio M3 Ultra (192GB Unified Memory)
192 GB UNIFIED MEMORY
~$5,500Q8_0, ~10 tok/s, full 192k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Llama 4 Scout takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUM2x NVIDIA RTX 3090 (Used) | SWEET SPOTMac Studio M2 Ultra (128GB Unified Memory) | FULL POWERMac Studio M3 Ultra (192GB Unified Memory) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~33 S★★★★★ | ~42 S★★★★★ | ~50 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~1.7 MIN★★★★★ | ~2.1 MIN★★★★★ | ~2.5 MIN★★★★★ |
| Build a websitea small landing page, markup and styles together | ~8.9 MIN★★★★★ | ~11 MIN★★★★★ | ~13 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~28 MIN★★★★★ | ~35 MIN★★★★★ | ~42 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~17 MIN★★★★★ | ~21 MIN★★★★★ | ~25 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly frustrated with Llama 4 Scout right now.
146 POSTS · 19 COMMENTS · BLUESKY · HACKER NEWS · GITHUB · LEMMY · REDDIT · DEV FORUMS · STACK OVERFLOW · OTHER FORUMS
0 thumbs up · 3 thumbs down
Checked picks first: tags say why you would switch, stars say how fully each one stands in for Llama 4 Scout.
Llama 4 Maverick/Scout 17B launched on Lambda API
Car Wash Test on 53 leading models: “I want to wash my car. The car wash is 50 meters away. Should I walk or drive?”
Llama 4 benchmarks !!
Llama 4 Scout with 10M tokens
How I use Cursor 10+ hours a day without torching my Claude Opus 4.6 limits
Seems like there was a lot of truth to this leak from 2 months ago llama 4 is beyond disappointing. it's a model that shouldn't have been released.
okay guys turn out the llama 4 benchmark is a fraud 10 million context window is fraud
DuckDuckGo installs are up 30% as users reject being ‘force-fed’ Google’s AI Search
Free Model List (API Keys)
Is RAG still relevant with 10M+ context length
QwQ-32b outperforms Llama-4 by a lot!
LLAMA 4 Scout on Mac, 32 Tokens/sec 4-bit, 24 Tokens/sec 6-bit