LOADING DATASHEET
LOADING DATASHEET
by Meta
RANKED #173 OF 228 ASSISTANTS · OVERALL #272 OF 6,618 · VIBE SCORE 5.5 · 400 VOICES
Llama3 available through cloud APIs.
45 mentions
45 mentions
26 weeks · 112 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Llama3 once and it's yours. No subscription, no rate limits, works offline.
8B parameters (Q4 file is ~4.7 GB) and 70B parameters (Q4 file is ~40 GB)
Nvidia GeForce RTX 3060 12GB
12 GB VRAM
~$250 USEDLlama-3-8B at Q4_K_M or Q8 quantizations, ~40 tok/s, 8k context fully offloaded to GPU
Apple Mac Mini M2 Pro
32 GB UNIFIED MEMORY
~$1,100 USEDLlama-3-8B FP16 at full speed or heavily quantized small models with large context windows
Dual Nvidia GeForce RTX 3090 24GB Setup
48 GB VRAM
~$1,400 USEDLlama-3-70B at Q4_K_M, ~15 to 20 tok/s with full GPU offload and full context window
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Llama3 takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNvidia GeForce RTX 3060 12GB | SWEET SPOTApple Mac Mini M2 Pro | FULL POWERDual Nvidia GeForce RTX 3090 24GB Setup |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~13 S★★★★★ | —★★★★★ | ~29 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~38 S★★★★★ | —★★★★★ | ~86 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~3.3 MIN★★★★★ | —★★★★★ | ~7.6 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~10 MIN★★★★★ | —★★★★★ | ~24 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~6.3 MIN★★★★★ | —★★★★★ | ~14 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. A DASH MEANS NO SPEED WAS RESEARCHED FOR THAT RIG. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People generally like Llama3, with some gripes. Biggest praise: help with code. Biggest gripe: speed.
346 POSTS · 54 COMMENTS · STACK OVERFLOW · REDDIT · GITHUB · DEV.TO · DEV FORUMS · HACKER NEWS · LEMMY · X · BLOG
5 thumbs up · 0 thumbs down
1 thumbs up · 2 thumbs down
2 thumbs up · 1 thumbs down
1 thumbs up · 2 thumbs down
“llama3 always hallucinating?.”
Checked picks first: tags say why you would switch, stars say how fully each one stands in for Llama3.
I trapped LLama3.2B into an art installation and made it question its own existence endlessly
We fine-tuned Llama 405B on AMD GPUs
Looks like we’re getting LLama3 405B this week
Ollama Models Ranked by VRAM Requirements
Your Ollama Servers Are So Open, Even My Grandma Could Use Them!
Google released ultra lightweight Gemma 2 models, the 27B one surpasses llama3 70B and 9B variant surpasses Claude 3 Haiku
A laptop bought over a decade before the AI boom (2010) running llama3:8b
DeepDive in everything of Llama3: revealing detailed insights and implementation
Knowledge graph using gemma3
Comment in r/singularity
"LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures"
Give your local Ollama models a personal knowledge bank (graph-based, not just vector search)