LOADING DATASHEET
LOADING DATASHEET
by Meta
RANKED #109 OF 230 ASSISTANTS · OVERALL #182 OF 6,596 · VIBE SCORE 5.9 · 74 VOICES
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
6 mentions
7 mentions
26 weeks · 81 voices
HONEYMOON FADING
Early praise is cooling off in recent weeks.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Llama 3.3 70B Instruct once and it's yours. No subscription, no rate limits, works offline.
70B parameters, a Q4_K_M GGUF is about 43 GB, unquantized 16-bit is around 140 GB
Nvidia GeForce RTX 3090 24GB paired with 64GB DDR4 system RA
24 GB VRAM PLUS 64 GB RAM
~$750 USEDIQ3_XXS or partial offload Q4_K_M, ~5 tok/s, 8k context
Dual Nvidia GeForce RTX 3090 24GB (48GB total)
48 GB VRAM
~$1500 USEDQ4_K_M fully in VRAM, ~20 tok/s, 16k context
Apple Mac Studio M2 Ultra (192GB Unified Memory)
192 GB UNIFIED MEMORY
~$5500 NEWQ8_0 or FP16 fully in memory, ~16 tok/s, full 128k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Llama 3.3 70B Instruct takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNvidia GeForce RTX 3090 24GB paired with 64GB DDR4 system RA | SWEET SPOTDual Nvidia GeForce RTX 3090 24GB (48GB total) | FULL POWERApple Mac Studio M2 Ultra (192GB Unified Memory) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~1.7 MIN★★★★★ | ~25 S★★★★★ | ~31 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~5 MIN★★★★★ | ~75 S★★★★★ | ~1.6 MIN★★★★★ |
| Build a websitea small landing page, markup and styles together | ~27 MIN★★★★★ | ~6.7 MIN★★★★★ | ~8.3 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~83 MIN★★★★★ | ~21 MIN★★★★★ | ~26 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~50 MIN★★★★★ | ~13 MIN★★★★★ | ~16 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Llama 3.3 70B Instruct yet to call it. We found 74 posts but people did not say much either way.
71 POSTS · 3 COMMENTS · STACK OVERFLOW · GITHUB · REDDIT · DEV FORUMS · LEMMY · X
Other models the crowd has fully reviewed, starting with text models like this one.
Comment in r/perplexity_ai
Hi folks, Llama 3.3 released, update Perplexity labs please 🥺
I think I made a prompt that breaks AI's tendency to push mainstream narratives. This interplay was pretty cool! One of my best yet.
I built a framework where multi-agent swarms are YAML files, not code.
We'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it.
8x AMD Instinct Mi50 Server + Llama-3.3-70B-Instruct + vLLM + Tensor Parallelism -> 25t/s
160+ frontier models behind one free api key and nobody talks about it nvidia nim gives you an openai-compatible endpoin
HuggingChat v2 has just nailed model routing!
Deepinfra sudden 2.5x price hike for llama 3.3 70b instruction turbo. How are others coping with this?
HuggingChat v2 has just nailed model routing!
Not a UI but a model question
Why no talk about Medium (size) Language Models? 70-200B