LOADING DATASHEET
LOADING DATASHEET
by Meta
RANKED #164 OF 271 ASSISTANTS · OVERALL #259 OF 6,567 · VIBE SCORE 5.7 · 38 VOICES
DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techn
5 mentions
2 mentions
26 weeks · 26 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download R1 Distill Llama 70B once and it's yours. No subscription, no rate limits, works offline.
70B parameters, a Q4 file is about 42 GB (requires ~48 GB memory to run)
Mac Studio M2 Max (64GB RAM)
64 GB UNIFIED MEMORY
~$1800 USEDQ4_K_M quant fully in unified memory, ~12 tok/s, 8k context
Dual NVIDIA RTX 3090 (24GB x 2)
48 GB VRAM
~$1500 USEDQ4_K_M quant split across GPUs, ~22 tok/s, 16k context
Dual NVIDIA RTX 4090 (24GB x 2)
48 GB VRAM
~$3400 NEWQ4_K_M or Q5_K_M with flash attention, ~35 tok/s, 32k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long R1 Distill Llama 70B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMMac Studio M2 Max (64GB RAM) | SWEET SPOTDual NVIDIA RTX 3090 (24GB x 2) | FULL POWERDual NVIDIA RTX 4090 (24GB x 2) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~42 S | ~23 S | ~14 S |
| Summarize a documenta long report boiled down to the points that matter | ~2.1 MIN | ~68 S | ~43 S |
| Build a websitea small landing page, markup and styles together | ~11 MIN★★★★★ | ~6.1 MIN★★★★★ | ~3.8 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~35 MIN★★★★★ | ~19 MIN★★★★★ | ~12 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~21 MIN★★★★★ | ~11 MIN★★★★★ | ~7.1 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about R1 Distill Llama 70B yet to call it. We found 38 posts but people did not say much either way.
36 POSTS · 2 COMMENTS · REDDIT · LEMMY · STACK OVERFLOW · GITHUB
Other models the crowd has fully reviewed, starting with text models like this one.
Testing Uncensored DeepSeek-R1-Distill-Llama-70B-abliterated FP16
DeepSeek released deepseek/deepseek-r1-distill-llama-70b via OpenRouter; Use it with Cursor now!
NEW DeepSeek-R1-Distill-Llama-70B vs Claude,o1-mini,4o
Tip: If you are using Deepseek R1 through OpenRouter and cannot see it "thinking", do this:
DeepSeek R1 70B on Cerebras Inference Cloud!
LLM in the Terminal
DeepSeek released deepseek/deepseek-r1-distill-llama-70b via OpenRouter; Use it with Cursor now!
Function Calling in Terminal + DeepSeek-R1-Distill-Llama-70B-Q_8 + vLLM -> Sometimes...
DeepSeek-R1-Distill-Llama-70B: how to disable these <think> tags in output?
Hosted deepseek-r1-distill-qwen-32b
Using Reasonder Models to Augment traditional LLM
DeepSeek-R1-Distill-Llama-70B: how to disable these <think> tags in output?