LOADING DATASHEET
LOADING DATASHEET
by Nvidia
RANKED #256 OF 278 ASSISTANTS · OVERALL #402 OF 6,579 · VIBE SCORE 5.1 · 27 VOICES
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
3 mentions
3 mentions
26 weeks · 30 voices
HONEYMOON FADING
Early praise is cooling off in recent weeks.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Nemotron Super once and it's yours. No subscription, no rate limits, works offline.
120B total parameters (12B active MoE), a Q4 quantized file is about 65 GB to 72 GB
Mac Studio M2 Ultra (128 GB Unified Memory)
128 GB UNIFIED RAM
~$3500 USEDQ4 GGUF quant, ~18 tok/s, 8k context across unified memory
Workstation with 3x NVIDIA GeForce RTX 3090 24GB
72 GB VRAM
~$2400 USEDQ4_K_M quant with full GPU offload, ~28 tok/s, 16k context
Workstation with 4x NVIDIA RTX 6000 Ada 48GB
192 GB VRAM
~$28000 NEWFP8 / NVFP4 high precision, ~50 tok/s, full 128k+ context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Nemotron Super takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMMac Studio M2 Ultra (128 GB Unified Memory) | SWEET SPOTWorkstation with 3x NVIDIA GeForce RTX 3090 24GB | FULL POWERWorkstation with 4x NVIDIA RTX 6000 Ada 48GB |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~28 S★★★★★ | ~18 S★★★★★ | ~10 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~83 S★★★★★ | ~54 S★★★★★ | ~30 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~7.4 MIN★★★★★ | ~4.8 MIN★★★★★ | ~2.7 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~23 MIN★★★★★ | ~15 MIN★★★★★ | ~8.3 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~14 MIN★★★★★ | ~8.9 MIN★★★★★ | ~5 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly frustrated with Nemotron Super right now.
22 POSTS · 5 COMMENTS · GITHUB · REDDIT · X
0 thumbs up · 5 thumbs down
Other models the crowd has fully reviewed, starting with text models like this one.
There is a very real possibility that Google, OpenAI, Anthropic, etc. will release their own super cheap versions of Grok-4-fast!
New Nvidia Llama Nemotron Reasoning Models
NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise?
What a joke...
You’re going to start seeing more and more companies catering open weights models based on the hardware out there. Nemot
Yes but then your stuck only using that model, that’s the issue. 27B is it, 2 x DGX Sparks gets you the ability to run D
Anyone gotten Nemotron 49B Running in Ollama?
Free Coding Agent with NVIDIA Nemotron (Open Source)
Unsloth, please create a Instruction-Following Dataset.
AI team: add Grok 4.6 via Bedrock BYOK, drop the retired GLM free rung
Grok 4.6 via Bedrock BYOK as the last-resort rung
fix(vllm): load Super-Omni RADIO final LayerNorm