LOADING DATASHEET
LOADING DATASHEET
by Nvidia
RANKED #183 OF 271 ASSISTANTS · OVERALL #292 OF 6,567 · VIBE SCORE 5.6 · 29 VOICES
NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per appl
2 mentions
1 mentions
26 weeks · 36 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Nemotron 3 Super once and it's yours. No subscription, no rate limits, works offline.
120B total parameters with 12B active (MoE), a Q4 quantized file is about 68 to 72 GB
Desktop PC with 96 GB DDR5 RAM and NVIDIA GeForce RTX 3060 1
12 GB VRAM PLUS 96 GB SYSTEM RAM
~$1100 BUILTPartial GPU offload Q4 quantization, ~3 tok/s, 8k context
Apple Mac Studio M2 Ultra (128 GB Unified Memory)
128 GB UNIFIED MEMORY
~$3400 USEDFull unified memory load Q4 or Q5_K_M quantization, ~22 tok/s, 32k context
Workstation with 4x NVIDIA GeForce RTX 3090 24GB
96 GB VRAM
~$3200 USEDFull GPU VRAM offload NVFP4 or Q4 quantization, ~45 tok/s, 64k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Nemotron 3 Super takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMDesktop PC with 96 GB DDR5 RAM and NVIDIA GeForce RTX 3060 1 | SWEET SPOTApple Mac Studio M2 Ultra (128 GB Unified Memory) | FULL POWERWorkstation with 4x NVIDIA GeForce RTX 3090 24GB |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~2.8 MIN★★★★★ | ~23 S★★★★★ | ~11 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~8.3 MIN★★★★★ | ~68 S★★★★★ | ~33 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~44 MIN★★★★★ | ~6.1 MIN★★★★★ | ~3 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~2.3 HR★★★★★ | ~19 MIN★★★★★ | ~9.3 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~83 MIN★★★★★ | ~11 MIN★★★★★ | ~5.6 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Nemotron 3 Super yet to call it. We found 29 posts but people did not say much either way.
25 POSTS · 4 COMMENTS · DEV FORUMS · REDDIT · DEV.TO · BLOG · GITHUB
Other models the crowd has fully reviewed, starting with text models like this one.
NVIDIA releases Nemotron 3 Nano 4B
NVIDIA releases Nemotron 3 Super!
Nemotron Labs Diff GGUF ?
Qwen 3.5 120B 4bit vs Nemotron 4 Super 4bit vs Minimax M2.5 3bit
As some of you know, I have the long-running habit of keeping a running list of research papers I want to read, revisit, or cite in future articles and projects. Last year, I shared two organized paper lists, one covering January to June and another one…
“We’re at an inflection point in cybersecurity,” Jensen Huang told a sold-out crowd at CrowdStrike’s Fal.Con 2026 in Las Vegas Tuesday. Attacks are now automated. Defense has to be, too. The NVIDIA founder and CEO joined CrowdStrike CEO and founder George…
Ollama Free Tier - Model status
agentic_coding: Ollama fleet on the six-task corpus, and a runner that kills the process group
Ghost Search — MCP-native, Tor-routed search layer for RAG (seeking research collaborators)
EngineCore Error with NVIDIA-Nemotron-3-Super-120B-A12B-FP8 on 2H100
Scope NVIDIA's quota rate to deepseek-v4-pro; make the 402 ladder monotonic
fix(dream): la nuit repart sur des primaires vivants — fin des 410 llama