LOADING DATASHEET
LOADING DATASHEET
by Nvidia
RANKED #179 OF 228 ASSISTANTS · OVERALL #278 OF 6,618 · VIBE SCORE 5.5 · 69 VOICES
Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including c
6 mentions
5 mentions
26 weeks · 120 voices
HONEYMOON FADING
Early praise is cooling off in recent weeks.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Nemotron 3 Ultra once and it's yours. No subscription, no rate limits, works offline.
550B total parameters (55B active MoE), a 4-bit quantized file (NVFP4 or MLX 4-bit) requires about 300 to 335 GB of memory
Custom PC with 512 GB system RAM and single RTX 4090 (24 GB)
512 GB DDR5 RAM + 24 GB VRAM
~$3500 BUILTThe model is too large for single consumer GPUs; CPU RAM offloading with 4-bit quant runs extremely slowly at around 1 t
Apple Mac Studio (M3/M4 Ultra) with max unified memory
512 GB UNIFIED MEMORY
~$7500 NEWMLX 4-bit quant, usable daily-driver speed at around 8 to 14 tok/s with moderate context
Multi-GPU workstation with 8x NVIDIA RTX 6000 Ada (or 8x H10
384 GB TO 640 GB VRAM
~$60000 BUILTFull FP8 / BF16 precision, native high-speed throughput up to 35+ tok/s and full long context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Nemotron 3 Ultra takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMCustom PC with 512 GB system RAM and single RTX 4090 (24 GB) | SWEET SPOTApple Mac Studio (M3/M4 Ultra) with max unified memory | FULL POWERMulti-GPU workstation with 8x NVIDIA RTX 6000 Ada (or 8x H10 |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | NO DATA | ~45 S | NO DATA |
| Summarize a documenta long report boiled down to the points that matter | —★★★★★ | ~2.3 MIN★★★★★ | —★★★★★ |
| Build a websitea small landing page, markup and styles together | —★★★★★ | ~12 MIN★★★★★ | —★★★★★ |
| Build a backendan API with routes, storage and tests | —★★★★★ | ~38 MIN★★★★★ | —★★★★★ |
| Build a gamea playable browser game in one file | —★★★★★ | ~23 MIN★★★★★ | —★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. A DASH MEANS NO SPEED WAS RESEARCHED FOR THAT RIG. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People generally like Nemotron 3 Ultra, with some gripes. Biggest praise: getting facts right. Biggest gripe: working when you need it.
60 POSTS · 9 COMMENTS · REDDIT · X · GITHUB · BLUESKY · LEMMY · BLOG
4 thumbs up · 2 thumbs down
4 thumbs up · 1 thumbs down
0 thumbs up · 5 thumbs down
3 thumbs up · 1 thumbs down
4 thumbs up · 0 thumbs down
1 thumbs up · 2 thumbs down
Other models the crowd has fully reviewed, starting with multimodal models like this one.
INCREDIBLE STUFF INCOMING
Inkling by Thinking Machines is the #1 US open weight model now
Top American AI Execs urge regulation on cheap chinese AI models. suggest a possibility of a global AI communism otherwise
Claude Sonnet 5 Artificial Analysis Results & Comparison
Artificial Analysis: Muse Spark 1.1 Results
Anthropic Leads top 10 models by $/spent
A trend we continue to see in open model releases is that the ecosystem is becoming more diverse, with an increasing number of organizations releasing a wide range of models. A year ago, open artifacts and the open model landscape more broadly were dominated…
As I’ve been recapping fundamentals of post-training to wrap up my RLHF / Post-training book I knew I needed to get Finbarr Timbers back on the podcast to talk about the state of play. Over the last few months we’ve had many discussions on what we’d need to…
Edit Jun. 11: Anthropic changed their silent model manipulation of AI research queries to also use a classifier like the other safety domains. This addresses a key concern I had in the mistreatment of “safety” in the release, and props to Anthropic for a…
It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models. DeepSeek-R1 followed about four months later, together with details of a reinforcement learning with verifiable rewards (RLVR) recipe to…
LLM Research Papers: The 2026 List (January to May) As some of you know, I have the long-running habit of keeping a running list of research papers I want to read, revisit, or cite in future articles and projects. Last year, I shared two organized paper…
Nemotron 3 Ultra is now available for Pro and Max subscribers on Perplexity and Computer