LOADING DATASHEET
LOADING DATASHEET
by MiniMax
RANKED #113 OF 231 ASSISTANTS · OVERALL #189 OF 6,583 · VIBE SCORE 5.9 · 57 VOICES
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,.
7 mentions
0 mentions
26 weeks · 54 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download MiniMax M2 once and it's yours. No subscription, no rate limits, works offline.
230B total parameters (MoE with 10B active parameters), a Q4 quant is about 130 GB to 140 GB
Model is too large for single consumer GPUs; requires Mac St
192 GB UNIFIED MEMORY / SYSTEM RAM
~$2800 USEDQ3_K or Q4_K_M quant offloaded to system memory or unified RAM, ~6 tok/s, 8k context
Apple Mac Studio M2 Ultra (192 GB Unified Memory)
192 GB UNIFIED MEMORY
~$3500 USEDQ4_K_M running fully in unified memory via MLX, ~16 tok/s, 32k context
4x NVIDIA RTX 4090 (24 GB each, 96 GB total) plus 128 GB Hos
192 GB VRAM
~$7500 USEDQ8 or FP8 quant across multi-GPU server via vLLM / SGLang, ~38 tok/s, 64k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long MiniMax M2 takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMModel is too large for single consumer GPUs; requires Mac St | SWEET SPOTApple Mac Studio M2 Ultra (192 GB Unified Memory) | FULL POWER4x NVIDIA RTX 4090 (24 GB each, 96 GB total) plus 128 GB Hos |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~83 S★★★★★ | ~31 S★★★★★ | ~13 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~4.2 MIN★★★★★ | ~1.6 MIN★★★★★ | ~39 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~22 MIN★★★★★ | ~8.3 MIN★★★★★ | ~3.5 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~69 MIN★★★★★ | ~26 MIN★★★★★ | ~11 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~42 MIN★★★★★ | ~16 MIN★★★★★ | ~6.6 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly happy with MiniMax M2. Biggest praise: price.
49 POSTS · 8 COMMENTS · REDDIT · GITHUB · BLOG · DEV FORUMS · LEMMY
2 thumbs up · 2 thumbs down
3 thumbs up · 0 thumbs down
“open-sourcing MiniMax M2 — Agent & Code Native, at 8% Claude Sonnet price, ~2x faster (best OS llm currently).”
Other models the crowd has fully reviewed, starting with text models like this one.
Deepseek V3.2 prices are insanely cheap
GPT-5.2 is the new champion of the Elimination Game benchmark, which tests social reasoning, strategy, and deception in a multi-LLM environment. Claude Opus 4.5 and Gemini 3 Flash Preview also made very strong debuts.
open-sourcing MiniMax M2 — Agent & Code Native, at 8% Claude Sonnet price, ~2x faster (best OS llm currently)
How many of you do use Q1 or Q2 of Big models(100-250B)? How's it?
Aligning to What? Rethinking Agent Generalization in MiniMax M2
LLM Research Papers: The 2026 List (January to May) As some of you know, I have the long-running habit of keeping a running list of research papers I want to read, revisit, or cite in future articles and projects. Last year, I shared two organized paper…
From DeepSeek R1 to MiniMax-M2, the largest and most capable open-weight LLMs today remain autoregressive decoder-style transformers, which are built on flavors of the original multi-head attention mechanism. However, we have also seen alternatives to…
Ollama Free Tier - Model status
I Made LLMs Play Texas Hold’em. The Smallest Model Beat a ~1T Model by Being Too Dumb to Fold
Testing Devstral 2 vs MiniMax M2 vs Grok Code Fast for AI code review
Comment in r/ollama
Need some help figuring which quant sizes of some of the big MoEs would fit properly (with how much context + unquantized KV) on a future 256GB or 512GB dram + 48GB VRAM rig I might build later on, since I want to download and save them now (not later on when…