LOADING DATASHEET
LOADING DATASHEET
by Google
RANKED #156 OF 242 ASSISTANTS · OVERALL #256 OF 6,401 · VIBE SCORE 5.6 · 94 VOICES
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
14 mentions
7 mentions
26 weeks · 59 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Gemma 3 27B once and it's yours. No subscription, no rate limits, works offline.
27B parameters, a Q4 file is about 16 to 17 GB
NVIDIA GeForce RTX 3060 12GB plus 32GB system RAM
12 GB VRAM + 32 GB RAM
~$220 USEDQ4_K_M partial GPU offload, ~3 tok/s, 8k context
NVIDIA GeForce RTX 3090 24GB
24 GB VRAM
~$650 USEDQ4_K_M fully in VRAM, ~18 tok/s, 16k context
Apple Mac Studio M2 Ultra (64GB)
64 GB UNIFIED MEMORY
~$2999 NEWQ8_0 or full precision FP16, ~22 tok/s, 128k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Gemma 3 27B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce RTX 3060 12GB plus 32GB system RAM | SWEET SPOTNVIDIA GeForce RTX 3090 24GB | FULL POWERApple Mac Studio M2 Ultra (64GB) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~2.8 MIN★★★★★ | ~28 S★★★★★ | ~23 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~8.3 MIN★★★★★ | ~83 S★★★★★ | ~68 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~44 MIN★★★★★ | ~7.4 MIN★★★★★ | ~6.1 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~2.3 HR★★★★★ | ~23 MIN★★★★★ | ~19 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~83 MIN★★★★★ | ~14 MIN★★★★★ | ~11 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
The internet is split on Gemma 3 27B. Biggest praise: speed. Biggest gripe: price.
75 POSTS · 19 COMMENTS · GITHUB · X · REDDIT · DEV FORUMS · OTHER FORUMS · HACKER NEWS · LEMMY · BLOG
3 thumbs up · 0 thumbs down
0 thumbs up · 3 thumbs down
“Mistral Small 3.1 Mistral Small 3.1 24B , which was released in March shortly after Gemma 3, is noteworthy for outperforming Gemma 3 27B on several benchmarks (except for math) while being faster.”
Other models the crowd has fully reviewed, starting with multimodal models like this one.
Best models under 16GB
Cosmos Predict 2 & Chroma v42 (feat. Gemma-3)
We ran a predator's playbook on an AI - it folded using the same dynamics described in social psychology
ollama 0.12.11 Brings Vulkan Acceleration
I had originally planned to write about DeepSeek V4. Since it still hasn’t been released, I used the time to work on something that had been on my list for a while, namely, collecting, organizing, and refining the different LLM architectures I have covered…
From DeepSeek R1 to MiniMax-M2, the largest and most capable open-weight LLMs today remain autoregressive decoder-style transformers, which are built on flavors of the original multi-head attention mechanism. However, we have also seen alternatives to…
OpenAI just released their new open-weight LLMs this week: gpt-oss-120b and gpt-oss-20b, their first open-weight models since GPT-2 in 2019. And yes, thanks to some clever optimizations, they can run locally (but more about this later). This is the first time…
Last updated: Apr 2, 2026 (added Gemma 4 in section 23) It has been seven years since the original GPT architecture was developed. At first glance, looking back at GPT-2 (2019) and forward to DeepSeek V3 and Llama 4 (2024-2025), one might be surprised at how…
Optimize Gemma 3 Inference: vLLM on GKE 🏎️💨
When a translation model starts solving the problem instead of translating it (small rant)
The Claude D&D thing I've posted about here is on the App Store now
New Political Leaning Benchmark shows Grok 4.5 ranks as the most neutral model overall while MiniMax M3 leads for Open-Source