LOADING DATASHEET
LOADING DATASHEET
by Prism ML
RANKED #162 OF 233 ASSISTANTS · OVERALL #259 OF 6,558 · VIBE SCORE 5.6 · 32 VOICES
Bonsai-27B-AWQ-4bit on Hugging Face (image text to text). 2,448 downloads. Open weights for local or hosted use.
10 mentions
5 mentions
26 weeks · 43 voices
AGING WELL
The crowd is warmer now than it was at the start.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Bonsai 27B once and it's yours. No subscription, no rate limits, works offline.
27B parameters, 1-bit binary file is about 3.9 GB, ternary format is about 5.9 GB to 7.2 GB
NVIDIA GeForce RTX 3060 12GB
12 GB VRAM
~$220 USED1-bit binary GGUF, ~16 tok/s, 8k context
NVIDIA GeForce RTX 4070 12GB
12 GB VRAM
~$520 USEDTernary GGUF, ~30 tok/s, 16k context
Apple Mac Studio M2 Ultra
64 GB UNIFIED MEMORY
~$3100 USEDTernary / full context MLX with vision tower, ~55 tok/s, 64k+ context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Bonsai 27B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce RTX 3060 12GB | SWEET SPOTNVIDIA GeForce RTX 4070 12GB | FULL POWERApple Mac Studio M2 Ultra |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~31 S | ~17 S | ~9 S |
| Summarize a documenta long report boiled down to the points that matter | ~1.6 MIN★★★★★ | ~50 S★★★★★ | ~27 S★★★★★ |
| Build a websitea small landing page, markup and styles together | ~8.3 MIN★★★★★ | ~4.4 MIN★★★★★ | ~2.4 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~26 MIN★★★★★ | ~14 MIN★★★★★ | ~7.6 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~16 MIN★★★★★ | ~8.3 MIN★★★★★ | ~4.5 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Bonsai 27B yet to call it. We found 32 posts but people did not say much either way.
23 POSTS · 9 COMMENTS · REDDIT · X · GITHUB · LEMMY
Other models the crowd has fully reviewed, starting with multimodal models like this one.
Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB
I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM
Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
Stack update: Devin Max — $200 Cursor Ultra — $200 SuperGrok MiniMax Token Plan — annual Google AI Pro — annual Hermes A
Prism-ML's Bonsai-27B Benchmarks
1-bit / 2-bit / Ternary / Bitnet Models - Updates & Tracking
Why have 8B-12B models been dropped?
Using the Bonsai 27b 1b quant locally - regularly.
Built a from-scratch BitNet inference engine in pure C — 1.8× faster than bitnet.cpp on Xeon (36 tok/s), zero dependencies [BitNet & Bonsai CPU testers wanted]
PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
The "Galician Gene" Directive: Forcing LLMs to ask for context instead of hallucinating (and reducing token waste)
How to run Prism Bonsai 27B