LOADING DATASHEET
LOADING DATASHEET
by Ornith AI
RANKED #194 OF 282 ASSISTANTS · OVERALL #308 OF 6,622 · VIBE SCORE 5.6 · 26 VOICES
Ornith-1.5-9B-MLX-4bit on Hugging Face (text generation). 3,043 downloads. Open weights for local or hosted use.
4 mentions
0 mentions
26 weeks · 44 voices
AGING WELL
The crowd is warmer now than it was at the start.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Ornith 1.5 9B once and it's yours. No subscription, no rate limits, works offline.
9B parameters, a Q4 file is about 5.5 GB (around 18 GB in BF16)
Nvidia GeForce RTX 3060 12GB
12 GB VRAM
~$250 USEDQ4_K / Q5_K, ~35 tok/s, 8k context
Apple Mac mini M4 (16GB RAM)
16 GB UNIFIED MEMORY
~$599 NEWQ8_0 / Q6_K, ~40 tok/s, 16k context
Nvidia GeForce RTX 4090 24GB
24 GB VRAM
~$1750 USEDBF16 unquantized, ~110 tok/s, 32k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Ornith 1.5 9B takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNvidia GeForce RTX 3060 12GB | SWEET SPOTApple Mac mini M4 (16GB RAM) | FULL POWERNvidia GeForce RTX 4090 24GB |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~14 S | ~13 S | ~5 S |
| Summarize a documenta long report boiled down to the points that matter | ~43 S | ~38 S | ~14 S |
| Build a websitea small landing page, markup and styles together | ~3.8 MIN★★★★★ | ~3.3 MIN★★★★★ | ~73 S★★★★★ |
| Build a backendan API with routes, storage and tests | ~12 MIN★★★★★ | ~10 MIN★★★★★ | ~3.8 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~7.1 MIN★★★★★ | ~6.3 MIN★★★★★ | ~2.3 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Ornith 1.5 9B yet to call it. We found 26 posts but people did not say much either way.
21 POSTS · 5 COMMENTS · REDDIT · GITHUB · X · LEMMY
Other models the crowd has fully reviewed, starting with multimodal models like this one.
We have Q3.8 35B at home: 3x new Ornith 1.5 released
Ornith 1.5: 9B dense and 35B/397B MoEs
Does anyone have real experience with Ornith-1.5-9B for coding
Low to midrange systems (8-32 GB) vs. free cloud tiers
Comment in r/unsloth
Ornith 1.5 + Qwen 3.8 27B hybrid runtime + DFlash speculative decoding on Metal
While Everyone Is Excited About Qwen 3.8 27B, Here’s the Reality for a 16GB AMD GPU User
Evaluate Ornith 1.5 and Qwen 3.6 on the dev-workflow matrix (prior rejections predate the current stack)
ROCm decode is 2.4x behind llama.cpp on a DENSE model (Ornith-1.5-9B, qwen35): the gap is not MoE-specific
KERNEL-QUANT-CIQ-GEMM-ROCM: ROCm keep-quant GEMM (KQuantGemmK) has no MFMA arm — 61.5% of GPU time, ~4.4x behind llama.cpp mul_mat_q on gfx1200
feat(mcp): Lightweight MCP Server with Palace Query Language (PQL) & Principled Retrieval Engine
Models have gotten so good, it doesn't really matter what you use. Stock Pi agent + good MCP + good prompt -> decent