LOADING DATASHEET
LOADING DATASHEET
by K2 Fsa
RANKED #14 OF 26 AUDIO · OVERALL #305 OF 6,578 · VIBE SCORE 5.6 · 27 VOICES
OmniVoice-GGUF on Hugging Face (text to speech). 191,335 downloads. Open weights for local or hosted use.
12 mentions
1 mentions
26 weeks · 29 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download OmniVoice once and it's yours. No subscription, no rate limits, works offline.
0.8B parameters (~3.3 GB total weights, main model is ~2.5 GB in FP16/BF16)
NVIDIA GeForce RTX 3060 12GB
12 GB VRAM
~$220 USEDFP16 full precision, ~20x realtime generation
NVIDIA GeForce RTX 4070 12GB
12 GB VRAM
~$520 NEWFP16 / BF16 batch inference, ~40x realtime generation
NVIDIA GeForce RTX 4090 24GB
24 GB VRAM
~$1750 NEWFP16 / BF16 parallel multi-speaker batches, ~60x realtime generation
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long OmniVoice takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce RTX 3060 12GB | SWEET SPOTNVIDIA GeForce RTX 4070 12GB | FULL POWERNVIDIA GeForce RTX 4090 24GB |
|---|---|---|---|
| Narrate a scripta minute of spoken audio | ~3 S★★★★★ | ~2 S★★★★★ | ~1 S★★★★★ |
| Transcribe a recordinga ten minute recording turned into text | ~30 S★★★★★ | ~15 S★★★★★ | ~10 S★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about OmniVoice yet to call it. We found 27 posts but people did not say much either way.
16 POSTS · 11 COMMENTS · REDDIT · GITHUB · DEV.TO
Other models the crowd has fully reviewed, starting with audio models like this one.
Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0
[WIP] Working ComfyUI Omnivoice ,
OmniVoice: multilingual local TTS with 600+ languages, voice cloning, and an OpenAI-compatible server
(request) Is it possible to add more TTS models to the unsloth desktop?
Comment in r/StableDiffusion
Yann LeCun Quit Meta to Bet Against Pixels — Inside JEPA Breach Protocol Podcast
Qwen3-TTS help
Podcast about Yann LeCun's Energy-Based Models
native: transcribe.cpp 0.2.3 + audio.cpp 0.7.1 (1.0.2); roster: drop Cohere Arabic, add Qwen 3.5 4B and EuroLLM 1.7B
Comment in r/LocalLLaMA
docs(services): pipecat fleet review — voice MCP + realtime agents under hardware constraint
Follow-up to aceteam-ai/aceteam#9523 (S8 slice of aceteam-ai/aceteam#9498). aceteam-ai/aceteam#9523 defaulted vllm, llamacpp, bonsai, unlimited-ocr, and sglang's compose files to a loopback-only (127.0.0.1) host publish. The parent issue's full acceptance…