LOADING DATASHEET
LOADING DATASHEET
by OpenAI
RANKED #17 OF 22 AUDIO · OVERALL #343 OF 6,567 · VIBE SCORE 5.4 · 37 VOICES
Speech transcription model for accurate audio-to-text and captioning workflows
4 mentions
4 mentions
26 weeks · 63 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Whisper Large v3 Turbo once and it's yours. No subscription, no rate limits, works offline.
809M parameters, full precision FP16 is about 1.6 GB and INT8 quant is under 1 GB
Nvidia GTX 1650 (4 GB VRAM)
4 GB VRAM
~$75 USEDINT8 or INT4 quant via faster-whisper or whisper.cpp, ~6x realtime speed
Nvidia RTX 3060 12GB
12 GB VRAM
~$280 NEWFP16 or INT8 with large batch sizes and beam search, ~25x realtime speed
Nvidia RTX 4070 Ti Super 16GB
16 GB VRAM
~$799 NEWFP16 batched inference with FlashAttention-2, ~50x realtime speed
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Whisper Large v3 Turbo takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNvidia GTX 1650 (4 GB VRAM) | SWEET SPOTNvidia RTX 3060 12GB | FULL POWERNvidia RTX 4070 Ti Super 16GB |
|---|---|---|---|
| Narrate a scripta minute of spoken audio | ~10 S★★★★★ | ~2 S★★★★★ | ~1 S★★★★★ |
| Transcribe a recordinga ten minute recording turned into text | ~1.7 MIN★★★★★ | ~24 S★★★★★ | ~12 S★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Whisper Large v3 Turbo yet to call it. We found 37 posts but people did not say much either way.
36 POSTS · 1 COMMENTS · GITHUB · HACKER NEWS · REDDIT
Other models the crowd has fully reviewed, starting with audio models like this one.
Whisper large-v3-turbo model published - but not a better model yet
A simple way to transcribe audio to subtitle: gemini-2.0-flash-exp
Show HN: Building Table Canon, an AI Campaign Memory Engine for TTRPGs
fix(aarya): skip too-short audio before whisper STT
fix(ai): replace deprecated default model with Llama 4 Scout 17B (Hindi-supported)
feat(aarya): add resilient Gemini Live fallback and upgrade Workers AI speech pipeline
Voice companion v2 — natural tone + hands-free voice mode + Groq Whisper STT
feat: first-class Lemonade transcription backend with faster-whisper fallback
your coding agent asks you a question out loud, you answer. all of it runs on your machine
Whisper server reachable on the LAN (BIND option, optional bearer token) so HA Assist and the phones can use the house STT
fix(qwen): guard all backends against EOS-miss runaway audio
landing(voice-transaction-entry): fix unsupported claim — "Whisper-small", and "whisper.rn ships the model"