LOADING DATASHEET
LOADING DATASHEET
by OpenAI
RANKED #18 OF 19 AUDIO · OVERALL #323 OF 6,608 · VIBE SCORE 5.2 · 39 VOICES
Speech transcription model for accurate audio-to-text and captioning workflows
3 mentions
4 mentions
26 weeks · 40 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Whisper Large once and it's yours. No subscription, no rate limits, works offline.
1.55B parameters, an FP16 weights file is about 3.1 GB while Q4 is about 1.5 GB
NVIDIA GeForce GTX 1650
4 GB VRAM
~$75 USEDfaster-whisper INT8 or Q4, ~3x realtime transcription speed
NVIDIA GeForce RTX 4060
8 GB VRAM
~$299 NEWfaster-whisper FP16 / INT8 batching, ~15x realtime transcription speed
NVIDIA GeForce RTX 4090
24 GB VRAM
~$1,850 USEDFP16 full precision with large batch sizes, ~45x realtime transcription speed
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Whisper Large takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce GTX 1650 | SWEET SPOTNVIDIA GeForce RTX 4060 | FULL POWERNVIDIA GeForce RTX 4090 |
|---|---|---|---|
| Narrate a scripta minute of spoken audio | ~20 S★★★★★ | ~4 S★★★★★ | ~1 S★★★★★ |
| Transcribe a recordinga ten minute recording turned into text | ~3.3 MIN★★★★★ | ~40 S★★★★★ | ~13 S★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about Whisper Large yet to call it. We found 39 posts but people did not say much either way.
37 POSTS · 2 COMMENTS · STACK OVERFLOW · DEV.TO · HACKER NEWS · REDDIT · GITHUB
Other models the crowd has fully reviewed, starting with audio models like this one.
You can now train your own TTS voice models locally!
You can now train your own Text-to-Speech (TTS) models locally!
GPT-4o-transcribe outperforms Whisper-large
This Week in AI (Curated News)
Near Realtime speech-to-text with self hosted Whisper Large (WebSocket & WebAudio)
Image generation is now available alongside LLMs and Whisper in Lemonade v9.2
HuMo LipSync Model from ByteDance! Demo, Models, Workflows, Guide, and Thoughts
Fully offline multi-modal RAG for NASA Life Sciences PDFs + images + audio + knowledge graphs – best 2025 local stack?
AI is watching a film via Marlin(visuals), Whisper(audio), and Pallaidium. Input video by avataraim.
Whisper large v3 benchmark: 1 Million hrs transcribed for $5110 (11,736 mins per dollar) on RTX-series GPUs
This week in AI - all the Major AI developments in a nutshell
WhisperFile - extremely easy OpenAI's whisper.cpp audio transcription in one file