LOADING DATASHEET
LOADING DATASHEET
by OpenAI
RANKED #24 OF 88 IMAGES · OVERALL #102 OF 6,446 · VIBE SCORE 6.3 · 214 VOICES
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
9 mentions
7 mentions
26 weeks · 357 voices
TOO QUIET
Not enough weekly voices to call a trend yet.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download Whisper once and it's yours. No subscription, no rate limits, works offline.
1.55B parameters for large-v3 (FP16 weights ~3.1 GB, Q4 quantized ~1 GB, with tiny at 39M parameters)
Raspberry Pi 5 (8GB) or basic quad-core CPU laptop
8 GB RAM
~$80whisper.cpp tiny/base models on CPU, ~2x real-time transcription
NVIDIA GeForce RTX 3060 12GB
12 GB VRAM
~$250 USEDfaster-whisper large-v3 in FP16 or INT8, ~15x real-time transcription
NVIDIA GeForce RTX 4090 24GB
24 GB VRAM
~$1750batched large-v3 FP16 high-throughput audio pipelines, ~40x real-time transcription
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long Whisper takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMRaspberry Pi 5 (8GB) or basic quad-core CPU laptop | SWEET SPOTNVIDIA GeForce RTX 3060 12GB | FULL POWERNVIDIA GeForce RTX 4090 24GB |
|---|---|---|---|
| Make an imageone picture at normal settings | —★★★★★ | —★★★★★ | —★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. A DASH MEANS NO SPEED WAS RESEARCHED FOR THAT RIG. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
People are mostly frustrated with Whisper right now. Biggest praise: price. Biggest gripe: saying no too much.
198 POSTS · 16 COMMENTS · GITHUB · REDDIT · DEV.TO · BLOG · LEMMY
0 thumbs up · 11 thumbs down
2 thumbs up · 3 thumbs down
2 thumbs up · 1 thumbs down
“Measured on an M1 Pro with ggml-tiny.en and a real whisper-server, a warm request drops from ~300 ms to ~90 ms; the saving grows with the model size and with a CUDA or Vulkan context, where building the context is the expensive part.”
Other models the crowd has fully reviewed, starting with images models like this one.
Blazingly fast whisper transcriptions with Inference Endpoints
Speculative Decoding for 2x Faster Whisper Inference
Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers
Introducing ChatGPT and Whisper APIs
We’ve trained and are open-sourcing a neural net called Whisper that approaches human level robustness and accuracy on English speech recognition.
GPT-3.5 Turbo, DALL·E and Whisper APIs are also generally available, and we are releasing a deprecation plan for older models of the Completions API, which will retire at the beginning of 2024.
4o image generation is a new, significantly more capable image generation approach than our earlier DALL·E 3 series of models. It can create photorealistic output. It can take images as inputs and transform them.
GPT-4 Turbo with 128K context and lower prices, the new Assistants API, GPT-4 Turbo with Vision, DALL·E 3 API, and more.
We developed a safety mitigation stack to ready DALL·E 3 for wider release and are sharing updates on our provenance research.
DALL·E 3 system card
Rebuild the Whisper Android AAR for 16 KB memory pages
[CI/Build] Continue Whisper GPU capacity fix on current main