Whisper Large v3 Turbo: speed, accuracy and RAM on a Mac
Whisper Large v3 Turbo is OpenAI's speech-to-text model for 99 languages, with translation into English. In PocketWebTools for Mac it is the model for every language Parakeet does not cover. Below is what we measured on a 16 GB M1 Pro: speed, memory and accuracy on five test sets.
Which Mac runs it
- Mac memory
- 8 GB and up
- File
- 0.89 GB
One GGUF file, 8‑bit.
- Memory while transcribing
- 1.1 GB
Peak for the whole process on our test set.
Whisper Large v3 Turbo speed on an M1 Pro
Measured in the app's own engine on a 16 GB M1 Pro.
- Short clips, 5 to 20 s
- 8 to 16 times realtime
Across the five test sets below.
- Against Parakeet v3
- 4 to 8 times slower
Parakeet ran the same clips at 60 to 72 times realtime.
Whisper Large v3 Turbo accuracy next to Parakeet
Word error rate on 170 clips with reference transcripts, 28 minutes in all, scored with the Open ASR Leaderboard normaliser. Lower is better. The clips are short, 5 to 20 seconds, and there are 25 to 40 per set.
- Read English (LibriSpeech test-other)
- 2.88%
Parakeet v3: 2.38%.
- Meetings (AMI)
- 14.01%
Parakeet v3: 7.72%. Whisper's weakest set.
- Earnings calls (Earnings22)
- 10.76%
Parakeet v3: 9.78%.
- French audiobooks
- 4.07%
Parakeet v3: 7.45%. Whisper is the better of the two.
- German audiobooks
- 9.57%
Parakeet v3: 12.22%. Whisper is the better of the two.
- Punctuation and capitals
- On every set
Settings the app runs Whisper Large v3 Turbo with
- File
- whisper-large-v3-turbo-Q8_0.gguf
The transcribe.cpp project's own conversion, pinned to one revision and checked against its SHA-256 after download.
- Engine
- transcribe.cpp on Metal
- Quantization
- 8‑bit (Q8_0)
- Timestamps
- Per segment, not per word
What the app does with Whisper Large v3 Turbo
Transcribe
The pick for any of its 99 languages, with the language detected for you. Export as text, SRT, WebVTT or JSON.
Translate to English
Transcribes speech in another language straight into English text.
Clip
Transcribes videos in languages Parakeet does not cover, so the app can still pick and caption clips from them.
Where it falls short
- It is 4 to 8 times slower than Parakeet v3 on the same Mac.
- Meetings are its weak spot on our tests: 14.01% word error rate against 7.72% for Parakeet v3.
- Times come per segment, not per word. Word-by-word captions and the transcript-based video editor need Parakeet.
- The test clips are short. We measured English, French and German only, so we have no number of our own for its other 96 languages.
Run Whisper Large v3 Turbo in PocketWebTools for Mac
One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.
Whisper Large v3 Turbo questions
- How fast is Whisper Large v3 Turbo on an M1 Pro?
- 8 to 16 times realtime on short clips, on a 16 GB M1 Pro with transcribe.cpp on Metal. Parakeet v3 ran the same clips at 60 to 72 times realtime.
- How much RAM does Whisper Large v3 Turbo need?
- About 1.1 GB while transcribing, measured as the peak for the whole process. The 8‑bit file is 0.89 GB. PocketWebTools for Mac files it under 8 GB Macs.
- Whisper Large v3 Turbo or Parakeet v3?
- Parakeet v3 for English and when you need a time for every word: it was more accurate on all three of our English sets and 4 to 8 times faster. Whisper Large v3 Turbo for French and German, where it scored 4.07% and 9.57% against 7.45% and 12.22%, for any language outside Parakeet's 25, and for translation into English.
- Does Whisper Large v3 Turbo give word timestamps?
- Not in the app. The engine returns a time per segment for Whisper, so the app uses Parakeet where it needs one per word.
- Can I use Whisper Large v3 Turbo commercially?
- Yes. It is released under the MIT license, which allows commercial use. The app ships the license text and adds no restrictions of its own.