Whisper Notes Guide Hub

Whisper Models and Local Transcription Benchmarks

Updated September 28, 2026

Compare local speech recognition models by speed, accuracy, language coverage, and Apple Silicon performance.

Every local speech model we ship or have benchmarked — what each is good at, how fast it runs on Apple Silicon, and which one to pick for your language. All numbers below come from our published benchmark articles.

3 topic clusters10 guides and comparisons

Local Speech Models at a Glance (2026)

The short version: on a Mac, use Parakeet V3 for English and European languages (it's the default in Whisper Notes), SenseVoice for Chinese, Japanese, Korean, and Cantonese, and Whisper Large-v3 Turbo for everything else in its ~100-language long tail.

ModelBest forAccuracySpeed on Apple SiliconSpeed vs Whisper Large V3 TurboMemory neededLanguagesIn Whisper Notes
Parakeet V3 (NVIDIA, 0.6B)English & European languages6.32% English WER (Open ASR Leaderboard)~10× faster than Whisper Turbo7.2× faster (103× vs 14.3× realtime, M4 Pro)0.09 GiB peak25Yes — default on Mac
Whisper Large-v3 Turbo (OpenAI, 809M)Widest language coverage7.83% English WER (Open ASR Leaderboard)~5× faster than Large V3Baseline1.6 GB~100Yes — optional download (~1.6 GB)
SenseVoice Small (FunAudioLLM)Chinese, Japanese, Korean, CantonesePurpose-built for CJK27-min Chinese podcast in 13.83 s (M4 Pro)9.0× faster (118× vs 13.1× realtime, M4 Pro)0.95 GiB peakCJK focusYes — since v1.5.0
Whisper Large V3 (OpenAI, 1.55B)Reference baselineNear-identical to TurboBaseline (slowest here)About 5× slower than TurboNot measured — not shipped~100No — Turbo supersedes it
Voxtral (Mistral)Speech + language understandingSee our benchmark articleServer-class modelsNot measuredNot measuredMultilingualNo — covered in blog

WER = word error rate; lower is better. English WER comes from the Hugging Face Open ASR Leaderboard; 25-language averages from FLEURS. Speed figures are our in-house tests on an M4 Pro running Whisper Notes — full methodology in the linked benchmark articles below. All models listed as "in Whisper Notes" are included in the one-time price. Memory is each engine's peak process footprint in our own FLEURS run on an M5 MacBook Air — that is how Parakeet V3 and SenseVoice Small were measured. The 1.6 GB for Whisper Large V3 Turbo is the working-set figure the app documents, not a peak we measured. The two speed ratios are computed from realtime factors measured on the same audio on an M4 Pro.

4

Highest-traffic model benchmarks

Start here when the searcher is comparing local ASR model speed, decoder layers, silence hallucinations, and Apple Silicon performance.

3

Model-specific local transcription pages

These pages connect model searches to practical Mac, iPhone, and multilingual transcription workflows.

3

Use models in real workflows

Where these models actually get used: iPhone recording, Mac Fn-key dictation, and end-to-end offline transcription workflows.

3

More Guide Hubs

Explore the other Whisper Notes topic hubs for models, alternatives, offline workflows, and Mac dictation.