Qwen3-ASR 1.7B vs Whisper Large V3 Turbo: Which Languages Win (Mac Benchmark)

August 25, 2026
·
11 min read
·Whisper Notes Team

We wanted to answer one question: is there a local model that transcribes your language more accurately than Whisper Large V3? So we tested these models on the FLEURS dataset.

FLEURS is a Google dataset of read speech in 102 languages. We took about 150 sentences from each language and ran the same clips, in the same order, through five models:

  • Parakeet V3
  • Whisper Large V3 Turbo
  • SenseVoice Small
  • Whisper Small
  • Qwen3-ASR 1.7B

The results look like this.

Qwen3-ASR 1.7B against Whisper Large V3 Turbo — our own FLEURS run, one M5 MacBook Air, read speech, short clips.

Qwen3-ASR 1.7B Whisper Large V3 Turbo
Languages supported 30 101
Clear advantages Cantonese, Hindi, Thai, Vietnamese, French and 2 more (7) Finnish, Hungarian, Greek and 8 more (11)
Speed 9.6x realtime 2.6x realtime
Peak memory 2.43 GiB 1.87 GiB
Download about 2.5 GB about 1.6 GB

Outside each model's strong languages, the accuracy difference is small.

Whisper Large V3 vs Qwen3-ASR: the detailed comparison

All 30 languages, sorted from Qwen's largest win to its largest loss. Lower is better in both columns.

Our own FLEURS run, one M5 MacBook Air, read speech. CER for Chinese, Cantonese, Japanese, Korean and Thai, WER for the rest; green marks a lead wider than the paired 95% error bars.

Language Qwen3-ASR 1.7B Whisper Large V3 Turbo
Cantonese 6.44 36.75
Hindi 13.19 29.64
Thai 6.53 13.40
Vietnamese 4.91 8.10
Chinese 6.49 7.97
Indonesian 4.91 6.31
Arabic 14.32 15.72
French 4.57 5.97
German 2.85 4.05
English 4.49 5.49
Korean 3.54 4.25
Persian 30.32 30.93
Portuguese 5.17 5.44
Spanish 3.38 3.62
Italian 2.58 2.58
Japanese 5.74 5.30
Russian 5.65 5.13
Macedonian 18.57 17.46
Malay 10.62 9.22
Dutch 7.36 5.69
Turkish 8.81 5.95
Polish 12.59 5.45
Danish 22.24 13.63
Swedish 19.80 8.78
Romanian 20.20 8.36
Czech 24.59 11.94
Filipino 23.31 9.72
Greek 30.22 13.94
Finnish 27.71 7.11
Hungarian 36.63 14.59

Our own paired run on the FLEURS test split, revision 70bb2e84, about 150 sentences per language, one M5 MacBook Air.

The seven Qwen3-ASR 1.7B clearly wins. Error rate in percent, Qwen against Turbo: Cantonese (6.4 vs 36.8), Hindi (13.2 vs 29.6), Thai (6.5 vs 13.4), Vietnamese (4.9 vs 8.1), French (4.6 vs 6.0), German (2.9 vs 4.1), English (4.5 vs 5.5). Cantonese and Thai are scored by character, the rest by word. Its advantage sits with the character-scored languages: across Chinese, Cantonese, Japanese, Korean and Thai it averaged 5.7% against Turbo's 13.5%, while across the other twenty-five, scored by word, it averaged 14.4% against Turbo's 10.2%.

The eleven Whisper Large V3 Turbo clearly wins. Error rate in percent, Qwen against Turbo, all scored by word: Hungarian (36.6 vs 14.6), Greek (30.2 vs 13.9), Finnish (27.7 vs 7.1), Czech (24.6 vs 11.9), Filipino (23.3 vs 9.7), Danish (22.2 vs 13.6), Romanian (20.2 vs 8.4), Swedish (19.8 vs 8.8), Polish (12.6 vs 5.5), Turkish (8.8 vs 6.0), Dutch (7.4 vs 5.7).

The twelve that are close or level. Spanish (3.4 vs 3.6), Italian (2.6 vs 2.6), Russian (5.7 vs 5.1), Portuguese (5.2 vs 5.4); Chinese (6.5 vs 8.0), Japanese (5.7 vs 5.3) and Korean (3.5 vs 4.3) are scored by character. Then Arabic, Indonesian, Persian, Malay and Macedonian. Every gap here is small and sits inside the 95% band, so this run cannot separate the two models.

Cantonese is the one row that changes a decision: 6.44% against 36.75%, and SenseVoice Small — which we had published as the Cantonese answer — scored 37.23% CER on the same clips. On this read-speech set neither of those clears the bar, so the older recommendation now comes with that qualifier.

How fast is Qwen3-ASR 1.7B?

All five engines, measured in the same harness on one machine (M5 MacBook Air), so they are at least comparable to each other.

Median realtime factor across each engine's languages in the same FLEURS run, one M5 MacBook Air, read speech.

Engine Languages Median speed in this run (short clips)
SenseVoice Small 6 choices 162.3x
Parakeet TDT 0.6B v3 25 83.0x
Qwen3-ASR 1.7B 30 named 9.6x
Whisper Small 101 7.0x
Whisper Large V3 Turbo 101 2.6x

These are short-clip harness figures and they are lower than what the same engines do on a long recording. Use the column to compare the engines against each other inside this one run, and nothing else.

Parakeet V3 and SenseVoice Small are roughly an order of magnitude quicker than Qwen3-ASR in the same harness. On long files rather than short clips, Parakeet has run at 103x realtime on an M4 Pro and SenseVoice at 52x or better — a different measurement on different audio, pointing the same way.

Qwen3-ASR sits in the middle. It was not the slowest engine in this short-clip run — Whisper Large V3 Turbo was — but it varies more by language than the others, from 2.8x on Greek to 13.4x on Japanese.

What does Qwen3-ASR cost in memory and disk?

Peak memory is the maximum process footprint observed across items in the same FLEURS run, one M5 MacBook Air, read speech, short clips.

Engine Download Peak memory in this run
Parakeet TDT 0.6B v3 465 MB 0.09 GiB
Whisper Small 600 MB 0.87 GiB
SenseVoice Small 827 MB 0.95 GiB
Whisper Large V3 Turbo about 1.6 GB 1.87 GiB
Qwen3-ASR 1.7B about 2.5 GB 2.43 GiB

Qwen3-ASR is the heaviest of the five on both axes: about 2.5 GB to download and 2.43 GiB of peak process memory in this run, against about 1.6 GB and 1.87 GiB for Whisper Large V3 Turbo.

On an 8 GB machine there is a lighter path. Parakeet V3 covers 25 languages in 465 MB and 0.09 GiB, and Whisper Small covers 101 in 600 MB and 0.87 GiB.

How we measured: FLEURS, 30 languages, one Mac

Qwen3-ASR 1.7B ran as an MLX 8-bit build, Whisper Large V3 Turbo as a whisper.cpp ggml build on Metal. The data is the Google FLEURS test split at revision 70bb2e84 (CC-BY-4.0): 83 of its 102 languages, the same clips in the same order for every model; the paired comparison is Qwen's 30, about 150 sentences and 30 minutes of audio each. WER for space-delimited languages, CER for Chinese, Cantonese, Japanese, Korean and Thai, never merged into one number; error bars are a 10,000-resample item-level bootstrap. One M5 MacBook Air, 32 GB, macOS 27.0, no empty outputs or failures from either model.

FLEURS is read speech, not a meeting and not dictation. Read-speech rankings have failed to survive real audio twice in our own testing, so we do not extrapolate them to long recordings: these tables say whether a model is plausible for your language, not what accuracy you will get. The real-meeting evaluation of Qwen3-ASR 1.7B will be a separate post.

Should you use Qwen3-ASR instead of Whisper Large V3?

Language by language, this is how we would answer that inside Whisper Notes for Mac.

  • If you transcribe Cantonese, Hindi, Thai, Vietnamese, French, German or English: pick Qwen3-ASR 1.7B.
  • If you transcribe Chinese, Korean or Spanish: it led by a hair, and a hair is all we will claim. Either model is fine there, as it is for Japanese, Italian, Persian, Arabic, Indonesian, Russian, Portuguese, Macedonian and Malay.
  • If you transcribe Finnish, Hungarian, Greek, Czech, Swedish, Danish, Polish, Romanian, Filipino, Turkish or Dutch: stay on Whisper Large V3 Turbo. It won all eleven clearly.
  • If speed or disk space is what you care about: not this engine. Parakeet V3 and SenseVoice Small were an order of magnitude faster, at a fraction of the download, and Whisper Small covers 101 languages in 600 MB at 7.0x.
  • If you are on iPhone or the Mac App Store build: the engine is not there yet. SenseVoice is the Cantonese pick on those channels.

This is Whisper Notes for Mac's on-device port of Alibaba's open Qwen3-ASR: the open weights converted to Apple MLX at 8-bit, running on your Mac's own GPU, with no audio going to Alibaba's servers. It is not the same model as the 0.6B that SenseVoice replaced in Direct Download (DMG) 1.5.0.

How to use it (Mac Direct Download, DMG 1.6.0 and later): Settings, then Transcription Model, then Qwen3-ASR 1.7B. The model downloads inside the app on first use and is about 2.5 GB. It needs macOS 15 or later on Apple silicon, with 16 GB of memory recommended; the rest of the app runs on macOS 14. The per-engine language list is on our language support page. The local CLI and MCP server accept "qwen", "qwen3" and "qwen3-asr".

If you would rather not switch, nothing has to change: Parakeet V3 is still the default.

Frequently Asked Questions

Is Qwen3-ASR more accurate than Whisper Large V3 Turbo?

It depends on the language, and the split is close to even. On our own FLEURS run inside Whisper Notes for Mac, across the 30 languages both models cover, Qwen3-ASR 1.7B was ahead in 14 and Whisper Large V3 Turbo in 15; Italian lands on an exact tie — identical to every decimal we kept. Qwen led clearly on Cantonese (6.44% vs 36.75% CER), Hindi (13.19% vs 29.64% WER), Thai, Vietnamese, French, German and English. Turbo led clearly on Finnish (7.11% vs 27.71% WER), Hungarian, Greek, Czech, Swedish, Danish, Polish, Romanian, Filipino, Turkish and Dutch. Twelve of the thirty gaps sit inside the error bars and should be read as ties. Choose Qwen3-ASR if your language is in its winning column and you are on the Whisper Notes Mac Direct Download (DMG) build; choose Whisper Large V3 Turbo for everything else, and for every language outside Qwen's 30, since Whisper reaches all 101 choices.

Can I use Qwen3-ASR on iPhone or from the Mac App Store?

Today it is in the Mac Direct Download (DMG) build of Whisper Notes from version 1.6.0 onwards. The Mac App Store version is planned to pick up the newer engines in later releases, and we do not have a date for it. In the meantime the iPhone app and the Mac App Store build have three engine families — Whisper, Parakeet V3 and SenseVoice — and every language Qwen recognises is also covered by Whisper there. Counted as models, the Mac App Store build offers four today and the Direct Download build five, because Whisper ships as both Small and Large V3 Turbo. If you need Cantonese on iPhone, SenseVoice is the engine to use there.

Qwen3-ASR or SenseVoice for Chinese, Japanese and Korean?

Both are close on accuracy and very far apart on speed. On our FLEURS run, Qwen3-ASR scored 6.49% CER on Chinese, 5.74% on Japanese and 3.54% on Korean; SenseVoice Small scored 7.69%, 6.78% and 7.85% on the same clips. SenseVoice ran at a median 162x realtime in that run against Qwen's 9.6x, downloads 827 MB against about 2.5 GB, and is available on iPhone and both Mac builds. Choose SenseVoice unless you are transcribing Cantonese, where Qwen3-ASR scored 6.44% CER against SenseVoice's 37.23% and the gap is large enough to be worth the wait.

What are the system requirements for Qwen3-ASR 1.7B?

An Apple silicon Mac on macOS 15 or later, with 16 GB of memory recommended, plus about 2.5 GB of disk for the model. That is higher than the rest of Whisper Notes for Mac, which runs on macOS 14 and later. In our benchmark run the model peaked at 2.43 GiB of process memory, the highest of the five engines. On an 8 GB Mac, Whisper Small covers the same 101-language list in 600 MB and Parakeet V3 covers 25 European languages in 465 MB.

Does the audio leave my Mac when I use Qwen3-ASR?

No. Alibaba also offers Qwen3-ASR as a cloud service through the DashScope Real-time and FileTrans APIs, where audio is uploaded to their servers. Whisper Notes does not use those. It runs the open weights locally as an MLX 8-bit model on your Mac's GPU, with no account and no API key, and transcription works with the network off. That is true of every engine family in the app, on every channel.

Why do these numbers not match the accuracy I get on my recordings?

Because FLEURS is short, clean, read speech, and your recordings are not. Our run is a calibration on a public dataset: about 150 sentences per language, one M5 MacBook Air, the same clips for every model. It is useful for asking whether a model is plausible for a language, and it is not a prediction of your accuracy on a meeting, on noisy audio, on accented speech or on a two-hour file.

Try it

Qwen3-ASR 1.7B is in Whisper Notes for Mac 1.6.0 and later, Direct Download (DMG) only. Settings, then Transcription Model. The free tier covers 5,000 words a week, enough to run your own language through it and disagree with the tables above.

If you find a language where our numbers do not match your experience, email mac@whispernotes.app. Full release notes: whispernotes.app/changelog.

Sources for everything above: our FLEURS run on the Google FLEURS test split at revision 70bb2e84, the Qwen3-ASR 1.7B model card, and the per-engine language table on our language support page, which is generated from the app's own source.