Your recording never leaves the device

Most apps that say "Whisper" send your audio to a server and run the model there. Whisper Notes runs it on your iPhone or Mac instead, which is a harder way to build the app and a much simpler thing to explain.

Updated August 2026•10 min read

Why run the model on the device?

Every transcription app makes one decision before it makes any others: where the audio gets processed. Sending it to a server is the easier build and it buys you the largest models, which are still slightly more accurate. Running the model on the phone or the laptop is harder to get right, and it means the recording has nowhere to go.

We took the second option, and the reason is narrower than a privacy slogan. A voice recording is biometric data, and unlike a password you cannot reset it after a leak. Once the audio sits on someone else's infrastructure it is subject to their breach history, their retention policy, and whatever their terms say about training data — three things you have to keep track of for as long as the file exists.

Whisper Notes runs Whisper Large V3 Turbo, Parakeet V3 or SenseVoice on your device's Neural Engine, so there is no transcription server in the path at all. Nothing is queued, nothing is uploaded, and the app works the same in airplane mode as it does on wifi. That last part is worth testing yourself rather than taking our word for it.

What Happened to the Whisper App? Three Different Whispers

If you searched "what happened to the Whisper app," you are probably thinking of one of three very different things that share the same name.

The first is Whisper, the anonymous secret-sharing social app (whisper.sh) launched in 2012. It let people post confessions anonymously and was widely used in the mid-2010s, but it has largely shut down operations. If that is the app you were looking for, it is effectively gone.

The second is Whisper, the open-source speech recognition model OpenAI released in 2022. It never disappeared, because it was never an app in the first place. It is an AI model: the engine that converts spoken audio into text, freely available to developers but requiring technical setup to use directly.

The third is the category of native transcription apps built on that model, like Whisper Notes. These package the Whisper engine into software you can actually install and use on your iPhone or Mac, with no command line, no cloud account and no technical setup.

This page is about the third kind: what a local-first Whisper app is, how it works, and why we built ours to run entirely on-device.

The Three Whispers at a Glance

NameWhat It IsStatus
Whisper (social app)Anonymous secret-sharing app (whisper.sh), launched 2012Largely shut down operations
Whisper (OpenAI model)Open-source speech recognition AI model, released 2022Active and free, but a model rather than an app
Whisper NotesNative transcription app for iPhone & Mac built on the Whisper modelActive, on the App Store, $7.99 one-time

What does a free Whisper app cost?

A transcription tool that costs nothing usually works the same way: your audio goes to a server, the model runs there, and the recording stays long enough to be useful to whoever is running it. That is a reasonable trade for a lot of people. It is worth knowing you are making it.

Why a voice recording is different from a password

A leaked password is an afternoon of annoyance because you can change it. A leaked recording of your voice is permanent, because the thing that identifies you is the voice itself, and a few seconds of it carry enough acoustic signature to be recognised somewhere else entirely.

The part that changed recently is how little audio an attacker needs. Cloning a voice convincingly now takes seconds of sample rather than minutes, and in 2025 a cloned voice of Italy's defence minister was used in an attempted fraud against several businessmen. We are not claiming this is likely to happen to you. We are pointing out that it is the one category of leaked data you cannot rotate.

So when audio goes to a transcription service, what is being stored is not a document. It is a biometric, on infrastructure whose retention policy you do not set.

How transcription data actually leaks

The failure is rarely a dramatic break-in. It is that audio and transcripts pass through more systems than anyone drew on the diagram: the transcription vendor, the model provider behind it, a logging layer, an analytics pipeline, a backup. Each one is a place where a misconfiguration turns private into public, and the more of them there are, the less anyone can say with confidence where a given recording currently sits.

Healthcare has produced the clearest examples, where protected health information has been exposed through transcription integrations and misconfigured storage rather than through anything anyone would call an attack. Contact centres have produced another: transcripts streamed to a model while account numbers landed unmasked in debug logs.

None of that is an argument that cloud transcription is reckless. It is an argument that the number of systems holding a copy is the thing to count, and that on-device processing gets that number to one.

Which way is this heading?

In March 2025 Amazon removed the "Do Not Send Voice Recordings" setting from Echo devices. Everything said to Alexa now goes to Amazon's servers, and there is no longer a switch to turn that off.

One product removing one toggle is not a trend, but the incentive behind it is durable: training data is valuable, and a switch that reduces how much of it you collect is a switch under permanent commercial pressure. A privacy setting is a promise that can be revised in a release note.

Whisper Notes is built so there is no equivalent switch to revise. There is no server to send audio to, so "don't send it" is not a preference we are honouring; it is a description of what the app can do.

What the free tools cost over three years

Web tools that charge nothing generally reserve the right to use your audio to improve their models, which is disclosed in terms almost nobody reads. Metered APIs look cheap at $0.006 to $0.40 a minute until you transcribe an hour a week, at which point the annual bill is in the hundreds. Subscription services are the honest version of the same arrangement: Otter charges around $99 a year and tells you so.

Whisper Notes is $7.99 on the App Store and $14 for the Mac direct download, paid once. The table below is the whole pricing argument, and the fourth row is the point: the free option and the paid-once option cost about the same over three years, and they differ in what happens to the recording.

Cost per year

Service TypeYear 1Year 2Year 3Data Handling
Whisper Notes$7.99$0$0Never leaves device
Subscription Service$99$99$99Cloud processed
Per-Minute Cloud API$120-480$120-480$120-480Cloud processed
"Free" Web Tools$0$0$0Used for AI training

When you should use a cloud service instead

Frontier cloud models are still somewhat more accurate on clean audio, because they run at a size no phone or laptop can hold, and they can caption speech live in a way on-device transcription does not yet match. Those are real advantages and we are not going to pretend otherwise.

If you need live captions, if several people have to edit the same transcript while a meeting is running, or if the last fraction of accuracy matters more than where the audio sits, a cloud service is the better tool and we would rather you used one.

What we would push back on is treating that accuracy gap as the only number in the decision. For work under a confidentiality duty — clinical notes, privileged material, interviews with sources — "where does the audio go" is a question you have to be able to answer, and the shortest defensible answer is that it never went anywhere.

Browser tool, cloud API, or native app?

Searching for a Whisper app returns three quite different things wearing the same name: pages that run the model in your browser tab, APIs that run it on someone's server, and applications compiled for the device in your hand. They behave differently enough that the category matters more than the feature list.

Browser tools

When a browser tool says it processes locally, that is usually true: the audio stays in the tab. The limits are not about honesty, they are about what a tab can do. WebAssembly memory caps around 4GB in most browsers, which puts a ceiling on model size, and running inference through JavaScript costs speed that native code does not pay.

The constraint people actually run into is narrower and more annoying. A browser tool cannot keep working while you switch to another app, it reaches hardware acceleration awkwardly if at all, and closing the tab by accident loses the run with nothing to resume from. That is fine for trying a model out and frustrating as a place to keep your work.

ProcessingWebAssembly/TensorFlow.js in browser
Model SizeLimited by browser memory (~4GB)
SpeedSlower due to JavaScript overhead
PrivacyBetter than cloud, but browser has access
ReliabilityTab can crash, no background processing

Native apps

Whisper Notes is compiled for macOS and iOS and talks to the Neural Engine directly, which is the same dedicated silicon behind Face ID and computational photography. That is the whole reason the speed numbers on this page are possible: on an M4 Pro, Parakeet V3 gets through 35 minutes of audio in about 18 seconds.

The rest of the difference is unglamorous and matters more day to day. Transcription keeps running when you switch apps, the app picks up cleanly after an interruption, and the operating system's sandbox keeps it out of other apps' data. There is no upload path for your audio or its transcript, which you can check in about thirty seconds by turning on airplane mode and transcribing a file.

ProcessingDirect Apple Neural Engine access
Model SizeWhisper Large V3 Turbo, roughly 1.5GB
SpeedParakeet V3 at roughly 60x realtime on our M5 Air
PrivacySandboxed, no network permissions
ReliabilityBackground processing, system integration

Cloud APIs

Server resources are effectively unbounded, so a cloud API can run the largest models and offer things that need serious compute, live captioning among them. On clean audio it will be the most accurate option on this page.

What you give up is knowing where the recording is. It travels the internet, is processed somewhere you do not administer, and is retained under a policy you did not write and can be changed without asking you.

For a therapist under a confidentiality duty, a lawyer handling privileged material, or a journalist protecting a source, that is usually where the evaluation ends — not because the accuracy is worse, but because the question "can you tell me where this recording is" has no good answer.

ProcessingRemote servers (unlimited compute)
Model SizeLargest available models
SpeedDepends on internet and server queue
PrivacyAudio uploaded and potentially stored
ReliabilityRequires internet, subject to rate limits

What we gave up to build it this way

Building natively was the only way to make the privacy claim structural rather than contractual. "Processed locally, then synced" and "encrypted in transit" are both true statements about systems that still end up holding a copy of your audio; we wanted the version where no copy exists to hold.

The bill for that comes due in features. We cannot caption speech live while you record, because the models that do that well do not fit on a phone. We cannot run anything larger than your device can hold. And we cannot offer shared editing or team workspaces, because those need a server and we do not have one.

If any of those three is on your list, this is the wrong app and one of the cloud services is the right one. We would rather say that here than have you find out after buying.

Where is the official Whisper app for iPhone?

There isn't one. OpenAI published Whisper as an open-source model for developers rather than as a consumer product, so there is no official Whisper app on the App Store or anywhere else, and anything claiming to be one is a third-party app using the name.

On iPhone, Whisper runs through native apps that convert the model to Apple's Core ML format and execute it on the Neural Engine. Whisper Notes is one of these: on iPhone the Whisper route is Whisper Small — the same 101 language choices in a model sized for the phone — alongside Parakeet V3 and SenseVoice, fully offline on iPhone 12 or later. Whisper Large V3 Turbo is a Mac model, where there is headroom for something that size. You can record directly in the app, or import audio from Voice Memos, Files, or any app with a share sheet.

Whisper Notes is available as a one-time App Store purchase that includes the iPhone/iPad app and the Mac App Store version. No subscription or account is required. The direct-download Mac DMG is a separately sold license that is free for 5,000 words a week. There is no Android version, because the app is built on Apple's Neural Engine hardware, which has no Android equivalent.

Which model is doing the work?

The engines, and how to pick one

Whisper Notes ships three speech recognition engines in every version and runs all of them on the device. Parakeet V3 is the default. The Whisper route is Whisper Large V3 Turbo on the Mac and Whisper Small on iPhone — the same 101 language choices in a model sized for the phone. SenseVoice is the engine you select for Chinese, Cantonese, Japanese and Korean; nothing switches you to it automatically, so if you work in those languages, changing the engine is the first thing to do after installing. The Mac direct download adds a fourth from version 1.6.0, Qwen3-ASR 1.7B, still in Beta and covering 30 languages; the Mac App Store version and iPhone keep the three.
What each one is for: Parakeet V3 covers 25 European languages and posts 6.32% English word error rate on the Open ASR Leaderboard, the lowest of the three on that benchmark. On our own runs it gets through a file about five times faster than Whisper Large V3 Turbo does. Turbo's argument is breadth: 101 language choices, at 7.83% English WER on the same leaderboard. SenseVoice covers six choices — English, Simplified and Traditional Chinese, Cantonese, Japanese and Korean — and on our own runs it moves at roughly 52 times realtime, which is what makes it the one to pick for CJK material.
How it runs on the device: The models are converted to Apple's Core ML format and execute on the Neural Engine, the same silicon behind Face ID and computational photography. There is no transcription server in the path, so no network request is made while a file is being transcribed. That is a claim you can check rather than take: turn on airplane mode before you start.
Why three engines rather than one: No open model is currently best at everything. The fastest one on European languages is not the one covering a hundred languages, and neither of those is the one built for Chinese and Japanese. Shipping three and letting you choose is less tidy than naming a single best model would be, and it is the reason the app holds up outside one language family.

Specifications

AI ModelsParakeet V3 (default), Whisper Large V3 Turbo (Mac) or Whisper Small (iPhone), SenseVoice; the Mac direct download 1.6.0+ adds Qwen3-ASR 1.7B (Beta)
Languages101 choices via Whisper, 25 European via Parakeet V3, 6 via SenseVoice
Audio FormatsMP3, WAV, M4A, FLAC, AAC, OGG, WMA
Processing SpeedParakeet V3 around 60x realtime on our M5 Air; the Whisper Turbo route around 11x
File Size LimitNone imposed by the app
PlatformsiOS 18+ (iPhone 12+), macOS 11+ (Apple Silicon)

What can it actually do?

Three things carry most of the daily use: getting audio in, getting text out, and the constraint that keeps both on the device.

File Import and Batch Processing

Import audio files from any source for offline transcription. The app processes complete files rather than streaming, which allows the model to use full context for improved accuracy.

  • ✓Import from Files, Voice Memos, or any app that shares audio
  • ✓Process multiple files in sequence
  • ✓Background processing while using other applications
  • ✓Automatic organization by date and source

Export Formats

Multiple output formats for different professional workflows.

  • ✓Plain text with paragraph formatting
  • ✓SRT and VTT subtitle files for video
  • ✓Timestamped transcripts for reference
  • ✓Speaker labels for multi-person recordings
  • ✓Custom paragraph break settings

Privacy Architecture

The app is built so that your audio cannot leave your device, not as a setting but as a technical constraint.

  • ✓No network permissions requested or granted
  • ✓No cloud servers to connect to
  • ✓No analytics or telemetry collection
  • ✓Local storage only, encrypted by iOS/macOS
  • ✓No third-party processor to cover with a BAA

How accurate is it, and where does that number come from?

Published benchmarks, plus our own timings on named hardware

We do not run our own accuracy study, because a vendor grading its own homework is worth very little. The word error rates below come from the Hugging Face Open ASR Leaderboard and the FLEURS benchmark, both public and both run by people with no stake in this app. The speed figures are ours, measured on one machine we name.

Word error rate by model

ModelSourceError rateNotes
Parakeet V3 (default on Mac)Open ASR Leaderboard6.32% English WER25 European languages; 12.0% average across them on FLEURS
Whisper Large V3 TurboOpen ASR Leaderboard7.83% English WERThe widest coverage of the three: 101 language choices
SenseVoice SmallOur test, M4 Pro27-minute podcast in 13.83sBuilt for Chinese, Japanese, Korean and Cantonese

Key Findings

  • •Parakeet V3 transcribes 35 minutes of audio in 18 seconds on an M4 Pro; Whisper Turbo takes about three minutes for the same file
  • •On English, the gap between the three models is under two points of word error rate
  • •Word error rate is measured on read and prepared speech; your own recordings will do worse, and microphone and crosstalk move the number more than the choice of model does

Frontier cloud models still hold a small accuracy lead on clean audio, because they run at a size no phone or laptop can hold. That gap is the price of the audio never leaving your device. If your work does not require the last point of accuracy, the trade is usually worth it; if it does, a cloud service is the honest recommendation.

How does it compare with the alternatives?

Against cloud services, the tools already on your device, and enterprise software

Most of this table is checkable in a few minutes. The accuracy row is the one to read carefully, because the four categories are not measured on the same benchmark.

Feature Comparison

FeatureWhisper NotesCloud ServicesBuilt-in ToolsEnterprise Software
Accuracy (English WER)6.32-7.83% (Open ASR Leaderboard)Vendor-reported, mostly not on that leaderboardNot publishedNot published
Where audio is processedOn your deviceVendor serversVaries by vendorOn-premise option
Cost$7.99 iPhone, $14 Mac, paid once$0.006-0.40/minFree (limited)$500-2000/license
Languages101 choices50-100 languages10-30 languages20-50 languages
Internet RequiredNoYesSometimesDepends on deployment
File Length LimitNone imposed by the app1-2 hours typically5-10 minutesVaries

Market Position: The trade is legible once the three options sit side by side. A frontier cloud API is still a little more accurate on clean audio, and it costs $0.006 to $0.40 a minute with your recording on someone else's disk. On-premise enterprise transcription keeps the audio inside your network and starts around $500 a seat. Whisper Notes runs the open models on hardware you already own for $7.99 on iPhone or $14 on the Mac, and gives up live captioning, shared editing and the last fraction of accuracy to do it. If any of those three is on your list, one of the other two is the better purchase.

Which jobs is this a good fit for?

Three kinds of work where the recording cannot travel

Healthcare

Medical professionals use Whisper Notes for patient documentation, clinical notes, and research interviews. Because the audio never reaches a server of ours, there is no third-party processor for a BAA to cover. Whether the whole workflow meets HIPAA still depends on how the device itself is secured — and if colleagues need to open and comment on the same note, a cloud service with a signed BAA is the more practical answer.

Use Cases
  • •Patient consultation documentation
  • •Medical procedure notes
  • •Research interview transcription
  • •Telemedicine session records
  • •Clinical training content
Benefits
  • ✓No data-sharing policy to trust, because nothing is sent
  • ✓Custom vocabulary for clinical terms and drug names
  • ✓No PHI leaves the device, since there is no upload path
  • ✓A 20-minute consultation transcribes in well under a minute on an Apple Silicon Mac

Legal

Attorneys use Whisper Notes for depositions, client interviews and case preparation. The recording stays on the device, so there is no vendor holding a copy of it. What that does not do is settle your own obligations: whether a given workflow satisfies your jurisdiction's confidentiality rules still depends on how the device, its backups and its exports are handled.

Use Cases
  • •Client interview documentation
  • •Deposition transcription
  • •Case research notes
  • •Legal proceeding records
  • •Investigative interviews
Benefits
  • ✓We hold no copy of the recording or the transcript
  • ✓Custom vocabulary for case names and legal terms
  • ✓Timestamped transcripts, exported as text, SRT, VTT or JSON
  • ✓$7.99 on iPhone or $14 on the Mac, paid once, against per-minute outsourced transcription

Journalism

Reporters use Whisper Notes for source interviews and field recordings. There is no server holding the audio, so the only copies are the ones on your own devices — which moves source protection from our retention policy, where you cannot verify it, to your own device security, where you can. If the desk needs several people editing one transcript during a live event, that is a cloud tool's job.

Use Cases
  • •Source interview transcription
  • •Field recording documentation
  • •Press conference notes
  • •Research interview archives
  • •Podcast production
Benefits
  • ✓Nothing about the recording or the transcript is sent anywhere
  • ✓Works with the device offline, which is also how you check the claim
  • ✓No account, so there is no login history tying an interview to you
  • ✓Exports to text, SRT, VTT or JSON for the tools the newsroom already uses

How fast is it, and where does it fall short?

The numbers we measured, and the things the design rules out

What we measured, and on what

Speed depends on the machine and on the engine you pick, so a single multiplier for "the app" would be misleading. These are our own runs, wall-clock from start to final text, on hardware we name.

Parakeet V3, the default engine

A 17.5-minute meeting recording finished in 16.7 seconds on our M5 Air — roughly 60x realtime

25 European languages; our own run, one machine

The Whisper Turbo route

Whisper Large V3 Turbo runs closer to 11x realtime on the same path, which is the price of its 101 language choices

Our own runs; Turbo is the wide-coverage route, not the fast one

Older and smaller devices

An iPhone works through a file slower than an Apple Silicon Mac does, and the gap widens with the age of the chip. Long files still complete; they take proportionally longer

iPhone 12 or later, or an Apple Silicon Mac

Storage

Transcripts are tiny, around 0.1MB per hour of audio. The models are the real footprint: Whisper Large V3 Turbo is roughly 1.5GB, and the other two are smaller

Parakeet V3 is built into the iPhone app; the other models download inside the app

Known Limitations

Running the models on the device sets hard limits, and they are easier to check before buying than after. The Mac free tier covers 5,000 words a week for exactly that reason.

Device Requirements

Requires iPhone 12 or later, or an Apple Silicon Mac. Older hardware lacks the Neural Engine performance to make this practical.

Impact: Not compatible with devices more than 4-5 years old

Processing Time

Processing time scales with the length of the recording. There is no way around the physics of on-device computation.

Impact: At the rates above, a four-hour file is minutes on a Mac and longer on an iPhone

Audio Quality Dependency

Poor audio or loud background noise reduces accuracy, and the model cannot recover information that is not in the signal.

Impact: Microphone placement and crosstalk move the error rate more than the choice of model does

No Live Captioning

The app transcribes a recording after it finishes rather than captioning it as you speak. Live captioning of that quality needs a model class that does not fit on a phone, and processing the whole file also gives the model more context to work with.

Impact: If you need captions on screen during the event, this is the wrong tool

Single Language Per Recording

Rapid language switching within a single recording reduces accuracy. The model performs best with consistent language throughout.

Impact: Best results with one primary language per file

So who is this for?

Whisper Notes runs Whisper Large V3 Turbo, Parakeet V3 and SenseVoice on your own device. There is no transcription server in the path, which is the whole design and also the whole limitation.
What you get: Word error rates of 6.32% to 7.83% on English from the Open ASR Leaderboard, 35 minutes of audio transcribed in about 18 seconds on an M4 Pro, and a recording that has nowhere to go. $7.99 on the App Store, which covers iPhone, iPad and the Mac App Store version; $14 for the Mac direct download, which is free for 5,000 words a week.
What you give up: Live captions while you speak, shared editing during a meeting, and the last fraction of accuracy that only a frontier-sized cloud model can reach. None of the three is coming, because all three need a server.
The honest recommendation: If your team has to read and comment on transcripts together, use Otter or Notta. If you want the largest feature set on a Mac and never work from a phone, look at MacWhisper. If what you need is that the audio stays where it was recorded — clinical notes, privileged material, interviews with sources — that is the case this app was built around, and you can verify the claim in thirty seconds with airplane mode.

Download Whisper Notes

Transcription that runs on your iPhone or Mac. Free trial on Mac, so you can check the claim before paying for it.

Available on iOS (iPhone 12+) and macOS (Apple Silicon). $7.99 one-time purchase. No subscriptions. No in-app purchases.