Voice Studio

Why did my voice memo transcription go to the cloud?

It almost certainly was not a setting you changed. Apps built on Apple’s Speech framework are structured to try the on-device path first, every time, and only reach the server as a second attempt when the local one comes back empty or errors out. If a transcription went to the cloud, something about that specific recording made the first attempt fail.

This question usually comes up after noticing something indirect — a transcript that took longer than usual, a spike in cellular data, or just reading a privacy explanation and wondering whether it applied to a transcript you already have. The honest answer is that you cannot know which recordings went which way after the fact without having tested them at the time, but you can know exactly what would have caused it, which is usually more useful.

The order of operations, not a setting

This is not a toggle sitting in Settings that got flipped. It is a sequence built into how the app calls Apple’s speech recognition: attempt the transcription with on-device recognition requested, and only if that attempt returns nothing or throws an error, retry the same audio with on-device recognition turned off. The second attempt is what reaches Apple’s server. Most of the time the first attempt succeeds and the second one never runs at all — the times it does run are the ones this question is about.

That structure exists for a plain reason: the alternative to retrying is handing you an empty transcript and stopping there. An empty result is not a safer failure, just a more useless one. The retry trades a small amount of certainty about where the audio went for a transcript that actually exists.

What makes the first attempt fail

The language has no on-device model

On-device speech recognition is built by Apple language by language, not switched on universally. If the language you were transcribing in has no local model on your phone, the on-device attempt fails immediately, every time, regardless of how the recording sounds. There is nothing about audio quality that would have changed that outcome.

The model exists but was never downloaded

A language being supported and a language being present on your specific phone are two different things. The on-device model for a language typically arrives after you have added that language somewhere the system notices — the keyboard or dictation settings — and the phone has had time on Wi-Fi to fetch it. A language you have never used for dictation may simply not have its model on your device yet, even if that language has on-device support in general.

The audio itself defeated a smaller model

On-device models are compact enough to run on a phone, which makes them less forgiving than the server model behind them. Quiet or distant speech, heavy background noise, a strong accent, or several people talking over each other are the kinds of audio that a local model is more likely to return nothing for, while the larger server model handles the same clip without trouble. This is the case where the language is genuinely supported on-device and still ends up going to the server, purely because of how the recording sounds.

It is a per-recording decision, not a per-app one

Because the trigger is "did this specific attempt fail," two recordings in the same language, made minutes apart, can take different paths. A clear one made close to the phone can transcribe entirely on-device; a quieter one made from across the room, in the same language, can fail locally and fall back. Nothing in the settings changed between them — the audio itself is what decided it.

What was actually sent, if it did fall back

When the retry runs, the audio is what gets transmitted to Apple’s speech recognition service — not a summary, not just partial text. That service is a first-party Apple path, the same one behind keyboard dictation, rather than a third-party company the app chose. First-party is not the same claim as local, and it is worth keeping those two separate: the audio did leave the phone, even though it did not leave the ecosystem of a service Apple already documents and operates.

How to find out before it matters, rather than after

Since the split happens per recording, the only way to know in advance is to remove the option to fall back and see what happens: put the phone in Airplane Mode with Wi-Fi off, then transcribe. With no network reachable, the retry has nowhere to go. A transcript coming back means the on-device attempt succeeded on its own; nothing coming back means that recording — in that language, in that room — was always going to need the server.

That test is specific to the recording you run it on. A different room, a different language, or a different speaker can land on the other side of the same split, so it is worth repeating for anything where the answer genuinely matters rather than trusting one result to generalise.

How Voice Studio handles this

Voice Studio always attempts on-device transcription first, for every recording in every one of its 38 transcription languages, and only retries with on-device recognition turned off if that first pass comes back empty or throws an error. That retry is what sends audio to Apple’s speech service, and it is the only path on which audio leaves the phone at all — Voice Studio has no account, no server of its own, and no analytics SDK, so nothing about a recording goes anywhere else regardless of which attempt produced the transcript. If you want certainty for a particular recording rather than an explanation after the fact, transcribing it in Airplane Mode settles the question directly: either the on-device pass handles it, or you get nothing back, and either way no audio was sent.

Common questions

Can I turn off the cloud fallback and only ever get on-device transcripts?

Not as a setting, but you can force the same outcome by transcribing in Airplane Mode. With no network available, only the on-device attempt can produce a result — if it fails, you get an empty transcript instead of one sent to the server.

Does a transcription going to the cloud mean the app chose to send it somewhere?

It means the on-device attempt for that specific recording came back empty or errored, and the app retried without requesting on-device recognition. It is a response to a failed local attempt, not an independent decision made in advance.

Why did the same language transcribe locally yesterday and go to the server today?

Because the trigger is the individual recording, not the language as a whole. A clearer or closer recording can succeed on-device while a quieter or noisier one in the same language fails locally and falls back.

Is Apple’s server transcription sent to a third party?

No. It is a first-party Apple service, the same one used for keyboard dictation. The audio does leave the phone, which is the part worth knowing, but it is not handed to an outside company.

Try it in Voice Studio

Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.

Download on the App Store

Free to download · iPhone and iPad · iOS 16.4 or later