Voice Studio

How long does it take to transcribe a voice memo on iPhone?

Short answer: there is no fixed number, because the wait depends on which of two different processes ends up handling the recording — and that is not something you choose yourself.

Type into a note and the words appear as you go. Hand a recording to a phone for transcription and something else happens — sometimes the text is back before you have set the phone down, sometimes it takes noticeably longer, and there is no timer telling you which one you are in for. Neither outcome means anything is wrong. They are two different processes, and which one runs decides the wait.

Transcribing is a step you take, not something recording does automatically

A recording sitting in your library is just a recording until you ask for a transcript — turning speech into text is a separate action, not something that happens invisibly the moment you press stop. That matters here because the clock only starts when you actually request it, and everything between that request and the text appearing is the whole question this guide is about.

Two different paths, and they wait differently

The on-device path

When the language has a model already on the phone, the audio is processed right there, using the phone’s own processor and neural engine. Nothing is uploaded, so there is no network step to wait on — the entire wait is just however long it takes that hardware to work through the audio. A short, clear clip in a well-supported language is usually the fastest thing this process can do.

Apple’s server, as the fallback

This path only runs if the on-device attempt already happened and came back empty or errored. It means the audio actually leaves the phone this time: uploaded, processed on Apple’s side, and the text sent back. Three network steps instead of zero, so a slow connection or a weak signal shows up here in a way it never does on-device.

Why a fallback makes a file slower, not just different

The fallback is not a shortcut around the local attempt — it comes after it. A recording that ends up needing the server was tried locally first, that attempt failed or returned nothing, and only then does the second, network-based attempt begin. So a file that takes the fallback path is not transcribed instead of locally, it is transcribed twice, one after the other, which is exactly why it can take noticeably longer than a file the on-device pass handled cleanly on the first try.

What actually changes the wait

Long recordings specifically

A ninety-minute lecture is not obliged to sit through processing as one unbroken block — long audio is commonly split into segments by systems built this way, so a very long file can come back in pieces rather than all at once or not at all. That is more an implementation detail than something to plan around, but it is worth knowing mainly so a long recording taking noticeably longer than a two-minute note does not read as a sign that something has stalled.

What to actually do while it processes

For anything long, plug the phone in. Running a speech model taxes the processor and the neural engine, and a long file will warm the phone up and use real battery — the same reason a demanding game or a long video call drains a phone faster than scrolling does. Beyond that, there is genuinely nothing to click faster. It is a one-time cost per recording, not something that keeps running once the text is back.

How Voice Studio handles this

Voice Studio always attempts the on-device path first and only retries against Apple’s server if that local pass comes back empty or errors. A short recording in a well-covered language on capable hardware is usually the fastest case there is, and it never touches the network at all. Something in a less common language, or recorded somewhere noisy, is more likely to take the slower route — not because anything failed permanently, but because it is quietly being tried a second way after the first did not produce anything. Transcription covers 38 languages; on-device coverage among them is narrower and is decided by Apple, not by the app. Either way, there is no account and no server of ours anywhere in the process.

Common questions

Does transcription happen while I’m recording, or only after I stop?

After. Transcribing is a separate step you take once a recording exists — it does not run live while you talk, the way dictation into a text field does.

Does a longer recording always take proportionally longer to transcribe?

Roughly, but not exactly — language coverage, audio quality and whether the on-device attempt succeeds all affect the wait more than length alone does.

Why did one transcript come back almost instantly and another take much longer?

The fast one was very likely handled entirely on the phone. The slow one probably needed the server fallback, which means it was effectively transcribed twice — once locally, then again over the network after that first attempt came back empty.

Does transcribing a recording use noticeably more battery?

Running a speech model on the phone uses the processor and the neural engine, so a long file will warm the phone up and use some battery. Plugging in for a long transcription is a reasonable habit, the same as for anything else that keeps the processor busy.

Try it in Voice Studio

Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.

Download on the App Store

Free to download · iPhone and iPad · iOS 16.4 or later