Can you transcribe a conversation in two languages on iPhone?
Not in one pass, no. iOS speech recognition is told a single language to expect and holds to it for the whole recording, so a conversation that moves between two languages gets one of them transcribed and the other one forced into words that do not belong to it. The fix is not a setting — it is knowing which of two workarounds actually fits the recording you have.
This comes up in ordinary places: a call with a bilingual relative that drifts between English and Spanish mid-sentence, a lecture that opens in one language and takes questions in another, a language-learning session where you deliberately switch back and forth. The recording itself has no trouble with any of that — a microphone does not care what language is arriving. The transcript is where the trouble shows up.
Why iOS cannot just follow along
Speech recognition on iPhone is not listening for whichever language happens to be spoken and adjusting on the fly. It is told, before the pass starts, which one language to expect, and it spends the entire recording trying to fit what it hears into that language's sounds and vocabulary. That is true whether the pass runs on-device or falls back to Apple's server — the language is a setting made in advance, not something detected as the audio plays.
For a recording in one language throughout, that design is invisible. For a recording that switches languages, it means the setting is only ever right for part of the audio.
What the transcript actually looks like
Not an error, and not a gap. The stretches spoken in the language the recognizer was told to expect come back close to normal. The stretches spoken in the other language get mapped onto the nearest-sounding words in the expected language anyway, because that is the only vocabulary the recognizer has available to guess from — the same thing that happens when the whole recording is set to the wrong language, just confined to the portion that actually was. The result reads as real words in the wrong place, not as a visible error you can spot at a glance, which is what makes it worth checking a transcript against the audio before trusting a section you were not paying close attention to when it was spoken.
Regional variants of the same language work the same way for this purpose. English (US) and English (UK) are separate entries a transcription setting can point at, not accents one model absorbs, so a recording that moves between them has the same shape of problem as one that moves between two unrelated languages, just milder.
Picking which language to set
When one language clearly dominates the recording — a lecture mostly in one language with a few questions in another, a call that is mostly one person speaking their own language — set the transcription language to whichever one covers more of the audio. That gets the majority of the transcript right and leaves the minority stretches as the part to read against the recording and correct by hand, which is a smaller job than it sounds like when it is a few sentences rather than the whole thing.
When the two languages are closer to even, there is no setting that does better than the other — either choice gets roughly half the recording right and mangles the rest. That is the situation where splitting the recording, rather than picking a language and living with it, is worth the extra step.
When splitting the recording is worth it
Splitting only helps if the languages fall into distinct stretches of time rather than alternating word by word inside the same sentence — a lecture that is one language for the first half and another for the second, a conversation where one person speaks entirely in one language and the other entirely in theirs. Cut at the point where the language changes, transcribe each piece with its own language setting, and each half gets the same accuracy a single-language recording would. Fast back-and-forth switching inside a sentence does not split cleanly at all — there is no cut point that separates the languages, so a transcript of that kind of speech is going to need a hand-corrected pass regardless of how it is set.
How to actually do the splitting
Cutting an .m4a into pieces is a job for a trim tool, and it is worth knowing which apps have one before assuming the app you record in does. Apple's own Voice Memos includes a basic trim, and GarageBand handles more involved cuts. Once each piece is a separate file, importing it into whatever app you transcribe with and setting the correct language for that piece is the same ordinary transcription pass as any other recording — nothing about having been split changes how it runs.
Living with it instead
For a lot of what this comes up for — a family conversation, a quick voice note that drifts languages out of habit rather than structure — splitting the file is more effort than the transcript is worth. Setting the dominant language and hand-correcting the minority stretches afterward gets a usable, readable transcript with one pass and one edit, which for most personal recordings is the faster route to something you would actually reread.
How Voice Studio handles this
Voice Studio transcribes against a single language setting, chosen from 38 supported languages, and that setting applies to the whole recording it is run on — there is no per-segment or automatic-detection option, because iOS speech recognition itself does not offer one. The transcript is editable text, so correcting the stretches transcribed in the wrong language is a normal hand edit, the same as fixing any other wrong word. Voice Studio does not include a trim tool, so splitting a recording by language means cutting it in Voice Memos or GarageBand first and importing each piece back in — importing existing audio is supported, and each imported piece transcribes exactly like a recording made in the app. None of this is limited in the free version: recording, importing, transcribing and export are unrestricted.
Common questions
Does iOS auto-detect which language is being spoken?
No. A transcription pass is told one language to expect in advance and spends the whole recording trying to fit what it hears into that language, regardless of what is actually being spoken.
What happens to the part of a recording spoken in a language I did not set?
It gets mapped onto the nearest-sounding words in the language you did set, rather than failing outright or switching languages on its own. It reads as real text in the wrong place, so it is worth checking against the audio rather than assumed correct.
Is it worth splitting a recording into separate files by language?
Only if the languages fall into distinct stretches of time rather than alternating inside the same sentence. A lecture that changes language halfway through splits cleanly; fast back-and-forth switching does not, and needs a hand-corrected pass either way.
Does Voice Studio support more than one transcription language per recording?
No — one language setting applies to the whole recording, the same limit iOS speech recognition has generally. Importing a trimmed piece and transcribing it separately with its own language setting is the workaround.
Try it in Voice Studio
Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.
Free to download · iPhone and iPad · iOS 16.4 or later