Does iPhone transcription work differently for US English vs UK English?
Yes — they are two separate options in the transcription language list, not one "English" setting with an accent it figures out on its own. Picking the wrong regional variant will not wreck a transcript the way a genuinely different language does, but it costs you accuracy in specific, predictable places.
It is easy to assume "English" is a single box to tick, the same way the app has a single interface language. It is not. iOS speech recognition treats regional variants of a language as distinct options — English (US), English (UK) and English (India) are three separate entries a transcription setting can point at, not three accents one English model quietly handles at once.
Why English is more than one option
A speech recognizer is matching sound against a model trained on speech and text from a specific place, and a lot changes region to region beyond the accent itself: which words are common, how names and places are said, and which spellings are considered correct. Apple's speech framework covers 38 recognition locales in total, and several widely spoken languages are split into regional variants rather than offered once — English, Spanish, Portuguese and Chinese all work this way, not just English.
The setting you pick is not about where you happen to be standing. It is about which model most closely matches the speech actually going into the microphone, which is usually the region you learned to speak in, regardless of where you are recording from today.
What actually changes between regional variants
Vocabulary and names
A model trained on one region's speech has seen that region's place names, brand names and everyday words far more often than another region's. A recognizer set to English (US) has a weaker prior for a British place name or a term that is common in Indian English but rare in American English, and the reverse is just as true. Neither model is broken — each is simply better calibrated to what it was trained to expect.
Spelling conventions
This is the one people notice fastest and understand least: the text a recognizer produces generally follows the spelling conventions of the region it is set to, the same way iOS keyboard dictation does. A recognizer set to UK English tends toward "colour" and "organise"; one set to US English tends toward "color" and "organize". Neither is a mistake — it is the model doing exactly what its regional setting tells it to, and it is worth knowing about before you assume a spelling in a transcript was misheard rather than simply regional.
Whether the on-device model is even installed
Coverage for the on-device path is not identical across every regional variant of a language, and it is not something an app controls — it depends on Apple's current models and on whether that specific variant's language pack has actually been downloaded to your phone. A widely used variant may run locally while a less common one for the same language leans on the server more often. Which is which changes between iOS releases, so it is worth treating as a moving target rather than something you memorise once.
What it looks like when you have the wrong one selected
This failure mode is quieter than picking an entirely wrong language. Setting the recognizer to a genuinely different language against English speech tends to come back empty or as nonsense text that does not correspond to anything said. Setting it to the wrong regional variant of the right language usually still produces a mostly readable transcript — it just has more misses than it should, concentrated in exactly the words that differ most from the model's own region: place names, brand names, slang, and the odd vowel-heavy word an accent shapes differently than the model expects. If a transcript is broadly right but keeps stumbling on specific, recognisable kinds of words, the regional setting is worth checking before you blame the recording.
Picking the right one when you are not sure
Match the setting to how the speaker actually talks, not to where the phone happens to be. Someone from London working in New York should still pick English (UK) for their own voice memos; a recognizer set to English (US) is not going to become more accurate because the recording happened to be made on American soil. For a recording of several people from different regions, there is no setting that serves everyone at once — pick whichever region the loudest or most consistent voice in the recording actually speaks, the same trade-off as any recording that mixes languages.
If you genuinely do not know which region's English is the closest match, the broader, more widely used variant is usually the safer default over a narrower one — it has typically had more training data behind it, even if it is not a perfect match for the accent in the room.
It is not only English
The same pattern applies anywhere a language has more than one regional entry in the transcription list: Spanish splits into variants including Spain and Mexico, Portuguese into Portugal and Brazil, and Chinese includes both Mandarin and Cantonese regional entries. The underlying advice does not change with the language — the setting should match how the person actually speaks, and getting it slightly wrong degrades a transcript in specific, recognisable ways rather than breaking it outright.
How Voice Studio handles this
Voice Studio's transcription language setting is chosen from the app's full list of 38 recognition locales, which includes these regional splits rather than a single generic entry per language. It is a free setting, not something gated behind Pro, and it is independent of the app's own interface language — changing one does not move the other. Transcription always attempts the on-device path first for whichever locale is selected, and only retries against Apple's server if that local pass comes back empty or errors; the transcript stays ordinary editable text afterward, so a handful of misheard names or an unfamiliar spelling convention is a quick correction rather than a reason to start over.
Common questions
Does it matter if I leave transcription set to English (US) when I speak British English?
It will still produce a mostly readable transcript, but accuracy suffers in specific spots — place names, slang and words an American-trained model was not expecting — and the spelling will follow US conventions rather than UK ones. Matching the setting to how you actually speak avoids both.
Will choosing the wrong English variant give me an empty transcript?
Usually not. That failure mode is more typical of picking an entirely different language. A mismatched regional variant of the right language tends to degrade a transcript in specific places rather than break it completely.
Does the spelling in my transcript depend on which English I pick?
Yes, generally. A recognizer follows the spelling conventions of the region it is set to, the same way iOS keyboard dictation does, so UK and US settings can produce different spellings for the same spoken word.
Is this only an issue for English?
No. Spanish, Portuguese and Chinese are also split into regional variants in iOS speech recognition, and the same advice applies to all of them: match the setting to how the speaker actually talks.
Try it in Voice Studio
Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.
Free to download · iPhone and iPad · iOS 16.4 or later