Why does my iPhone voice memo transcript include "um" and "uh"?
Short answer: because you said them, and the recognizer's job is to write down sounds, not to decide which ones you meant to say. Whether a given "um" makes it into the text is closer to a coin flip than a rule — here is why, and what actually helps once the transcript is in front of you.
A transcript that reads "so, um, I think the, uh, the deadline is Friday" looks like the recognizer tripping over itself. It is not. "Um" and "uh" are real, recognizable sounds with their own acoustic shape, and a speech model that is good at telling words apart is, by the same mechanism, reasonably good at catching them too. The surprise is not that they show up — it is that they do not show up every time.
A recognizer does not know what a filler word is
Speech recognition maps sound to text. It is not told, as a rule, "these particular sounds are disfluencies, drop them" — that would require deciding what you meant to say versus what you actually said, which is a judgment about intent, not a transcription task. Left alone, the honest output is everything that was voiced, hesitations included, for the same reason a transcript keeps a repeated word or a sentence that trails off and restarts: the recognizer's job stops at "what sound was that," not "was that sound worth keeping."
Why it is inconsistent rather than all-or-nothing
A few things push a given "um" or "uh" one way or the other, and none of them are something you control from a settings screen:
- How clearly it was voiced. A short, clipped "uh" buried between two words can blend into the surrounding sound and never get resolved as its own unit, the same way a mumbled real word can vanish.
- What came immediately around it. A filler word right before a pause tends to survive; one swallowed mid-phrase, with no gap on either side, is more likely to get merged into whatever it is next to or dropped.
- How it was said, not just what. A drawn-out, emphasized "ummm" behaves differently from a quick "um" clipped almost to nothing — acoustically they are not the same event, even though a person hears both as filler.
None of that is a bug to report. It is the same pattern-matching-under-uncertainty behind every other word in the transcript, applied to sounds that happen to be hesitations instead of content.
Does on-device vs. the server fallback change this?
iOS attempts on-device speech recognition first and only falls back to Apple's server if that local pass comes back empty or errors. Both are guessing at the same acoustic evidence with the same lack of instruction to filter hesitations, so switching which one ran on a given recording is not a reliable way to get more or fewer filler words in the output — the inconsistency lives in the audio and the model, not in which of the two ran.
When it is worth leaving them in
Before editing every "um" out, it is worth asking whether this transcript needs to. A verbatim record of a deposition, an interview, or a difficult conversation is sometimes more useful with the hesitations left in — a pause and a stumble before someone answers a question can be part of what the transcript is for. And if the recording exists to help you hear your own speaking habits — rehearsing a talk, practicing an answer — the filler words are the whole point of listening back; scrubbing them from the text before you have noticed the pattern in the audio defeats the purpose.
Cleaning it up when you do not need them
For most everyday use — turning a voice memo into notes, a draft, a message — filler words are just clutter to remove, and the fastest way through is a single read-through rather than hunting each one individually: read the transcript at normal pace and delete "um," "uh," and the odd repeated word as your eye hits them, the same pass you would do for punctuation or a misheard name. It goes quickly because you are not correcting content, you are trimming sounds that were never meant to carry meaning.
How Voice Studio handles this
Transcription runs through Apple's speech engine, on-device first, and Voice Studio does not run a separate pass to detect or strip filler words — whatever the recognizer wrote down is what shows up in the transcript. What it does give you is a transcript that stays editable afterward, for free: delete an "um," fix a misheard word, and the correction sticks everywhere that text is used — the in-app search, and a TXT or JSON export carries the cleaned-up version, not the original.
Common questions
Can I turn off filler words in Voice Studio's transcripts automatically?
No — there is no setting that strips "um" and "uh" before they reach the transcript. The text reflects whatever Apple's recognizer produced, and removing filler words is a manual edit, the same as fixing any other word.
Why does the transcript catch "um" the first time someone says it but miss it later in the same recording?
Because each occurrence is recognized independently from the surrounding sound, not tracked as a pattern. How clearly a particular "um" was voiced, and what pause sat around it, decides whether that instance gets written down — not whether earlier ones were.
Does switching to server-based transcription remove filler words?
No. iOS tries on-device recognition first and only falls back to the server if that pass is empty or errors, and both are working from the same acoustic evidence without any instruction to filter hesitations, so which one ran does not reliably change how many filler words show up.
Should I edit filler words out of every transcript?
Only if the transcript is meant to read cleanly. A verbatim record — an interview, a deposition, a rehearsal you are reviewing for your own pacing — can be more useful with them left in, since the hesitations themselves are sometimes the information you wanted.
Try it in Voice Studio
Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.
Free to download · iPhone and iPad · iOS 16.4 or later