Voice Studio

Standard vs high vs ultra recording quality on iPhone

The setting exists, it has three options, and nothing in the app tells you what changes between them. Here is what each one actually does, what it costs, and a rule of thumb that covers almost every situation you will actually be in.

A quality picker with three unlabelled tiers invites the same reflex as any other settings screen: pick the top one, on the theory that more must be safer. For a voice recording that reflex mostly wastes storage rather than protecting anything, and it is worth knowing exactly why before you default to it out of caution.

What actually changes between the three

On Voice Studio, the three levels — standard, high and ultra — set the sample rate: how many times per second the microphone signal is measured. Standard runs at 22.05 kHz, high at 44.1 kHz, and ultra at 48 kHz. All three record mono AAC into an .m4a file. Nothing else moves between them — same codec, same channel count, same file type.

A higher sample rate captures a wider slice of the audible spectrum and, because there is more measured signal to encode, a somewhat larger file for the same length of recording. That is the entire trade: some amount of extra frequency range and file size, against nothing you asked to give up.

Why "wider frequency range" barely matters for a voice

A human voice, including everything that makes speech intelligible, sits mostly below about 8 kHz. To capture a frequency cleanly a recorder needs to sample at roughly double it, so even the standard setting, at 22.05 kHz, sits comfortably above what a spoken voice needs. High and ultra capture more of a room's upper-frequency detail — a chair creak, the hiss of an air vent, the top end of a musical instrument — not more of the voice itself, because the voice was already fully there at the lowest of the three settings.

This is not a claim that the tiers sound identical on everything. It is specifically a claim about spoken-word audio, which is what a phone microphone in someone's hand or on a desk is recording the overwhelming majority of the time.

What actually degrades a recording, and what does not

None of the three settings touches the things that genuinely make a voice recording hard to listen to or hard to transcribe afterwards: how far the phone sits from whoever is speaking, how much the room echoes, background noise from a fan or traffic, and audio clipped from speaking too close or too loud. All of that reaches the recording the same way regardless of which tier is selected. Moving the phone six inches closer to a speaker does more for a recording than jumping from standard to ultra ever will.

The one place a higher tier earns its size

There is a genuine, narrow reason to reach for ultra: a recording that is heading into a video project running at 48 kHz. Matching the audio's native sample rate to the video timeline's avoids a resampling step later, which is a convenience for that specific pipeline rather than anything you would hear as a quality difference on its own. Outside of that case, ultra is mainly for someone who wants the highest-fidelity file the app can produce as a matter of principle, independent of whether the difference is audible.

What it costs to go up a tier

A higher sample rate means more measured signal reaching the encoder, and that shows up as a larger file for the same length of recording — the exact difference depends on the audio and the encoder, but the direction is consistent: standard produces the smallest files of the three, ultra the largest. Across a long recording, or a habit of recording often, that difference is the one that actually shows up later, as free space on the phone rather than as anything you notice on playback.

A simple way to choose

There is no wrong choice here in the sense of a ruined recording — all three produce a clean, usable .m4a. The only real cost of picking the wrong one is storage spent on frequency range a voice was never going to use.

How Voice Studio handles it

The quality setting is a free choice, not something held back for Pro — standard, high and ultra are all available from the start, and you can change the setting between recordings whenever you like. High is the default, which matches the recommendation above for most spoken-word recording. Whichever tier you pick, transcription behaves the same way: it attempts the on-device pass first regardless of the sample rate the recording was made at, since the quality setting is not part of what a transcript is built from.

Common questions

What is the actual difference between standard, high and ultra quality?

The sample rate: 22.05 kHz for standard, 44.1 kHz for high, and 48 kHz for ultra. All three record mono AAC into the same .m4a file type — nothing else changes.

Does a higher quality setting make a voice recording sound clearer?

Not for speech, in any way you would notice. Speech intelligibility lives in a frequency range that even the standard setting already covers. What genuinely affects clarity is distance from the speaker, room noise and echo — none of which a quality setting changes.

Which setting should I use by default?

High is a sensible default for almost any spoken-word recording. Standard is a reasonable choice if storage is tight, since it does not cost you anything speech needs. Ultra is mainly worth it if the audio is going into a 48 kHz video project.

Does the quality setting affect how big the file is?

Yes — a higher sample rate means more measured signal for the encoder to store, so the file grows somewhat from standard to high to ultra for the same length of recording. It is the main real trade-off between the three.

Try it in Voice Studio

Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.

Download on the App Store

Free to download · iPhone and iPad · iOS 16.4 or later