Voice Studio

How to make a voice memo accessible to someone who is deaf or hard of hearing

The instinct is to share the recording, the same way you would with anyone else. For someone who cannot hear it, that hands over nothing at all. The part that actually reaches them is the transcript, and getting that right matters more here than it does when a hearing recipient can always play the audio if something reads oddly.

This usually comes up in an ordinary, unplanned moment: a voicemail from a doctor’s office, a voice note from a relative, a recording of a meeting someone missed. Forwarding the audio feels like the obvious, generous thing to do. For a deaf or hard-of-hearing recipient it is not generous, it is just inert — a file that plays sound into a channel they cannot use, no different from sending a photo to someone who is blind and calling it a description.

The recording was never the accessible part

A voice memo is, by construction, audio — a signal meant for ears. That is true no matter how clearly it was recorded, how good the microphone was, or how quiet the room was. None of the things that make a recording sound good change what it fundamentally is. The only version of a voice memo that is accessible to someone who cannot hear it is the one made of words on a page, and that means a transcript is not an optional extra here the way it might be for a hearing recipient who could just as easily listen. It is the entire deliverable.

Which export format to actually send

A transcript export comes in two shapes, and they are built for different readers. A TXT export puts the title, date, time and location above the transcript, then the words themselves as plain, continuous prose below — the same shape as a text message or an email, opening cleanly in Notes, Mail, Messages or any text editor on any kind of device. A JSON export carries the same details as separate labelled fields and adds a few more, like duration, mood and tags, wrapped in the punctuation a program expects to parse rather than a person expects to read.

For handing a transcript to a person, TXT is almost always the right choice. It is what a plain reading of "what did they say" wants, and it asks nothing of the recipient beyond being able to open a text file, which every phone and computer can already do. JSON has a real purpose — moving structured details somewhere else, or feeding a transcript into a program that expects fields — but reading one directly means wading through braces, quotation marks and field names sitting in between you and the actual words.

What a transcript still cannot do

Handing over the words is not the same as handing over synced captions. Neither a TXT nor a JSON transcript export carries per-word or per-line timestamps — the text comes out as continuous prose, not as timed lines that could scroll along with audio playback the way a caption track does. There is no SRT or VTT export either, which is the format that would actually be needed to show text in sync with sound as it plays. What you can hand someone is the full text of what was said, to read on its own; what you cannot hand them is a version of the recording where the words appear in time with it, because that is a different kind of file than a transcript export produces.

In practice this is rarely the gap it sounds like. Most of the situations a voice memo actually gets shared for — a voicemail, a message from a relative, notes from a meeting someone missed — are read once, at the recipient’s own pace, not watched in real time alongside a video. The plain text does the job that mattered. Synced captions matter more for video, where the picture and the timing are doing real work together, and that is a different problem from a voice recording being made readable at all.

Read it before you send it, not after

A transcript is editable text sitting on the recording, not a locked printout, and that matters more for this than for almost any other use of a transcript. A hearing recipient who hits a garbled sentence can just play the audio and hear what was actually said. A deaf or hard-of-hearing recipient has no such fallback — the text is the only version of the recording that reaches them, so a misheard name or a nonsense word is not a minor rough edge, it is the entire message they receive at that point in the recording. Give the transcript one read-through against your own memory of the conversation before sending it, and fix anything that looks like a guess rather than a word, the same quick pass worth doing before sharing a transcript with anyone, but here it is the difference between them getting the message and getting a gap.

When it is more than one recording

A single voicemail is a one-off export. A recurring situation — a family member who is deaf and wants the text of every voice note the family group chat sends, or a student who needs a transcript of every recorded lecture — is really the same job done repeatedly: export as TXT, read it once, send it. There is no batch step that changes this, and there does not need to be one; a transcript costs a handful of kilobytes and a minute of reading regardless of how many times you do it, which is a far smaller ask than it sounds the first time.

The audio is still worth keeping

None of this means the .m4a should be thrown away or left out. Sending both the transcript and the original recording costs nothing extra and covers the cases text alone does not: a hearing person in the same conversation who wants the audio, a recipient who lipreads and wants to watch a video call recording rather than only read it, or simply wanting the original kept somewhere in case a transcript ever needs a second look. The transcript is what makes the recording accessible to someone who cannot hear it; the audio is not competing with that, it is just not the part doing that particular job.

How Voice Studio handles this

Exporting a transcript as TXT or JSON, or sharing a recording as its .m4a file, are ordinary, unrestricted parts of the free app — none of it sits behind Pro. A transcript inside Voice Studio is editable text on the recording, so correcting a misheard name or a garbled phrase before sending is a normal edit, not a workaround. There is no SRT or VTT export and no per-word timing in either transcript format, so what you send is the full text of a recording, not captions synced to it — worth knowing up front rather than discovering after the fact.

Common questions

Does sharing the audio file itself help someone who is deaf or hard of hearing?

No. The recording is sound, and sharing it hands over nothing they can use on its own. The transcript is what actually reaches them — the audio matters for other reasons, but not this one.

Should I send the transcript as TXT or JSON?

TXT, for handing words to a person. It opens as plain, readable prose on any device. JSON carries the same details as separate structured fields plus a few more, like duration and tags, which suits moving data into another program more than it suits being read directly.

Can I make the text scroll along with the audio like captions?

No. Neither transcript export format carries per-word or per-line timestamps, and there is no SRT or VTT export, so there is no way to sync text to playback. What you can send is the full transcript to read on its own.

Why does it matter more to proofread this transcript than a normal one?

Because the recipient has no audio to fall back on if a word looks wrong. A hearing person can just replay the recording; for someone who cannot hear it, the text is the entire message, so a misheard word is a real gap rather than a minor slip.

Try it in Voice Studio

Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.

Download on the App Store

Free to download · iPhone and iPad · iOS 16.4 or later