Voice Studio

How to combine several voice memo transcripts into one document

A transcript app transcribes one recording at a time, because that is the unit it actually has — the audio it captured in one continuous take. When the thing you actually want is a week of journal entries, or both halves of an interview, read as one piece, that is text and text combines. It just does not combine itself.

This comes up more than the "merge" framing suggests. It is rarely one recording that got mysteriously split in two. More often it is several recordings that were always separate — a second interview session on another day, a lecture series where each class is its own file, a project you have been talking notes about for a week — and you only want them read together after the fact, as one document rather than a shelf of individually-titled entries.

When this is actually several recordings, not a broken one

Worth being precise about, because it changes what you are looking for. A single continuous recording that is interrupted by a phone call does not need combining — the app pauses on the interruption and resumes into the same file, so you get one recording with a gap, not two. Two separate files for what felt like one session usually means one of a few specific things: the phone ran out of storage or battery mid-way and the rest had to be a fresh recording once that was sorted, or a call interruption landed at a moment recording genuinely could not resume from. Those are worth reading about on their own terms if either just happened to you. What this page is about is the far more common case: recordings that were never one take to begin with, and that you now want assembled into a single document anyway.

There is no batch-merge feature, and that is fine

Nothing in a voice recorder app watches for related recordings and offers to stitch their transcripts together — that would require guessing which recordings belong together, which is a judgment only you can make. What every recording does have is a transcript you can get out as plain text, and assembling plain text into one document is a five-minute job with tools you already have, not a missing feature.

Fix each transcript before you combine them, not after

Do this first, because it is far more tedious once several transcripts are already pasted into one document. Speech recognition has no memory between recordings — a name it mis-heard in Monday's session gets guessed at fresh in Thursday's, and there is no reason to expect it lands on the same wrong spelling twice. Read each transcript on its own and fix what is wrong while it is still one file and you still remember what was actually said; catching an inconsistent spelling of the same name across three separate transcripts, after they are all one document, is a much harder read.

Get the order right before you paste anything

Chronological is the right default for almost everything this applies to — an interview across two sessions, a week of journal entries, a lecture series. Sort the library by date to see the actual order things happened in, rather than relying on memory or on whatever order the recordings happen to sit in on screen; the order you recorded in and the order you happen to be looking at them in are not always the same thing.

Exporting each one

For a document meant to be read, export as TXT rather than JSON. The TXT export already puts the title, date, time and location above the transcript as plain lines — that header is doing useful work here even outside its original purpose, because once several transcripts are pasted into one document, it is what tells you where one recording ends and the next begins. Keep it rather than stripping it out for the sake of a cleaner-looking page.

For two or three recordings, exporting one at a time through the share sheet and pasting each into a notes or document app as it comes is the whole job. For a longer series — a full lecture course, a month of journal entries — pulling everything out through the Files app at once is faster than repeating the share sheet for each recording individually, since it hands you every exported file in one place rather than one at a time.

Assembling the document

Paste each transcript in, in the order you settled on, and leave its TXT header in place as a section break — you now have a single document where "Tuesday's session" and "Thursday's session" are both visible and both attributable, rather than one undifferentiated block of text that used to be two separate recordings. A blank line or a rule between sections is enough; nothing more elaborate is needed for a document whose whole purpose is being read straight through.

When JSON is the better source instead

Pasting by hand is the right tool for a handful of recordings read by a person. If you are instead building something that has to process a large number of transcripts on a schedule — importing them into a spreadsheet, feeding a script that assembles the document for you — JSON is the format built for that. It keeps the date, duration and transcript as separate fields a program can sort and stitch together reliably, rather than text a person has to read and place by eye. Which one is right depends entirely on whether a person or a program is doing the assembling, not on which format is "better".

What does not carry over into a combined document

A marker dropped during one of the original recordings is a line on that recording's own waveform, meant to jump playback back to a moment in the audio — it has no equivalent in plain text, so it does not appear in the export and cannot appear in a document built from several exports pasted together. If a moment marked during recording matters enough to keep in the combined document, that is a manual note you add yourself while assembling it, not something that survives the export automatically.

This is not the same job as a full backup

A full backup is one JSON file covering every recording on the phone, built so the whole library can be restored, not so a handful of related recordings can be read as one document. It has no sense of which recordings belong together and it is not meant to be opened and read the way a combined document is — it is insurance for the day the phone is lost, not a shortcut for this. Building the document you actually want to read still means picking the relevant recordings out and exporting them individually, whether or not you also keep a full backup running.

How Voice Studio fits into this

There is no built-in way to select several recordings and export one combined document — each recording exports its transcript as TXT or JSON on its own, or its audio as the original .m4a, and the library can be sorted by newest, by length or by word count to get the recordings in the order you need before you start. The transcript stays editable right up until you export it, so fixing a name in Monday's session before you paste it alongside Thursday's costs a moment there and saves a much longer read later. Nothing about any of this needs Pro — export, and the editing that makes it worth exporting, are both in the free app.

Common questions

Is there a way to select multiple recordings and export one combined file?

No. Each recording exports its transcript or audio on its own — as TXT, JSON or the original .m4a. Combining several into one document means exporting each and assembling them yourself, which for a handful of recordings is a short job.

What order should the transcripts go in?

Chronological, in almost every case this comes up — an interview across sessions, a lecture series, a run of journal entries. Sort the library by date rather than relying on the order the recordings happen to be sitting in on screen.

Will a marker I tapped during recording show up in the combined document?

No. A marker is a position on that recording's own waveform, used for jumping playback back to a moment — it has no text equivalent and does not appear in a TXT or JSON export. Note the moment yourself if it needs to survive into the combined document.

Should I use the full backup instead of exporting each recording separately?

No. The full backup is one JSON covering the entire library, meant for restoring everything if the phone is lost, not for assembling a handful of related recordings into a readable document. Export the specific recordings you need individually instead.

Try it in Voice Studio

Voice Studio records, transcribes on your iPhone, and files each note by time and place — so the thought you had in the car is still findable next month.

Download on the App Store

Free to download · iPhone and iPad · iOS 16.4 or later