Speaker diarization

How to Transcribe Audio With Speaker Labels (iPhone & Android)

Updated Jul 5, 2026·8 min read

On this page

You recorded a conversation — a meeting, an interview, a call — and the auto-transcript came back as one long, speakerless paragraph. Unreadable. What you want is a transcript that shows who said what, with each line labeled by speaker. This guide shows you exactly how to transcribe audio with speaker labels on iPhone and Android, how to make the labeling accurate, and how to do it all privately on your device.

What you'll need

Getting a labeled transcript takes two capabilities working together:

  • Transcription — turning speech into text (the words).
  • Speaker diarization — separating the voices and tagging each line (the turns).

Most built-in phone transcription does the first but not the second, which is why you get a wall of text. To get labels, you need a tool that does both. For the concepts behind this, see what is speaker diarization and transcription with speaker labels.

Step by step

Here's the full workflow, start to finish.

1. Record the conversation cleanly

Label accuracy is decided here, before any software runs. The three rules:

  • Mic close to the speakers. A phone on the table between two people works well; across the room does not. See recording two people or a room.
  • Quiet room. Background noise muddies the voices and confuses labeling — how to reduce background noise.
  • One voice at a time. Overlapping speech is the hardest case for diarization; a little social nudging helps a lot.

An always-on recorder is ideal — start it before the conversation and it captures everything in the background so you're not fiddling with equipment.

2. Transcribe the recording

Convert the audio to text with word-level timestamps (those timestamps are what let speaker labels attach to the right words). The key decision here is where transcription happens — on a server (upload) or on your device. For conversations, on-device is the private choice; more on that below.

3. Diarize — attach the speaker labels

The tool separates the voices and tags each utterance: Speaker 1, Speaker 2, and so on. Good tools do this in the same step as transcription, so you get a finished labeled transcript in one action rather than stitching two tools together.

4. Rename labels to real people

Diarization gives you anonymous labels. Rename "Speaker 1" to "Alex," "Speaker 2" to "Priya," and the transcript reads naturally. You supply the names — the app never needs to recognize anyone, which keeps it private (that's the difference between diarization and speaker recognition).

5. Use the transcript

Now it's genuinely useful: read it, search it by person, quote the right speaker, or hand it to an AI for a summary and action items.

iPhone vs Android: the same problem

Both platforms have built-in transcription that mostly falls short for this task:

  • iPhone — Voice Memos (iOS 18+) can transcribe on newer devices, but it generally gives you one block of text with no speaker labels, and it's limited by hardware and language.
  • Android — built-in options vary by manufacturer and are inconsistent, and speaker labeling is rarely included.

So on either phone, you need a dedicated tool for labeled transcripts. The advantage of an app that ships its own transcription and diarization is that the workflow is identical on iPhone and Android — you're not dependent on what your particular handset supports. BlackBox works the same on both.

The private method: on-device, no upload

Here's the part that matters for anything sensitive. The common way to get labeled transcripts is to upload your audio to a cloud service (Otter, Fireflies, Rev) or a developer API. Your recording leaves your device. But labeled transcripts are most useful on conversations — meetings, interviews, personal calls — your most private audio.

BlackBox keeps the whole thing on your phone:

  • Record with an always-on recorder.
  • Transcribe on-device — no upload, works offline (on-device transcription).
  • Label speakers on-device — "who said what," computed locally.
  • Everything stays put — no account, behind Face ID.

You get labeled transcripts with none of the cloud exposure — and, unlike the open-source route, no command line, GPU, or developer account required. For a broader tool comparison, see the best speaker diarization apps.

Making labels more accurate

If your transcript mislabels speakers, the fix is almost always the audio:

  • Overlap — the top cause of errors. Re-record or edit if two people constantly talk over each other.
  • Similar voices — two alike voices can merge into one label; nothing beats close, clean audio here.
  • Too many speakers — accuracy drops past six or seven; for big groups, expect a longer review.
  • Distance and noise — move the mic closer and quiet the room.

Then do a quick review pass: rename labels, and fix any lines assigned to the wrong speaker. It's far faster than transcribing by hand, and it gets you a clean, quotable transcript.

What kind of audio can you label?

Speaker labeling works on essentially any recording with multiple voices, but some sources are easier than others:

  • In-person conversations captured with the phone on the table — a strong case when the mic is central and the room is quiet.
  • Online meetings recorded from your phone with the call on speaker — see recording Zoom, Teams, or Meet. Both sides come through the speaker.
  • Interviews — a two-person setup is the easiest and most accurate (interview speaker labels).
  • Existing recordings you already have — as long as you can get the audio into the app, it can be transcribed and labeled after the fact.

The common thread is capture quality: whatever the source, close and clean audio labels far better than distant or noisy audio. If you're recording specifically to label speakers, control the capture; if you're working with audio you already have, expect a heavier review pass when the recording is rough.

Fitting it into your routine

The nice thing about labeled transcription is how little it asks of you in the moment. You don't type during the conversation, run a bot, or manage a live transcript — you just record, then let the phone do the transcription and labeling afterward, often while it's idle or charging. That "capture now, label later" rhythm means you can stay fully present in the meeting or interview and still walk away with a clean, attributed transcript. Make it a default for any multi-person recording and, over time, you build a searchable, quotable archive of your conversations without any extra effort per session — all of it computed on your own device.

A quick example

Before (transcription only):

okay so I think we push to the 15th does that work yeah but marketing needs the assets by the 8th right can you own that sure

After (transcription + speaker labels):

Alex: Okay, so I think we push to the 15th — does that work? Priya: Yeah, but marketing needs the assets by the 8th. Alex: Right — can you own that? Priya: Sure.

Same audio, transformed from an unreadable blob into a conversation you can actually work with.

Troubleshooting bad speaker labels

If your labeled transcript comes out messy, work through these in order — most problems are fixable and most are about the audio, not the software:

  • Everyone merged into one speaker? The voices are too similar or the audio is too quiet/distant. Re-record with the mic closer, or accept a heavier manual split.
  • One person split into several speakers? Their volume or tone varied a lot (leaning in and out, phone vs speaker). Merge the labels; closer, steadier mic placement prevents it next time.
  • Wrong speaker on scattered lines? Almost always overlap. Those lines sit at moments two people talked at once — jump to the timestamp and reassign.
  • A phantom extra speaker? Background noise or a TV got detected as speech. Record in a quieter room (reduce background noise).
  • Labels drift over a long recording? Very long sessions with many speakers strain clustering; expect a review pass for hour-plus group audio.

Exporting and using the labeled transcript

A labeled transcript is only useful if you can get it out of the app and put it to work. Look for the ability to export the text to a file or folder so you can:

BlackBox can export timestamped transcripts to a folder you choose, so your labeled conversations become real, searchable text files rather than staying trapped in an app — all while the audio and the transcription stay on-device.

A note on real-time vs after-the-fact

Some tools label speakers live during a call; others (including the on-device approach here) transcribe and label the finished recording afterward. After-the-fact diarization is generally more accurate, because the system sees the whole conversation before deciding who's who — and for most uses you want a clean transcript after the meeting, not live labels during it. So the "record now, transcribe later" flow isn't a limitation; it's usually the more accurate choice.

The bottom line

To transcribe audio with speaker labels, record cleanly, then use a tool that does transcription and diarization so each line is tagged by speaker — then rename the labels to real people. Built-in iPhone and Android transcription usually skips the labels, so you need a dedicated tool; and because conversations are sensitive, prefer one that runs on-device. BlackBox records, transcribes, and labels speakers on your phone — iPhone and Android, offline, free, with nothing uploaded.

Frequently asked questions

How do I transcribe audio with speaker labels?

Record the conversation cleanly, then transcribe it with a tool that also does speaker diarization. It converts the speech to text and tags each line with a speaker. With BlackBox you record and transcribe on-device, and speakers are labeled on your phone — so you get a 'who said what' transcript without uploading the audio.

Can iPhone or Android label speakers automatically?

The built-in Voice Memos and phone transcription features generally produce a single block of text without speaker labels. To get labeled transcripts you need a tool with speaker diarization. BlackBox adds on-device speaker labeling on both iPhone and Android, so the workflow is the same on either phone.

How do I make speaker labeling more accurate?

Accuracy is decided at recording time: put the mic close to the speakers, record in a quiet room, and encourage one person to talk at a time. Overlapping voices and background noise are the main causes of mislabeled speakers. A quick review to rename labels and fix the occasional line finishes the job.

Record your day with BlackBox

Always-on, on-device and private. Free on iPhone and Android.

Keep reading