Speaker Labels & Diarization

How speaker detection works and what affects accuracy.

Local Transcriber uses on-device speaker diarization to label who said what in your transcripts. No cloud processing is involved — the model runs entirely on your Mac.

How It Works

During a recording, the app captures audio and runs two models:

  1. Speech recognition — converts audio to text in real time using Apple’s on-device speech engine.
  2. Speaker diarization — analyzes voice characteristics to assign speaker labels (Speaker 1, Speaker 2, etc.).

Live transcription shows text as it’s spoken. After you stop recording, post-processing re-analyzes the full audio for more accurate speaker assignment and improved text.

What Affects Accuracy

Speaker diarization works best when:

  • Speakers take turns. The model struggles with heavy crosstalk or people talking over each other.
  • Audio quality is good. A quiet room with a decent microphone helps. Built-in laptop mics work, but external mics or headsets are better.
  • Speakers have distinct voices. The model uses voice characteristics (pitch, cadence, tone) to distinguish speakers. Very similar voices may be grouped together.
  • Sessions aren’t extremely long. Accuracy is consistent for most meeting lengths, but very long recordings (3+ hours) may see some drift.

Tips for Better Results

  • Use headphones or a headset. This prevents speaker audio from bleeding into your microphone, which helps the diarization model separate voices more cleanly.
  • Mute when not speaking. If you’re in a noisy environment, muting your mic between turns reduces background noise in the recording.
  • Let people finish. Overlapping speech is the hardest case for diarization. Natural turn-taking produces the best results.

Post-Processing

When you stop a recording, the app runs post-processing automatically. This takes a few seconds (roughly 15 seconds for a 30-minute meeting) and:

  • Re-runs the full audio through a higher-accuracy transcription model.
  • Re-analyzes speaker assignments across the entire recording for consistency.
  • Generates an AI title and summary (if Apple Intelligence is available).

You’ll see a “Processing” indicator in the sidebar while this runs. The transcript updates in place when it’s done.

Reprocessing

If speaker labels don’t look right, you can reprocess a recording:

  • From the app — right-click a recording and select Reprocess, or use the URL scheme:
    open "transcriber://reprocess?file=2026-04-03-1400"
  • From the CLI — your AI agent can trigger reprocessing too.

Reprocessing re-runs the full post-processing pipeline on the original audio file.

Try it yourself.

Download Local Transcriber, join a call, and see the transcript appear in real time. No account needed.

Download Now

14-day free trial · $20 one-time · macOS Sonoma 14.2+ · Apple Silicon native