Speaker Diarization Online

Upload a recording with more than one speaker and get a transcript that labels who's talking, with timestamps you can search and export.

Upload an audio file to get a transcript.

MP3, WAV, M4A, AAC, OGG, FLAC, AMR, WMA

Or paste an audio link

YouTube, TikTok, Instagram, Vimeo, Google Drive, Dropbox, direct links

Recordings are encrypted in transit and at rest and are not used to train AI models. Privacy · Security

Label Speakers in a Transcript, Get Timestamps

Upload audio or video with more than one speaker, or paste a link to one, and ScreenApp returns a transcript that labels each speaker’s turns with a timestamp. The recording does not need to be uploaded from your device: a link from Dropbox, Google Drive, or a podcast host works too.

A labeled transcript means you can find what a specific person said without listening back to the whole recording.

What you get:

The system assigns speakers by voice, not by matching faces or names, so every recording starts with generic labels (Speaker 1, Speaker 2) that you rename afterward.

How to Identify Speakers in Audio

  1. Upload or paste a link: Drag in a file (MP3, M4A, MP4A, M4B, AAC, WAV, OGG, OPUS, FLAC, AIFF, WMA, WEBMA, MKA, AC3, EAC3, WV, AMR, DSF, DFF) or paste a link from Dropbox, Google Drive, or a podcast host.
  2. The transcript separates by speaker: Each voice gets its own label, with a timestamp on every turn.
  3. Rename and export: click a speaker label to rename it, match it to a team member, or reassign a single segment. Export the finished transcript as PDF, DOCX, TXT, SRT and VTT.

Clear audio with distinct voices separates best, and overlapping speech can confuse the diarization system, so keep speakers from talking over each other when you can. See how we measure accuracy.

Speaker Diarization vs Other Apps

FeatureScreenAppAssemblyAIOtter.aiNotta
Free plan2 transcriptions, up to 45 min each$50 in credits300 min/month, 30 min/conversation120 min/month, 3 min/conversation
Speaker labelsIncluded+$0.02/hr add-onIncluded from Basic planIncluded
Languages100+Not stated658
Export formatsPDF, DOCX, TXT, SRT and VTTNot statedmp3 and txt on the free plan, plus pdf, docx and srt from Pro, with bulk export from BusinessTXT, DOCX, PDF, or XLSX
Paid plan$19/month annual$0.15/hour$8.33/month annual$8.17/month annual

Sources, checked 2026-09-24: screenapp.io/accuracy#languages, ScreenApp pricing, assemblyai.com/pricing, otter.ai/pricing, notta.ai/en/pricing, notta.ai/en, notta.ai/en/lecture-summarizer

  • vs AssemblyAI: AssemblyAI is a developer API, priced at $0.15 an hour of audio with speaker diarization as a +$0.02/hr add-on on top, and $50 in free credits for a new account. ScreenApp is a finished app: upload a file and get the labeled transcript back, no integration required.
  • vs Otter.ai: Otter’s free plan is 300 minutes a month, capped at 30 minutes per conversation, with speaker identification from the Basic plan up. ScreenApp’s free plan is 2 transcriptions of recordings up to 45 minutes each, with speaker labels included from the start.
  • vs Notta: Notta’s free plan gives 120 minutes a month in clips of up to 3 minutes, with speaker identification included. Notta transcribes 58 languages; ScreenApp transcribes 100+.

Who Uses Speaker Diarization

Podcasters upload the raw episode and get a transcript split by host and guest, ready to paste into show notes or a transcript page for the episode.

Meeting and interview notes benefit when the audio alone still shows who said what. Interviewers use the speaker split to separate their own questions from the answers without re-listening.

Researchers running focus groups or qualitative interviews use consistent speaker labels to track who contributed what, instead of labeling turns by hand.

Legal and healthcare professionals need speaker-labeled transcripts of depositions, client calls, and consultations, with a timestamp attached to each turn.

FAQ

What is speaker diarization?

Speaker diarization is the process of working out who spoke when in a recording. The system groups the audio by voice and assigns each segment a speaker label, so the transcript reads as Speaker 1, Speaker 2, and so on, each turn timestamped.

How accurate is speaker diarization?

Accuracy depends on the recording: clear audio with distinct voices separates best, and overlapping speech can confuse the diarization system. See how accuracy is measured.

How many speakers can it identify?

Diarization assigns a label to every distinct voice it detects in the recording, from two speakers up to a full group conversation. More speakers, especially with similar voices or overlapping speech, make the separation harder.

Does it work for podcasts?

Yes. Upload the episode file or paste a link from your podcast host, and each host and guest gets a separate speaker label that you can rename in the transcript.

Does it do real-time diarization?

No, this is an upload tool: it processes a file or link after the recording is finished, not while it happens. For live meetings, use the meeting recorder, which records and transcribes as the call happens.

Can it identify specific people by name?

Not automatically. The system assigns generic labels (Speaker 1, Speaker 2) from voice alone; it does not match a voice to a known identity. Afterward, click a speaker label to rename it, match it to a team member, or reassign a single segment.

Is there a free tier?

Yes. 2 transcriptions of recordings up to 45 minutes each, with speaker labels, timestamps, and export included. The Pro plan at $19 a month billed annually removes the recording cap.

Real Results from Real Users

works like a charm, notes have been super helpful alongside chat function. loveeeee
KS
Katharine Suy
Chrome Web Store, October 17, 2024
nice app for important summary
S
SAHIL
Google Play, January 30, 2026

Ready to boost your productivity?

Start Speaker Diarization with a free account. Usage limits apply; app downloads require a paid plan.

Start Free →

Start using in 60 seconds