Label Speakers in a Transcript, Get Timestamps
Upload audio or video with more than one speaker, or paste a link to one, and ScreenApp returns a transcript that labels each speaker’s turns with a timestamp. The recording does not need to be uploaded from your device: a link from Dropbox, Google Drive, or a podcast host works too.
A labeled transcript means you can find what a specific person said without listening back to the whole recording.
What you get:
- A transcript with speaker labels and timestamps, in 100+ languages
- click a speaker label to rename it, match it to a team member, or reassign a single segment
- Export as PDF, DOCX, TXT, SRT and VTT
- Ask AI questions about the recording and get answers that point to the moment it happened
- Free plan: 2 transcriptions of recordings up to 45 minutes each, with speaker labels included
The system assigns speakers by voice, not by matching faces or names, so every recording starts with generic labels (Speaker 1, Speaker 2) that you rename afterward.
How to Identify Speakers in Audio
- Upload or paste a link: Drag in a file (MP3, M4A, MP4A, M4B, AAC, WAV, OGG, OPUS, FLAC, AIFF, WMA, WEBMA, MKA, AC3, EAC3, WV, AMR, DSF, DFF) or paste a link from Dropbox, Google Drive, or a podcast host.
- The transcript separates by speaker: Each voice gets its own label, with a timestamp on every turn.
- Rename and export: click a speaker label to rename it, match it to a team member, or reassign a single segment. Export the finished transcript as PDF, DOCX, TXT, SRT and VTT.
Clear audio with distinct voices separates best, and overlapping speech can confuse the diarization system, so keep speakers from talking over each other when you can. See how we measure accuracy.
Speaker Diarization vs Other Apps
| Feature | ScreenApp | AssemblyAI | Otter.ai | Notta |
|---|---|---|---|---|
| Free plan | 2 transcriptions, up to 45 min each | $50 in credits | 300 min/month, 30 min/conversation | 120 min/month, 3 min/conversation |
| Speaker labels | Included | +$0.02/hr add-on | Included from Basic plan | Included |
| Languages | 100+ | Not stated | 6 | 58 |
| Export formats | PDF, DOCX, TXT, SRT and VTT | Not stated | mp3 and txt on the free plan, plus pdf, docx and srt from Pro, with bulk export from Business | TXT, DOCX, PDF, or XLSX |
| Paid plan | $19/month annual | $0.15/hour | $8.33/month annual | $8.17/month annual |
Sources, checked 2026-09-24: screenapp.io/accuracy#languages, ScreenApp pricing, assemblyai.com/pricing, otter.ai/pricing, notta.ai/en/pricing, notta.ai/en, notta.ai/en/lecture-summarizer
- vs AssemblyAI: AssemblyAI is a developer API, priced at $0.15 an hour of audio with speaker diarization as a +$0.02/hr add-on on top, and $50 in free credits for a new account. ScreenApp is a finished app: upload a file and get the labeled transcript back, no integration required.
- vs Otter.ai: Otter’s free plan is 300 minutes a month, capped at 30 minutes per conversation, with speaker identification from the Basic plan up. ScreenApp’s free plan is 2 transcriptions of recordings up to 45 minutes each, with speaker labels included from the start.
- vs Notta: Notta’s free plan gives 120 minutes a month in clips of up to 3 minutes, with speaker identification included. Notta transcribes 58 languages; ScreenApp transcribes 100+.
Who Uses Speaker Diarization
Podcasters upload the raw episode and get a transcript split by host and guest, ready to paste into show notes or a transcript page for the episode.
Meeting and interview notes benefit when the audio alone still shows who said what. Interviewers use the speaker split to separate their own questions from the answers without re-listening.
Researchers running focus groups or qualitative interviews use consistent speaker labels to track who contributed what, instead of labeling turns by hand.
Legal and healthcare professionals need speaker-labeled transcripts of depositions, client calls, and consultations, with a timestamp attached to each turn.


