Transcribe a Recording with Multiple Speakers, Get Each One Labeled
Upload a recording with several people talking, or paste a link, and ScreenApp turns it into a transcript with each speaker labeled and a timestamp on every line. Record the conversation live with the browser recorder or the iOS and Android apps, or connect the meeting bot for Zoom, Google Meet and Microsoft Teams for a call, then upload the file when you have one.
A labeled transcript replaces re-listening to a recording to work out who said what.
What you get:
- A transcript with speaker identification, in 100+ languages
- AI chat with any recording, so you can ask what a specific speaker said instead of scanning the whole transcript
- AI notes: action items with an owner and deadline, and decisions listed separately
- Export as PDF, DOCX, TXT, SRT and VTT
- Free plan: 2 transcriptions of recordings up to 45 minutes each, with AI chat on both
How to Transcribe Multiple Speakers
- Upload or paste a link: Drag in the audio or video file (MP4, M4V, MOV, AVI, WEBM, MKV, FLV, TS, MTS, M2TS, 3GP, 3GPP, 3G2, WMV, ASF, VOB, OGV, RM, RMVB, MPG, MPEG, M2V, F4V, MXF), or paste a link. No account setup is needed to start.
- Let it transcribe: The audio is converted to text and each person is split into their own speaker label, with a timestamp on every sentence.
- Review and export: Read the transcript in the browser, rename a speaker label if needed, and export as PDF, DOCX, TXT, SRT and VTT.
Accuracy depends on the audio and the language spoken. 25 of the 100+ supported languages carry word-level speaker diarization; the rest transcribe to text without per-speaker labels. A single microphone with participants close by and limited cross-talk also gives the clearest split between speakers. See how we measure accuracy.
Multi-Speaker Transcription vs Other Apps
| Feature | ScreenApp | Otter.ai | Rev | Sonix |
|---|---|---|---|---|
| Free plan | 2 transcriptions, up to 45 min each | 300 min/month | 45 min/month (AI) | 30-min trial |
| Speaker identification | Included | Basic | Included | Included |
| Speaker rename | click a speaker label to rename it, match it to a team member, or reassign a single segment | Not stated | Not stated | Not stated |
| AI notes | Included | Pro, $8.33/month billed annually, or $16.99 monthly | Not stated | Not stated |
| Export formats | PDF, DOCX, TXT, SRT and VTT | mp3 and txt on the free plan, plus pdf, docx and srt from Pro, with bulk export from Business | Not stated | DOCX, PDF, TXT, SRT, VTT |
| Languages | 100+ | 6 | Not stated | 54+ |
| Paid plan | $19/month annual | $8.33/month annual | $25.49/month annual | $25/month |
Sources, checked 2026-10-01: screenapp.io/accuracy#languages, ScreenApp pricing, otter.ai/pricing, rev.com/pricing, sonix.ai/pricing, Editing speaker labels
- vs Otter.ai: Otter’s free plan caps you at 300 minutes a month, and its speaker identification tier on the free plan is Basic. ScreenApp includes speaker labels on every plan and lets you rename a label after the fact.
- vs Rev: Rev’s free tier gives 45 AI minutes a month, then charges $25.49/month (billed annually) per minute for an AI transcript or $1.99 per minute for a human one. A recurring series of multi-speaker recordings adds up fast at either rate.
- vs Sonix: Sonix gives a one-time 30-minute trial, then a $25-a-month plan. ScreenApp’s free plan renews rather than being a one-time trial.
Who Uses Multi-Speaker Transcription
Journalists transcribe interviews and panel discussions with several sources, then quote each speaker directly with a timestamp for fact-checking.
Podcast producers transcribe episodes with a host and multiple guests, splitting the conversation by speaker to pull show notes and clips.
Researchers transcribe group interviews and focus groups, then read the transcript by speaker to see how each participant answered.
Recruiting and HR teams transcribe panel interviews with several interviewers and a candidate, so anyone reviewing the file later can see who asked what.
Legal and compliance teams transcribe depositions, multi-party calls, and recorded statements where knowing which person said which line matters.
Documentary and video teams transcribe interview footage with more than one subject, using the speaker labels to log who is on camera at each point.


