Transcribe a Recording with Multiple Speakers, Get Each One Labeled
Upload a recording with several people talking, or paste a link, and ScreenApp turns it into a transcript with each speaker labeled and a timestamp on every line. Record the conversation live with the browser recorder or the iOS and Android apps, or connect the meeting bot for Zoom, Google Meet and Microsoft Teams for a call, then upload the file when you have one.
A labeled transcript replaces re-listening to a recording to work out who said what.
What you get:
- A transcript with speaker identification, in 99 languages
- AI chat with any recording, so you can ask what a specific speaker said instead of scanning the whole transcript
- AI notes: action items with an owner and deadline, and decisions listed separately
- Export as PDF, DOCX, TXT, SRT and VTT
- Free plan: one recording of up to 45 minutes, with AI chat on it
How to Transcribe Multiple Speakers
- Upload or paste a link: Drag in the audio or video file (MP4, WebM, MOV, AVI, MKV, WMV, and FLV), or paste a link. No account setup is needed to start.
- Let it transcribe: The audio is converted to text and each person is split into their own speaker label, with a timestamp on every sentence.
- Review and export: Read the transcript in the browser, rename a speaker label if needed, and export as PDF, DOCX, TXT, SRT and VTT.
Accuracy depends on the audio and the language spoken. 25 of the 99 supported languages carry word-level speaker diarization; the rest transcribe to text without per-speaker labels. A single microphone with participants close by and limited cross-talk also gives the clearest split between speakers. See how we measure accuracy.
Multi-Speaker Transcription vs Other Apps
| Feature | ScreenApp | Otter.ai | Rev | Sonix |
|---|---|---|---|---|
| Free plan | one recording, up to 45 min | 300 min/month | 45 min/month (AI) | 30-min trial |
| Speaker identification | Included | Basic | Included | Included |
| Speaker rename | click a speaker label to rename it, match it to a team member, or reassign a single segment | Not stated | Not stated | Not stated |
| AI notes | Included | Pro, $8.33 to $16.99/month | Not stated | Not stated |
| Export formats | PDF, DOCX, TXT, SRT and VTT | mp3, txt, pdf, docx, srt (Pro, Business and Enterprise plans, with bulk export) | Not stated | DOCX, PDF, TXT, SRT, VTT |
| Languages | 99 | 6 | Not stated | 54+ |
| Paid plan | $19/month annual | $8.33/month annual | $25.49/month annual | $25/month |
Sources, checked 2026-09-13: screenapp.io/accuracy#languages, github.com/screenappai/screenapp-new/blob/main/apps/web/src/app/api/files/[id]/transcript/export/route.ts, screenapp.io/pricing, screenapp.io/help/how-many-minutes-can-i-record, otter.ai/pricing, rev.com/pricing, sonix.ai/pricing, screenapp.io/help/how-to-edit-speaker-labels
- vs Otter.ai: Otter’s free plan caps you at 300 minutes a month, and its speaker identification tier on the free plan is Basic. ScreenApp includes speaker labels on every plan and lets you rename a label after the fact.
- vs Rev: Rev’s free tier gives 45 AI minutes a month, then charges $0.25 per minute for an AI transcript or $1.99 per minute for a human one. A recurring series of multi-speaker recordings adds up fast at either rate.
- vs Sonix: Sonix gives a one-time 30-minute trial, then a $25-a-month plan. ScreenApp’s free plan renews rather than being a one-time trial.
Who Uses Multi-Speaker Transcription
Journalists transcribe interviews and panel discussions with several sources, then quote each speaker directly with a timestamp for fact-checking.
Podcast producers transcribe episodes with a host and multiple guests, splitting the conversation by speaker to pull show notes and clips.
Researchers transcribe group interviews and focus groups, then read the transcript by speaker to see how each participant answered.
Recruiting and HR teams transcribe panel interviews with several interviewers and a candidate, so anyone reviewing the file later can see who asked what.
Legal and compliance teams transcribe depositions, multi-party calls, and recorded statements where knowing which person said which line matters.
Documentary and video teams transcribe interview footage with more than one subject, using the speaker labels to log who is on camera at each point.
FAQ
What does “speaker identification” mean in a transcript?
It means the transcript marks who is speaking at each moment, so instead of one block of text you get separate, labeled lines for each person.
How does ScreenApp identify speakers?
speaker identification splits the audio into separate voices and labels each one in the transcript, in 99 languages.
Does speaker identification work in every language?
No. Word-level speaker diarization runs on 25 of the 99 supported languages, marked with a dagger on the language list. Recordings in the other languages still transcribe to text, without per-speaker labels.
Can I rename the speaker labels?
Yes. click a speaker label to rename it, match it to a team member, or reassign a single segment.
How does ScreenApp handle people talking over each other?
overlapping speech can confuse the diarization system. A single, central microphone with limited cross-talk gives the clearest split between speakers.
Does ScreenApp reduce background noise before transcribing?
No. ScreenApp does not apply noise reduction. A quieter recording environment and a microphone placed close to the speakers produces a cleaner transcript than a noisy room.
Can I export a multi-speaker transcript?
Yes, as PDF, DOCX, TXT, SRT and VTT, with the speaker labels and timestamps kept in the exported file.
Can I ask questions about what a specific speaker said?
Yes, AI chat with any recording lets you ask about the discussion and get an answer that points to the moment in the recording, including who said it.
Is a multi-speaker recording kept private?
Recordings are encrypted in transit and stored with AES-256, on SOC 2 Type 2 infrastructure, and are not used to train AI models. Security controls are published on the trust center.


