Methodology

Accuracy, speed, and trust signals (with receipts)

This page is the source of truth for every accuracy, speed, and language claim on ScreenApp.io. Numbers come from our internal test corpus, the Groq engineering case study, OpenAI's Whisper benchmarks, and xAI's published Grok Speech-to-Text benchmarks. Provider stack last verified against the codebase: September 2026.

Last updated: 2026-09-16.

The transcription model stack

The repository's default batch order is xAI, then Mistral, then Modulate, then Google, then Groq. Live captions use Groq, then xAI, then Mistral, then Modulate, then Google. Credentials, language fit and deployment overrides can change which provider handles a job. These are repository defaults checked on 2026-09-16, not a record of every production request.

ProviderDefault model or endpointLanguage routing declarationSpeaker labels
xAIGrok STT (/v1/stt)25 languages declared by the adapterSupported by adapter
Mistralvoxtral-mini-latest13 languages declared by the adapterSupported by adapter
Modulatevelma-2-stt-batch57 languages declared by the adapterSupported by adapter
Googlegemini-2.5-flashNo fixed language list in the adapterSupported by adapter
Groqwhisper-large-v3-turbo24 languages in the app proficiency list, not the model totalNot supported by adapter

The linked provider documentation describes each service. Provider language totals are not added together to calculate ScreenApp's supported-language count. In particular, Groq's internal routing list is narrower than Whisper's multilingual model coverage.

When a language is selected, batch routing prefers providers that declare it before degraded fallbacks. Live routing excludes providers marked unsuitable for that language. Automatic detection does not provide the same explicit-language filter.

No accuracy or latency benchmark for this current provider order is published here. See the dated measurements below for their original test conditions. Data handling is described in the Privacy Policy.

Speed: the Groq case study

This section is history, not a current benchmark. In 2025, ScreenApp moved from a self-hosted Whisper deployment on AWS to Groq's inference infrastructure. Groq published the case study below; the numbers are from their engineering team's measurements of that pipeline, which was the default at the time. The current default models and routing are listed above. No speed or latency figures have been published yet for the current xAI-led provider chain, or for live captions.

MetricBefore GroqAfter GroqChange
20-minute transcription job~20 minutes~15 seconds20x faster
Per-minute transcription costbaseline1/15th15x cheaper
Free-to-paid conversionbaseline+30%uplift
Annual recurring revenue (year-over-year)baseline+405%growth attributed to the speed and cost gains

Source: ScreenApp + Groq case study (groq.com).

Historical, for the 2025 Groq Whisper pipeline: a 60-minute meeting completes in roughly 3 minutes end-to-end (transcription, diarization, summary generation). A 2-hour video processes in about 6 minutes. These are end-to-end times that include summarization and chaptering, not just raw transcription. No speed figures have been published yet for the current xAI-led chain.

YouTube caption import (not transcription)

When a YouTube video already publishes a caption track, ScreenApp reads that track instead of transcribing the audio. No speech recognition runs at all, which is why it is fast. The text is written by YouTube, so none of the word error rates on this page describe it. Videos without a caption track go through the AI transcription path above, which is slower and is what our benchmarks measure.

Path Imports (n) Median p90
Download + AI transcription (before, Jul 8 to 9 2026)6141.1 sec159 sec
Caption import (after, Jul 16 to 17 2026)1,77414.7 sec29 sec

Method: both windows measure the same endpoint (import start to transcript indexed) on the same population, on production traffic. We state the gain as about 3x rather than a precise figure: the before-window sample is small (n=61), which puts the true median speedup somewhere between 1.9x and 3.7x at 95% confidence. The mean (3.58x) is inflated by a long tail of large videos, so the median is the honest central figure.

Two limits worth stating plainly. First, caption import currently serves free-tier YouTube imports; paid accounts still take the download and transcription path, which measured a 60.5 second median before and 44.7 seconds after (that difference is sampling noise, not a change in the product). Second, 14.7 seconds reflects the current implementation, which reads the captions already published with the video. Measured: July 8 to 17, 2026.

Accuracy: word error rate benchmarks

Word error rate (WER) counts substitutions + deletions + insertions per 100 reference words. Lower is better. Baseline figures below come from the published benchmarks for each underlying model; the per-condition rows are from our own April 2026 retest on 18 hours of public-domain audio per language across three conditions: studio (single speaker, treated room), conference (multi-speaker, room mic), and field (handheld phone mic, ambient noise).

Headline figure: up to 95.8% transcription accuracy on clear English audio in our April 2026 test, which is 100 minus the 4.2% English (US) Studio WER in the table below. This is measured on the April 2026 provider mix used for that retest, before xAI Grok Speech-to-Text became the default provider. It has not been re-measured on the current default provider.

Published baselines

Per-language WER (April 2026 retest)

Language Locale Studio WER Conference WER Field WER Speakers tested
English (US)en-US4.2%7.8%12.4%4
Spanish (Latin Am.)es-4195.1%9.2%14.6%3
Spanish (Spain)es-ES5.4%9.8%15.1%3
Portuguese (BR)pt-BR5.8%10.1%15.8%3
Portuguese (PT)pt-PT6.4%11.2%17.0%2
Frenchfr-FR5.9%10.4%16.2%3
Germande-DE6.1%10.8%16.5%3
Italianit-IT6.3%11.0%17.1%3
Japaneseja-JP7.8%13.5%19.8%2
Koreanko-KR7.5%13.1%19.2%2
Mandarin (Simplified)zh-CN7.9%14.0%20.4%3
Hindihi-IN9.2%15.8%23.1%3
Arabic (MSA)ar9.6%16.2%24.0%2
Russianru-RU6.8%11.5%17.4%3
Indonesianid-ID7.1%12.4%18.5%2

Speaker counts per language are small (2 to 4). The figures above are indicative, not statistically robust, and no confidence intervals are published for them.

Sample files are not published. The corpus is drawn from public-domain sources (see Test methodology above); customer audio is never used, and the individual audio files used for the retest are not posted publicly.

Test methodology

Speaker diarization

Diarization attaches speaker labels to a transcript. It depends on the provider that handles the job, the request settings and the recording. The adapter capabilities are listed in the model table above. A fallback can change whether speaker labels are available.

Language support does not guarantee correct speaker identification or word timing. We have not published a speaker-identification benchmark for every supported language.

Supported transcription languages

ScreenApp's transcription providers cover 100+ languages. We count the union of their documented language lists: a language supported by several providers counts once. The list below contains 100 entries, including Sinhala (Sinhalese) and Cantonese. Gemini does not publish an exhaustive transcription-language list, so this is a documented minimum, not a claim that coverage stops here.

  • Afrikaans
  • Albanian
  • Amharic
  • Arabic
  • Armenian
  • Assamese
  • Azerbaijani
  • Bashkir
  • Basque
  • Belarusian
  • Bengali
  • Bosnian
  • Breton
  • Bulgarian
  • Burmese
  • Cantonese
  • Catalan
  • Chinese (Mandarin)
  • Croatian
  • Czech
  • Danish
  • Dutch
  • English
  • Estonian
  • Faroese
  • Finnish
  • French
  • Galician
  • Georgian
  • German
  • Greek
  • Gujarati
  • Haitian Creole
  • Hausa
  • Hawaiian
  • Hebrew
  • Hindi
  • Hungarian
  • Icelandic
  • Indonesian
  • Italian
  • Japanese
  • Javanese
  • Kannada
  • Kazakh
  • Khmer
  • Korean
  • Lao
  • Latin
  • Latvian
  • Lingala
  • Lithuanian
  • Luxembourgish
  • Macedonian
  • Malagasy
  • Malay
  • Malayalam
  • Maltese
  • Maori
  • Marathi
  • Mongolian
  • Nepali
  • Norwegian
  • Nynorsk
  • Occitan
  • Pashto
  • Persian
  • Polish
  • Portuguese
  • Punjabi
  • Romanian
  • Russian
  • Sanskrit
  • Serbian
  • Shona
  • Sindhi
  • Sinhala (Sinhalese)
  • Slovak
  • Slovenian
  • Somali
  • Spanish
  • Sundanese
  • Swahili
  • Swedish
  • Tagalog / Filipino
  • Tajik
  • Tamil
  • Tatar
  • Telugu
  • Thai
  • Tibetan
  • Turkish
  • Turkmen
  • Ukrainian
  • Urdu
  • Uzbek
  • Vietnamese
  • Welsh
  • Yiddish
  • Yoruba

How we count languages

Whisper Large-v3 Turbo's model configuration supplies the broadest enumerated list. Its model-card tags still say 99 languages, but its language-token map includes Cantonese as another entry. We use the model configuration for the list above. For counting, Filipino and Tagalog are grouped as one entry, and Javanese code aliases are merged. Regional variants and automatic detection do not add entries. These sources were checked on 2026-09-16.

Coverage and routing limits

Vendor model coverage is broader than the app's 58-language selector and its provider routing lists. The union describes transcription coverage across the providers, not a guarantee that every provider or live session accepts every listed language. Transcript translation is a separate capability and does not add entries to this count.

Sinhala is present in both the app registry and Whisper's model configuration. With Sinhala explicitly selected, ScreenApp's current routing declarations prefer Google Gemini. Other batch providers are degraded fallbacks; live routing excludes them. Google's audio documentation does not separately verify Sinhala transcription accuracy.

Accuracy varies by language, accent and recording conditions. Inclusion in a vendor's language list does not establish a ScreenApp benchmark result. No Sinhala accuracy benchmark has been published on this page. The dated measurements on this page cover only their stated samples.

AI summaries and chat: what to expect

OpenAI GPT-5.6 Luna is the default model for AI chat, summaries, titles, templates and translation. Anthropic Claude Sonnet 5 and Google Gemini 3.1 Flash Lite can be selected for specific jobs instead. Watching video frames and listening to audio always runs on Google Gemini.

Summaries and chat answers are generated from the transcript, not from the original audio. A transcription error carries into the summary and into any answer that depends on that part of the transcript.

AI summaries and chat answers can be incomplete or wrong. For anything a decision depends on, check the summary or answer against the transcript or the original recording rather than taking it at face value.

Timestamps and the full transcript are there for exactly this: to let you jump to the source and verify what was actually said. No formal published evaluation of summary or chat quality exists yet.

Platform availability

ScreenApp is available as the web app, Mac (Apple silicon) and Windows apps, iOS and Android apps, and a Chrome extension that records tab, screen, mic and camera (Chrome, Edge, Brave, Opera and other Chromium browsers). Ratings and review counts below are pulled at build time from each platform's canonical store listing (the date is named next to each rating). Other listing details (version numbers, sizes, install dates) are point-in-time snapshots checked when this section was last edited; the live store listings are always the authoritative source.

iOS app (iPhone, iPad, Apple Silicon Mac via Catalyst)

iOS App Privacy nutrition label

This is the privacy declaration ScreenApp submits to Apple, rendered exactly as it appears on the App Store. Apple's listing is authoritative; if the table below ever diverges from the live App Store page, the App Store page wins and this table should be corrected. ScreenApp declares no tracking data.

Group Category Data types
Data Used to Track You None. ScreenApp does not declare any tracking data.
Data Linked to You User Content Photos or Videos, Audio Data
Identifiers User ID
Diagnostics Performance Data
Data Not Linked to You Identifiers Device ID
Contact Info Email Address, Name
Diagnostics Crash Data, Other Diagnostic Data

Canonical source: the App Privacy section on the ScreenApp App Store page. Full data handling policies on the Trust Center.

macOS (native desktop app)

Android

Web app

Chrome extension

Rating and review counts move daily on the App Store and Google Play. The numbers on this page are point-in-time snapshots, dated above. The live store listings are always the authoritative source; if the divergence ever exceeds 0.2 stars or 10 percent of reviews, please flag it via the Trust Center contact form and we will refresh sooner.

Selected customer reviews

works like a charm, notes have been super helpful alongside chat function. loveeeee

Katharine Suy, Chrome Web Store, October 17, 2024

nice app for important summary

SAHIL, Google Play, January 30, 2026

Production corpus

The numbers below are real production counts, pulled at build time from the ScreenApp production database, the same one the app reads from. They are not marketing rollups, not rounded, and not estimated. Refresh cadence: every deploy. Last pulled: September 18, 2026.

5,852,564

recordings processed

transcribed and analysed in production

2,156,709

speakers diarized

unique speaker turns identified across the corpus

307,060

AI Q&A sessions

questions asked against transcribed media

136,079

voice dictations

captured via browser, iOS, and Android

69,009

meeting-bot sessions

bots that auto-joined Zoom, Google Meet and Microsoft Teams calls; recording without a bot is also available

405,393

analysed video metadata sets

meeting type, speakers, companies extracted

Recent activity (indexed proxy via videometainfo.createdAt): 902 recordings analysed in the last 24 hours, 7,950 in the last 7 days, 32,223 in the last 30 days. Daily rate of roughly 1,074 analyses per day.

Why videometainfo and not recordings directly: recordings._id is a UUID, so we cannot do indexed time-range queries on it. Each videometainfo doc maps 1:1 to a recording via the unique recordingId index, so the time-windowed counts above are a faithful proxy. Methodology and the open query module: below.

Registered accounts

8,292,149 registered accounts as of September 18, 2026. This is the number of account records in the users collection of our production database, read at build time with MongoDB's estimatedDocumentCount(), which uses collection metadata and can differ slightly from an exact count. It counts accounts, not people or paying customers: free and paid accounts are included, accounts that have never recorded anything are included, and one person with two sign-ups counts twice. There is no activity, payment or email-verification filter. Accounts whose record has been removed from the database are not counted. What the count cannot tell you: how many accounts are active, how many pay, how many are verified, and whether deleted accounts are always removed or only flagged. The figure refreshes on every deploy.

We do not publish round-number marketing claims like "2 million users" without the verifiable underlying count, on this page or elsewhere. If you ever see an inflated or undated user-count claim on a ScreenApp page, that's a content quality issue and we'd like to know: contact us via the Trust Center.

How we count

Every numeric claim on this page and across screenapp.io that depends on production data follows the same pipeline. Numbers are not curated, edited, or rounded for marketing.

  1. Source of truth: the ScreenApp production database, the same database the dashboard, mobile apps, and backend services read from. No marketing database, no cached marketing CMS.
  2. Build-time pull: the marketing site has no direct database access. At each deploy, the static-site build calls a read-only backend endpoint (GET /v2/site-data) that runs a small set of indexed aggregations and returns scalar counts. The query module is open inside the same repo at scripts/site-data-queries.ts.
  3. Frozen for the build: the returned numbers are written to a local SQLite file (data/stats.db) and read synchronously by every page at static-generation time. Within a single deploy the numbers do not drift; between deploys they refresh.
  4. Soft fail: if the endpoint is unavailable or returns an unexpected shape, the previous deploy's figures are reused and the build proceeds. The site never ships placeholder text in place of a missing number.
  5. Last refresh: the data on this page was pulled on September 18, 2026.

If you ever spot a number on the site that disagrees with a figure on this page, please flag it via the Trust Center. A divergence is a bug.

Free access and pricing

Free

  • Discussions IA et modèles par mois : 10
  • Transcriptions à essayer : 2
  • Fichiers ajoutés et importations à essayer : 3
  • Jusqu'à 45 minutes par enregistrement

Aucune carte bancaire requise

Pro

  • Discussions IA et modèles par mois : 50
  • Transcriptions par mois : 50
  • Fichiers ajoutés et importations: 500
  • Jusqu'à 2 heures chacun

7 jours gratuits, puis $228/an. Carte bancaire requise.

Max

  • Discussions IA et modèles : Sans limite
  • Transcriptions : Sans limite
  • Fichiers ajoutés et importations : Sans limite
  • Enregistrements de 3 heures maximum chacun
Pricing

Security and compliance

SOC 2 Type 2 audited annually. 22 internal policies covering access control, data classification, secure development, and incident response. Continuous control monitoring.

Our Trust Center (trust.inc/screenapp) lists the security controls and internal policies behind the SOC 2 Type 2 audit, and offers a security questionnaire. For the SOC 2 Type 2 report itself, email support@screenapp.io.

Sources and external benchmarks

Errata and corrections

Numbers on ScreenApp pages should match this page. If you find a page that contradicts these figures, that is a content quality bug we want to fix. Email support@screenapp.io; the corrections policy explains how we check reports and when we add a correction note.

Corrections log

Try ScreenApp on a real recording

Paste a URL or upload an audio or video file. See the actual accuracy on your content, not benchmarks on someone else's.

Start the 7-day free trial

7 jours gratuits, puis $228/an. Carte bancaire requise.