works like a charm, notes have been super helpful alongside chat function. loveeeee
Katharine Suy, Chrome Web Store, October 17, 2024
Methodology
This page is the source of truth for every accuracy, speed, and language claim on ScreenApp.io. Numbers come from our internal test corpus, the Groq engineering case study, OpenAI's Whisper benchmarks, and xAI's published Grok Speech-to-Text benchmarks. Provider stack last verified against the codebase: September 2026.
Last updated: 2026-09-16.
The repository's default batch order is xAI, then Mistral, then Modulate, then Google, then Groq. Live captions use Groq, then xAI, then Mistral, then Modulate, then Google. Credentials, language fit and deployment overrides can change which provider handles a job. These are repository defaults checked on 2026-09-16, not a record of every production request.
| Provider | Default model or endpoint | Language routing declaration | Speaker labels |
|---|---|---|---|
| xAI | Grok STT (/v1/stt) | 25 languages declared by the adapter | Supported by adapter |
| Mistral | voxtral-mini-latest | 13 languages declared by the adapter | Supported by adapter |
| Modulate | velma-2-stt-batch | 57 languages declared by the adapter | Supported by adapter |
| gemini-2.5-flash | No fixed language list in the adapter | Supported by adapter | |
| Groq | whisper-large-v3-turbo | 24 languages in the app proficiency list, not the model total | Not supported by adapter |
The linked provider documentation describes each service. Provider language totals are not added together to calculate ScreenApp's supported-language count. In particular, Groq's internal routing list is narrower than Whisper's multilingual model coverage.
When a language is selected, batch routing prefers providers that declare it before degraded fallbacks. Live routing excludes providers marked unsuitable for that language. Automatic detection does not provide the same explicit-language filter.
No accuracy or latency benchmark for this current provider order is published here. See the dated measurements below for their original test conditions. Data handling is described in the Privacy Policy.
This section is history, not a current benchmark. In 2025, ScreenApp moved from a self-hosted Whisper deployment on AWS to Groq's inference infrastructure. Groq published the case study below; the numbers are from their engineering team's measurements of that pipeline, which was the default at the time. The current default models and routing are listed above. No speed or latency figures have been published yet for the current xAI-led provider chain, or for live captions.
| Metric | Before Groq | After Groq | Change |
|---|---|---|---|
| 20-minute transcription job | ~20 minutes | ~15 seconds | 20x faster |
| Per-minute transcription cost | baseline | 1/15th | 15x cheaper |
| Free-to-paid conversion | baseline | +30% | uplift |
| Annual recurring revenue (year-over-year) | baseline | +405% | growth attributed to the speed and cost gains |
Source: ScreenApp + Groq case study (groq.com).
Historical, for the 2025 Groq Whisper pipeline: a 60-minute meeting completes in roughly 3 minutes end-to-end (transcription, diarization, summary generation). A 2-hour video processes in about 6 minutes. These are end-to-end times that include summarization and chaptering, not just raw transcription. No speed figures have been published yet for the current xAI-led chain.
When a YouTube video already publishes a caption track, ScreenApp reads that track instead of transcribing the audio. No speech recognition runs at all, which is why it is fast. The text is written by YouTube, so none of the word error rates on this page describe it. Videos without a caption track go through the AI transcription path above, which is slower and is what our benchmarks measure.
| Path | Imports (n) | Median | p90 |
|---|---|---|---|
| Download + AI transcription (before, Jul 8 to 9 2026) | 61 | 41.1 sec | 159 sec |
| Caption import (after, Jul 16 to 17 2026) | 1,774 | 14.7 sec | 29 sec |
Method: both windows measure the same endpoint (import start to transcript indexed) on the same population, on production traffic. We state the gain as about 3x rather than a precise figure: the before-window sample is small (n=61), which puts the true median speedup somewhere between 1.9x and 3.7x at 95% confidence. The mean (3.58x) is inflated by a long tail of large videos, so the median is the honest central figure.
Two limits worth stating plainly. First, caption import currently serves free-tier YouTube imports; paid accounts still take the download and transcription path, which measured a 60.5 second median before and 44.7 seconds after (that difference is sampling noise, not a change in the product). Second, 14.7 seconds reflects the current implementation, which reads the captions already published with the video. Measured: July 8 to 17, 2026.
Word error rate (WER) counts substitutions + deletions + insertions per 100 reference words. Lower is better. Baseline figures below come from the published benchmarks for each underlying model; the per-condition rows are from our own April 2026 retest on 18 hours of public-domain audio per language across three conditions: studio (single speaker, treated room), conference (multi-speaker, room mic), and field (handheld phone mic, ambient noise).
Headline figure: up to 95.8% transcription accuracy on clear English audio in our April 2026 test, which is 100 minus the 4.2% English (US) Studio WER in the table below. This is measured on the April 2026 provider mix used for that retest, before xAI Grok Speech-to-Text became the default provider. It has not been re-measured on the current default provider.
| Language | Locale | Studio WER | Conference WER | Field WER | Speakers tested |
|---|---|---|---|---|---|
| English (US) | en-US | 4.2% | 7.8% | 12.4% | 4 |
| Spanish (Latin Am.) | es-419 | 5.1% | 9.2% | 14.6% | 3 |
| Spanish (Spain) | es-ES | 5.4% | 9.8% | 15.1% | 3 |
| Portuguese (BR) | pt-BR | 5.8% | 10.1% | 15.8% | 3 |
| Portuguese (PT) | pt-PT | 6.4% | 11.2% | 17.0% | 2 |
| French | fr-FR | 5.9% | 10.4% | 16.2% | 3 |
| German | de-DE | 6.1% | 10.8% | 16.5% | 3 |
| Italian | it-IT | 6.3% | 11.0% | 17.1% | 3 |
| Japanese | ja-JP | 7.8% | 13.5% | 19.8% | 2 |
| Korean | ko-KR | 7.5% | 13.1% | 19.2% | 2 |
| Mandarin (Simplified) | zh-CN | 7.9% | 14.0% | 20.4% | 3 |
| Hindi | hi-IN | 9.2% | 15.8% | 23.1% | 3 |
| Arabic (MSA) | ar | 9.6% | 16.2% | 24.0% | 2 |
| Russian | ru-RU | 6.8% | 11.5% | 17.4% | 3 |
| Indonesian | id-ID | 7.1% | 12.4% | 18.5% | 2 |
Speaker counts per language are small (2 to 4). The figures above are indicative, not statistically robust, and no confidence intervals are published for them.
Sample files are not published. The corpus is drawn from public-domain sources (see Test methodology above); customer audio is never used, and the individual audio files used for the retest are not posted publicly.
Diarization attaches speaker labels to a transcript. It depends on the provider that handles the job, the request settings and the recording. The adapter capabilities are listed in the model table above. A fallback can change whether speaker labels are available.
Language support does not guarantee correct speaker identification or word timing. We have not published a speaker-identification benchmark for every supported language.
ScreenApp's transcription providers cover 100+ languages. We count the union of their documented language lists: a language supported by several providers counts once. The list below contains 100 entries, including Sinhala (Sinhalese) and Cantonese. Gemini does not publish an exhaustive transcription-language list, so this is a documented minimum, not a claim that coverage stops here.
Whisper Large-v3 Turbo's model configuration supplies the broadest enumerated list. Its model-card tags still say 99 languages, but its language-token map includes Cantonese as another entry. We use the model configuration for the list above. For counting, Filipino and Tagalog are grouped as one entry, and Javanese code aliases are merged. Regional variants and automatic detection do not add entries. These sources were checked on 2026-09-16.
Vendor model coverage is broader than the app's 58-language selector and its provider routing lists. The union describes transcription coverage across the providers, not a guarantee that every provider or live session accepts every listed language. Transcript translation is a separate capability and does not add entries to this count.
Sinhala is present in both the app registry and Whisper's model configuration. With Sinhala explicitly selected, ScreenApp's current routing declarations prefer Google Gemini. Other batch providers are degraded fallbacks; live routing excludes them. Google's audio documentation does not separately verify Sinhala transcription accuracy.
Accuracy varies by language, accent and recording conditions. Inclusion in a vendor's language list does not establish a ScreenApp benchmark result. No Sinhala accuracy benchmark has been published on this page. The dated measurements on this page cover only their stated samples.
OpenAI GPT-5.6 Luna is the default model for AI chat, summaries, titles, templates and translation. Anthropic Claude Sonnet 5 and Google Gemini 3.1 Flash Lite can be selected for specific jobs instead. Watching video frames and listening to audio always runs on Google Gemini.
Summaries and chat answers are generated from the transcript, not from the original audio. A transcription error carries into the summary and into any answer that depends on that part of the transcript.
AI summaries and chat answers can be incomplete or wrong. For anything a decision depends on, check the summary or answer against the transcript or the original recording rather than taking it at face value.
Timestamps and the full transcript are there for exactly this: to let you jump to the source and verify what was actually said. No formal published evaluation of summary or chat quality exists yet.
ScreenApp is available as the web app, Mac (Apple silicon) and Windows apps, iOS and Android apps, and a Chrome extension that records tab, screen, mic and camera (Chrome, Edge, Brave, Opera and other Chromium browsers). Ratings and review counts below are pulled at build time from each platform's canonical store listing (the date is named next to each rating). Other listing details (version numbers, sizes, install dates) are point-in-time snapshots checked when this section was last edited; the live store listings are always the authoritative source.
This is the privacy declaration ScreenApp submits to Apple, rendered exactly as it appears on the App Store. Apple's listing is authoritative; if the table below ever diverges from the live App Store page, the App Store page wins and this table should be corrected. ScreenApp declares no tracking data.
| Group | Category | Data types |
|---|---|---|
| Data Used to Track You | None. ScreenApp does not declare any tracking data. | |
| Data Linked to You | User Content | Photos or Videos, Audio Data |
| Identifiers | User ID | |
| Diagnostics | Performance Data | |
| Data Not Linked to You | Identifiers | Device ID |
| Contact Info | Email Address, Name | |
| Diagnostics | Crash Data, Other Diagnostic Data | |
Canonical source: the App Privacy section on the ScreenApp App Store page. Full data handling policies on the Trust Center.
ScreenApp-latest.dmg and always serves the current production build, so there is no separate version string to copy here.io.screenapp.screenapp_mobile.Rating and review counts move daily on the App Store and Google Play. The numbers on this page are point-in-time snapshots, dated above. The live store listings are always the authoritative source; if the divergence ever exceeds 0.2 stars or 10 percent of reviews, please flag it via the Trust Center contact form and we will refresh sooner.
works like a charm, notes have been super helpful alongside chat function. loveeeee
Katharine Suy, Chrome Web Store, October 17, 2024
nice app for important summary
SAHIL, Google Play, January 30, 2026
The numbers below are real production counts, pulled at build time from the ScreenApp production database, the same one the app reads from. They are not marketing rollups, not rounded, and not estimated. Refresh cadence: every deploy. Last pulled: September 18, 2026.
5,852,564
recordings processed
transcribed and analysed in production
2,156,709
speakers diarized
unique speaker turns identified across the corpus
307,060
AI Q&A sessions
questions asked against transcribed media
136,079
voice dictations
captured via browser, iOS, and Android
69,009
meeting-bot sessions
bots that auto-joined Zoom, Google Meet and Microsoft Teams calls; recording without a bot is also available
405,393
analysed video metadata sets
meeting type, speakers, companies extracted
Recent activity (indexed proxy via videometainfo.createdAt): 902 recordings analysed in the last 24 hours, 7,950 in the last 7 days, 32,223 in the last 30 days. Daily rate of roughly 1,074 analyses per day.
Why videometainfo and not recordings directly: recordings._id is a UUID, so we cannot do indexed time-range queries on it. Each videometainfo doc maps 1:1 to a recording via the unique recordingId index, so the time-windowed counts above are a faithful proxy. Methodology and the open query module: below.
8,292,149 registered accounts as of September 18, 2026. This is the number of account records in the users collection of our production database, read at build time with MongoDB's estimatedDocumentCount(), which uses collection metadata and can differ slightly from an exact count. It counts accounts, not people or paying customers: free and paid accounts are included, accounts that have never recorded anything are included, and one person with two sign-ups counts twice. There is no activity, payment or email-verification filter. Accounts whose record has been removed from the database are not counted. What the count cannot tell you: how many accounts are active, how many pay, how many are verified, and whether deleted accounts are always removed or only flagged. The figure refreshes on every deploy.
We do not publish round-number marketing claims like "2 million users" without the verifiable underlying count, on this page or elsewhere. If you ever see an inflated or undated user-count claim on a ScreenApp page, that's a content quality issue and we'd like to know: contact us via the Trust Center.
Every numeric claim on this page and across screenapp.io that depends on production data follows the same pipeline. Numbers are not curated, edited, or rounded for marketing.
GET /v2/site-data) that runs a small set of indexed aggregations and returns scalar counts. The query module is open inside the same repo at scripts/site-data-queries.ts.data/stats.db) and read synchronously by every page at static-generation time. Within a single deploy the numbers do not drift; between deploys they refresh.If you ever spot a number on the site that disagrees with a figure on this page, please flag it via the Trust Center. A divergence is a bug.
Aucune carte bancaire requise
7 jours gratuits, puis $228/an. Carte bancaire requise.
SOC 2 Type 2 audited annually. 22 internal policies covering access control, data classification, secure development, and incident response. Continuous control monitoring.
Our Trust Center (trust.inc/screenapp) lists the security controls and internal policies behind the SOC 2 Type 2 audit, and offers a security questionnaire. For the SOC 2 Type 2 report itself, email support@screenapp.io.
Numbers on ScreenApp pages should match this page. If you find a page that contradicts these figures, that is a content quality bug we want to fix. Email support@screenapp.io; the corrections policy explains how we check reports and when we add a correction note.
Paste a URL or upload an audio or video file. See the actual accuracy on your content, not benchmarks on someone else's.
7 jours gratuits, puis $228/an. Carte bancaire requise.