How to measure caption accuracy across your video library with Word Error Rate
A question we hear constantly from organizations managing large video libraries: "We need a way to measure the accuracy and consistency of captions across our library of training videos. They've been created by folks across the organization, so quality varies. Does anyone use a tool for this?"
The short answer: yes, and Recap has it built in. Here's how Word Error Rate tooling works and why it matters when caption quality varies across your team.
The problem: inconsistent captions at scale
Most organizations don't have a single person creating all their captions. Training videos, onboarding materials, recorded meetings, and compliance content get captioned by different teams, different tools, and at different times. The result is predictable: some files are near-perfect, others are riddled with errors, and nobody knows which is which until someone watches the whole thing.
Manual spot-checking doesn't scale. You can't review 300 training videos by reading every caption line. And generic "accuracy percentages" from vendors tell you about their system's average performance, not about your specific files.
What Word Error Rate actually measures
Word Error Rate (WER) compares your caption text against a verified reference transcript, word by word. It counts three types of errors:
- Substitutions: a word was replaced with a different word ("principle" instead of "principal")
- Deletions: a spoken word is missing from the caption entirely
- Insertions: extra words appear in the caption that were never spoken
The formula: WER = (substitutions + deletions + insertions) / total reference words
A 4% WER means 96% of words are correct. For most organizations, anything under 5% is considered good for AI-generated captions; under 1.5% is the target for compliance-critical content with human review.
For a deeper comparison of WER with Formatted Error Rate (which also covers punctuation, speaker labels, and non-speech elements), see our FER vs WER explainer.
How Recap's WER tooling works
Recap includes Word Error Rate computation directly in the platform. You don't need external scripts, spreadsheets, or third-party tools. Here's what it gives you:
Library-wide analysis
Upload reference transcripts for your files, and Recap computes WER for each one. You get a dashboard view showing every file's accuracy score at a glance, with color-coded indicators:
- Pass (green): WER below your threshold, no action needed
- Review (yellow): WER is marginal; worth a quick check
- Fail (red): WER exceeds your threshold; needs re-captioning or editing
Error breakdown per file
For each file, you can see exactly how many substitutions, deletions, and insertions occurred. This helps you diagnose patterns: if a file has high substitutions, the ASR engine might have struggled with accents or terminology. High deletions often indicate audio quality issues.
Word-level diff view
The most actionable feature: a side-by-side diff that highlights exactly which words diverge from the reference. You can see at a glance whether the errors are concentrated in one segment (suggesting a bad audio section) or scattered throughout (suggesting a poor-quality ASR model was used).
Custom accuracy thresholds
Every organization has different standards. A marketing team might accept 95% accuracy for internal meeting recordings. A compliance team might require 99% for training materials that employees must certify they completed. Set your own pass/fail thresholds to match.
WER + caption audits: the full picture
WER measures word-level accuracy, but caption quality involves more than just getting the words right. Recap also offers comprehensive caption audits that evaluate your entire library against professional captioning standards.
Word Error Rate analysis
Caption audits evaluate your entire library against professional captioning standards, including:
- Whether captions exist at all (many videos in large libraries have none)
- Synchronization and timing accuracy
- Proper speaker identification
- Punctuation, capitalization, and formatting
- Non-speech elements like [applause] or [music]
Think of it this way: WER tells you how accurate your captions are. A caption audit tells you how complete and compliant your accessibility posture is. Used together, they give you an objective, defensible quality baseline across your entire library.
For a deep dive on audit methodology and how institutions use audit results to prioritize remediation, see our guide on comprehensive caption audit best practices.
When to use WER tooling
WER analysis is most valuable when:
- You've inherited a library of captioned videos from multiple contributors and need to triage which ones need work
- You're switching ASR providers and want to objectively compare quality on your actual content
- You're preparing for a compliance audit and need to demonstrate that your captions meet a measurable accuracy standard
- Different teams contribute captions and you want consistency across departments
- You've updated your custom dictionary and want to verify that re-transcription improved accuracy for domain-specific terms
Getting started
If you already have reference transcripts (even for a sample of your library), you can run WER analysis today. If you don't have reference transcripts, a common approach is to have a human reviewer verify a representative sample of files, then use those as your reference set for ongoing quality monitoring.
For organizations that need the full picture, start with a caption audit to identify which videos lack captions entirely or have obvious quality issues, then use WER analysis for a deeper accuracy assessment on the files that matter most.
Book a demo to see WER tooling and caption audits in action on your own content.