Skip to main content

Multi-provider live captioning: Deepgram and ElevenLabs join Recap's speech recognition lineup

· 6 min read
AI-Powered Video Accessibility Solutions

Live captioning for public meetings, lectures, and broadcasts depends entirely on the speech recognition engine running behind the scenes. If that engine goes down, captions stop. For organizations with accessibility obligations under ADA Title II or WCAG 2.1 AA, an outage during a city council session or commencement ceremony is not just inconvenient; it is a compliance gap.

Recap now supports three live speech recognition providers: Deepgram, ElevenLabs, and Google. When one provider experiences errors, Recap automatically fails over to the next, keeping captions running without manual intervention.

Why multiple speech providers matter​

Cloud services go down. Even the largest providers experience outages that last hours or, in some cases, days. When your live captioning depends on a single speech recognition API, an outage means your audience loses access at the worst possible time: during the live event itself.

This is not a hypothetical risk. Major cloud speech APIs have experienced extended outages that affected customers globally. For a university streaming a graduation ceremony to thousands of families, or a city council broadcasting a public hearing that residents rely on for civic participation, losing captions mid-event undermines the accessibility commitment those organizations have made.

Adding multiple providers solves this in two ways:

  • Automatic failover. If the active provider returns errors during a live session, Recap detects the failure and transitions to the next available provider in succession. Captions resume within seconds, and the switch is invisible to viewers.
  • Provider choice during setup. When configuring a live captioning stream, you can select which provider to use as your primary engine. This lets you match the provider's strengths to your content type.

What each provider brings​

Each speech recognition engine has characteristics that make it a better fit for certain content. We tested all three extensively across a range of real-world audio, and here is what we found.

Deepgram​

Deepgram excels at segmenting continuous speech into readable caption cues. For long-winded audio with few natural pauses, Deepgram is significantly better at splitting the stream into short, well-timed captions without requiring manual timer-based commits. This matters for content like lectures, sermons, or public comment periods where speakers talk at length without stopping.

Deepgram also delivers consistently high confidence scores and handles overlapping speech well, making it a strong default for most live captioning scenarios.

ElevenLabs​

ElevenLabs produces notably better punctuation, capitalization, and quotation placement. When a speaker says something like "I was like, 'Now what?'" ElevenLabs captures the quotes and formatting more accurately than other providers. This is particularly valuable for:

  • City council and board meetings where speakers quote ordinances, policies, or public comments
  • Lectures that reference published material or dialogue
  • Any content where proper punctuation affects comprehension of the transcript

For organizations that use live session transcripts as official records or meeting minutes drafts, the improved formatting from ElevenLabs reduces the amount of post-event editing required.

Google​

Google's speech recognition remains a solid, well-established engine with broad language support and strong general-purpose accuracy. It continues to serve as a reliable option in the provider lineup, particularly for multilingual captioning scenarios.

How automatic failover works​

Recap monitors the active speech recognition provider throughout every live session. If the provider begins returning 500 errors or connection failures, the system automatically transitions to the next provider in the configured order.

The failover sequence is configurable (connect with support to make changes), but the default prioritizes Deepgram or ElevenLabs first, with the remaining providers as backups. Here is how the process works in practice:

  1. Primary provider fails. The active provider returns errors during the live stream.
  2. Recap detects the failure. The system identifies consecutive errors and initiates a switch.
  3. Next provider activates. The backup provider begins processing the audio stream.
  4. Captions resume. Viewers see captions continue with minimal interruption.

If the second provider also fails, the system moves to the third. This three-deep failover chain provides a level of resilience that a single-provider setup cannot match.

Beyond automated failover, human caption moderators always retain full control. At any point during a live session, a moderator can switch AI captioning off and immediately resume with manual captioning or re-speaking. This means that even in a scenario where all three providers experience issues simultaneously, your team can continue delivering captions without interruption.

Choosing a provider during stream setup​

When setting up a live captioning stream, you can now select your preferred speech recognition provider. The choice appears in the stream configuration interface alongside other settings like custom dictionaries and caption display options.

Consider these guidelines when choosing:

ScenarioRecommended providerWhy
Long continuous speech with few pausesDeepgramBetter automatic segmentation into readable cues
Content with quotes, dialogue, or formal languageElevenLabsMore accurate punctuation and quotation handling
Multilingual events or general-purpose captioningGoogleBroad language coverage and consistent baseline
Maximum resilience (no preference)Deepgram or ElevenLabs as primaryStrong accuracy with automatic failover to remaining providers

Whichever provider you choose as your primary, the other two remain available as automatic failover targets.

What this means for compliance​

Organizations subject to WCAG 2.1 Success Criterion 1.2.4 (Captions Live) need captions to be available during live synchronized media. An outage that takes down your speech recognition API does not create a compliance exception; the obligation to provide live captions still applies.

Multi-provider failover gives procurement teams and accessibility coordinators a concrete answer to the resilience question that appears in many RFPs for real-time captioning services: "What happens if the primary system fails during a live event?" The answer is now straightforward: Recap switches to an alternative provider automatically.

This also aligns with the operational continuity expectations in government and education contracts, where service level agreements often require backup plans for critical accessibility infrastructure.

Available now​

Multi-provider live captioning with automatic failover is available to all Recap customers using AI-tier or hybrid live captioning. No additional configuration is required for failover; it is enabled by default. Provider selection is available in the stream setup interface.

Next steps​

  1. If you are already using Recap for live captioning, try selecting a different provider for your next session and compare the output quality for your content type.
  2. If you are evaluating live captioning vendors, book a demo to see multi-provider failover in action with your own audio.
  3. Review our real-time closed captioning services guide for a full comparison of AI, hybrid, and CART captioning tiers.
  4. For government buyers, see our city council quick start guide for setup recommendations.