Skip to main content

Introducing Voice Mode: Create Audio Descriptions and Edit Transcripts by Speaking

· 7 min read
Scott Griffin
Scott Griffin
Co-Founder & CEO, Recap Innovations

Recap Innovations already offers fully automated AI audio descriptions that can process thousands of videos with same-day turnaround. For many institutions, that is the right solution. But not every team can use AI-generated content. Some operate under data governance policies that prohibit sending video to third-party AI models. Others require human-authored descriptions for quality assurance, subject-matter accuracy, or institutional policy reasons. And some simply want a faster way to review and refine AI output before publishing.

Today we are introducing voice mode: a new way to create audio descriptions and edit transcripts by speaking instead of typing. It gives human describers the productivity benefits of automation without requiring any AI processing of video content.

Who Voice Mode Is For​

Voice mode is designed for teams that need a human in the loop, whether by choice or by policy.

Data governance and privacy constraints​

Government agencies, healthcare organizations, and some higher education institutions operate under policies that restrict sending video content to external AI services. If recordings contain student data protected by FERPA, patient information covered by HIPAA, or classified material, cloud-based AI processing may not be an option. Voice dictation keeps video content in the browser; only microphone audio is sent for speech recognition.

Subject-matter expertise​

For specialized content like medical procedures, legal proceedings, advanced mathematics, or indigenous language instruction, a subject-matter expert will produce more precise descriptions than a general-purpose AI model. A chemistry professor describing a titration experiment captures pedagogical context that AI cannot infer from visual frames alone.

Contractual or accreditation requirements​

Some procurement contracts, grant-funded projects, or accreditation bodies specify that accessibility accommodations must be human-authored. Voice dictation produces genuinely human-created descriptions efficiently.

Reviewing and correcting AI output​

Even teams that use AI-generated audio descriptions often need to review the output. Voice dictation makes that review process significantly faster by letting editors speak corrections instead of typing them.

The Voice AD Studio​

The Voice AD Studio is a dedicated workspace for creating or editing audio description tracks using your voice. It presents the video as a series of timed cues, each paired with a thumbnail and a text area for the description.

Creating a new track from scratch​

When no audio description track exists yet, the studio generates blank cues spread across the video at a configurable interval. You choose the spacing between cues (3, 5, 10, 15, or 30 seconds) and the duration of each cue. The studio generates the full set of timed slots, and you fill them by speaking.

The workflow is continuous:

  1. Review the thumbnail to see what is happening visually at that point in the video
  2. Speak your description into the microphone; words appear in the text field in real time
  3. Press Alt+Down to advance to the next cue and keep speaking

Voice mode maintains a persistent dictation stream. Moving between cues with keyboard shortcuts instantly redirects your speech to the active cue without stopping and restarting the microphone. You can describe an entire video in a single unbroken session.

Editing an existing track​

The same tools work for refining descriptions that were generated by AI or created previously. Open any audio description track in the studio, click the microphone icon on a specific cue, and speak your correction. Recognized text is inserted at the cursor position, so you can append to existing descriptions or replace selected text.

Inline Voice Dictation for Transcript Editing​

Voice dictation extends beyond audio descriptions to every caption cue in Recap's transcript editor. Each cue now includes a microphone button. Click it or press Cmd+Shift+M (Ctrl+Shift+M on Windows) to begin dictating directly into that cue.

This is useful for:

  • Correcting misrecognized words in AI-generated transcripts, where speaking the correct term is faster than selecting and retyping
  • Adding context or speaker labels to caption cues
  • Filling in gaps where the AI transcript missed content due to poor audio quality or overlapping speakers
  • Mobile and tablet editing where typing in small text fields is cumbersome

Interim results appear in real time as you speak, and finalized text is saved to the caption cue automatically.

Efficiency Gains​

Audio description creation: two to three times faster​

Most people speak two to three times faster than they type. The difference is especially pronounced for descriptive writing, where the goal is to capture what is visible in natural language. Speaking mirrors how a sighted person would naturally explain a scene.

A typical five-minute educational video might contain 20 to 30 audio description cues. Typing each description takes an experienced describer roughly 30 to 45 minutes. With voice mode, the same describer can complete the track in 10 to 20 minutes, spending time watching and speaking rather than watching, pausing, and typing.

For teams processing dozens or hundreds of videos per semester, this reduction compounds significantly. A media accessibility coordinator who previously spent 40 hours per week on audio description can handle the same volume in 15 to 20 hours, freeing time for quality review.

Faster AI review cycles​

Teams that use AI-generated descriptions as a starting point benefit from voice dictation during the review phase. Instead of clicking into a text field, selecting incorrect text, and typing a correction, the reviewer can simply speak the replacement. Attention stays on the content rather than on the mechanics of editing.

Reduced repetitive strain​

For accessibility professionals who spend hours per day typing descriptions, voice dictation offers an ergonomic benefit. Alternating between typing and speaking distributes the physical workload and reduces the risk of repetitive strain injuries.

Keyboard Shortcuts​

The Voice AD Studio is designed for keyboard-driven workflows:

  • Cmd+Shift+M (Ctrl+Shift+M on Windows): toggle voice mode
  • Alt+Up/Down arrow: navigate between cues without stopping dictation
  • Cmd+Enter: finalize the current cue and advance to the next
  • Escape: stop dictation

These shortcuts work even when a text area is focused.

Getting Started​

To create an audio description track with voice dictation:

  1. Open any file in your Recap workspace
  2. Navigate to the audio descriptions section and select Create with voice from the menu
  3. Configure your preferred cue interval and duration, then click Start creating
  4. Toggle voice mode with Cmd+Shift+M and begin describing each cue

For inline transcript editing, open any transcript in edit mode and click the microphone icon on any caption cue.

If your team processes video at scale and wants to start with AI-generated descriptions instead, see our companion article on AI vs. human audio description workflows for guidance on choosing the right approach.

Voice dictation is currently in preview. We welcome feedback from teams using it in production workflows. Reach out to our team with suggestions or questions.


About Recap Innovations​

Recap Innovations provides comprehensive video accessibility solutions for educational institutions and enterprises. Our platform combines AI-powered automation with professional-grade quality for captions, transcripts, audio descriptions (standard and extended modes), translations, voiceovers, and live captioning. With native integrations across major video platforms, cloud storage providers, conferencing tools, and Learning Management Systems, we help organizations meet accessibility requirements while serving diverse global audiences. We are the leading platform to offer AI solutions for all time-based media requirements of the Web Content Accessibility Guidelines (WCAG) at the 2.1 AA level.

Learn more about our audio description capabilities, explore our platform features, or contact our team to see voice dictation in action.