ASR captions could be costing you a third of your viewers

A television screen with a coffee table in the foreground. A TV remote and a popcorn-filled wooden bowl on the table convey how captions improve viewers’ experience.

Why relying on Automated Speech Recognition for your captions could be losing you 33% of your audience, and how editing fixes it

Automated Speech Recognition (ASR) tools are pretty impressive. They can turn hours of audio into text in minutes, and they’re constantly improving. It’s no wonder that so many of us turn to ASR tools for a fast, low-effort way to generate captions and transcripts.

When you’re producing video content, captions and transcripts should absolutely form part of your workflow. They’re essential for accessibility, user experience and discoverability.

But here’s the catch: speed doesn’t equal accuracy. If you need your video content to meet accessibility standards, relying on ASR tools alone isn’t enough.

As audiovisual translation specialists, we have extensive experience in editing ASR transcripts. To show how easily ASR errors accumulate, even when confidence scores are high, here’s what we found when editing an ASR transcript.

Key takeaways from this article:

  • ASR confidence scores don’t guarantee accuracy – human editing remains essential.
  • Unedited ASR captions exclude one-third of your potential audience.
  • Professional transcript editing prevents costly localisation errors downstream.

The assignment: ASR transcript editing

Editing an ASR transcript to create accessible, broadcast-ready English captions.

  • Video duration: 45 minutes
  • Number of speakers: 5
  • Total word count: 8,070
  • Average number of words per minute: 179
  • The ASR tool’s confidence score: 95%

Let’s dig into that last figure first.

An ASR confidence score of 95%: high confidence, low usability

Who isn’t just a little bit wowed by a 95% confidence score? On paper, it sounds fantastic. Let’s explore what that really means in practice.

It doesn’t mean your transcript is 95% accurate – it means the algorithm is 95% confident in its guess. Confidence and correctness are not the same thing.

What seems like an acceptable error rate actually represents a huge number of inaccuracies over longer video durations.

In our example, the ASR transcript contained:

  • 614 timing errors
  • 173 punctuation errors
  • 70 misheard words
  • 58 missing words
  • 12 homophone mix-ups

These errors can distort meaning and make captions hard to follow. For deaf and hard-of-hearing viewers, those gaps aren’t inconveniences. They’re barriers to understanding.

614 timing errors: when captions become unreadable

Our ASR transcript contained 614 timing errors. Typically, captions should appear on screen for at least one second. In our transcript, some captions only stayed on screen for a tenth of a second. Nobody reads that quickly.

Many captions appeared on screen a few frames before a shot change. You can’t spot this error just by looking at the transcript. But when the captions appear on screen, they cause an uncomfortable flashing effect.

Timing issues like these don’t just make captions irritating; they make them inaccessible. These issues become more apparent when your video has many speakers and lots of rapid, back-and-forth dialogue.

For viewers who rely on captions, this is the digital equivalent of being shut out of the room.

ASR tools don’t take into consideration readability or cognitive load. Human editors do.

173 punctuation errors that distort meaning

Punctuation isn’t decorative – it’s essential for meaning. ASR tools don’t reliably reflect tone, pauses or sentence boundaries. A comma changes the meaning entirely in the two sentences below:

  • Let’s eat, Grandma.
  • Let’s eat Grandma.

Imagine similar punctuation issues scattered across a 45-minute video, and you begin to see how clarity breaks down.

70 misheard words that change the message

The ASR tool misheard words 70 times. When captions say something different from what’s spoken, you lose your audience’s trust.

Here are just a few examples that appeared in our 45-minute video:

ASR outputWhat was actually spoken
mischiefmystery
jobdrop
himthem
tidytiny

A single incorrect word can derail the message or misrepresent someone. It’s a serious issue in news, training, healthcare and customer-facing content.

58 missing words: when captions fail to capture crucial information

Our ASR transcript missed 58 words. This often happens when the audio quality isn’t perfect. But poor audio quality is one of the key reasons viewers switch on captions in the first place.

If your captions aren’t conveying key information, they’re not serving their primary purpose: accessibility.

12 homophone mix-ups that sound right but read wrong

ASR has come a long way with contextual recognition, but it still makes mistakes. Here are a few examples that appeared in our script:

ASR outputWhat was actually spoken
twotoo
forfour
bybuy

Small errors like these can make your content come across as unpolished and unprofessional. Not the impression you want to make with your brand content.

Why accurate captions matter

Unedited ASR transcripts exclude 33% of your target audience

One in three adults in the UK are deaf, have hearing loss or live with tinnitus. When captions are inaccurate, poorly timed or confusing, they actively exclude people:

  • Deaf and hard-of-hearing viewers
  • Neurodivergent viewers
  • People watching on low volume or in noisy environments

ASR errors reinforce bias and harm your brand’s reputation

Speech-to-text tools were trained on vast amounts of data. But ASR tools reveal bias in unexpected ways. Accuracy is known to drop for:

  • Strong regional accents
  • Faster speakers
  • Quiet speakers
  • Non-native speakers
  • People with speech disorders

Neglecting to edit ASR transcripts can reinforce bias and contribute to a sense of “othering”. This can lead to unfair or inaccurate representation of speakers – something no brand wants. A professional ASR transcript editor makes sure that the voices of marginalised speakers are heard.

What you can do to improve ASR output

There are some straightforward steps you can take to reduce these risks. Most of them start before you generate the ASR transcript.

Start with the recording environment

ASR accuracy drops when you record in a noisy environment. Find a quiet space with no background noise. Make sure your speakers are using good-quality microphones, and ask them all to speak clearly.

Be mindful when you’re editing your video

Make sure that the audio quality in your video is as high as it can possibly be. Even adding background music to your video can lower the accuracy of ASR.

Use the ASR tool’s custom glossary

You can influence how ASR recognises acronyms and proper names by adding their correct spellings to a custom glossary. This helps avoid confusing or embarrassing renderings.

Edit the ASR output

You can do this yourself if you speak the language. But having a professional editor check and edit the ASR output will make all the difference. A professional will not only remove all of these kinds of errors. They’ll also:

  • Time the captions around shot changes for a comfortable viewing experience
  • Add labels to identify speakers and important non-speech sounds
  • Segment captions and adjust line breaks to minimise cognitive load
  • Make sure that the reading speed is appropriate

ASR transcript editing is an important shift left strategy

Shift-left approaches help prevent issues from snowballing later in the content life cycle. They also reduce costs and the need for rework.

Meet accessibility requirements

Making digital content accessible is increasingly becoming a legal requirement. Avoid unnecessary compliance risks by incorporating ASR transcript editing into your workflow.

Here’s what the Web Accessibility Initiative has to say:

“Automatically generated captions do not meet user needs or accessibility requirements, unless they are confirmed to be fully accurate. Usually they need significant editing.”

Be localisation ready

Planning to subtitle your video into other languages? Any errors that appear in your ASR transcript could be reflected in the translated subtitles. Unchecked ASR captions multiply the cost and complexity downstream.

The difference between a smooth localisation workflow and costly delays? Professional ASR transcript editing before localisation.

Key takeaways

ASR is a great starting point for speeding up transcription. It’s become an important part of a streamlined multimedia localisation workflow. But it still lags behind on accuracy. Uncorrected errors compromise accessibility, clarity and audience experience. Human editors remain essential for:

  • Caption readability
  • Accuracy
  • Correct punctuation
  • Speaker identification
  • Sensitivity to context
  • Accessibility compliance
  • Protecting your brand

If you care about your audience, your brand and your message, invest in professional ASR transcript editing. Get in touch with us today to find out how we can help you transform your ASR transcripts into accessible captions that win your audience’s trust.

Do you need help with ASR captions?

More Automated Speech Recognition questions? Here are the questions people often ask after reading this post:

How much does professional ASR transcript editing cost compared to starting from scratch?

Professional ASR transcript editing typically costs less than creating captions from scratch because the editor works from an existing base. The exact cost depends on video length, audio quality, and turnaround time. Contact a specialist captioning provider for a tailored quote.

How long does it take to professionally edit ASR-generated captions?

Editing time varies depending on audio complexity, number of speakers, and technical terminology. A skilled editor typically needs two to four times the video duration to produce broadcast-ready captions. Rush turnarounds are often available for time-sensitive projects at additional cost.

What qualifications should I look for in an ASR transcript editor?

Look for editors with captioning or subtitling experience, knowledge of accessibility standards like WCAG, and familiarity with your industry’s terminology. Certifications such as CPACC demonstrate accessibility expertise. Native-language proficiency and experience with broadcast-standard timing are also essential.

Do edited ASR captions meet legal accessibility requirements in the UK and EU?

Properly edited captions can meet accessibility requirements, including WCAG guidelines and the European Accessibility Act. However, automatically generated captions alone do not meet these standards. Professional editing ensures compliance by correcting errors, timing captions appropriately, and adding speaker identification.

Can ASR transcript editing improve my video’s SEO performance?

Yes. Accurate captions and transcripts help search engines understand your video content, improving discoverability. Error-filled ASR output can confuse search algorithms and damage your brand’s credibility. Clean, professionally edited transcripts boost both accessibility and search rankings simultaneously.

About the author

Bethan Thomas CPACC has worked in the language services industry for 20 years. At Planet Languages, she manages audiovisual translation projects for companies that embrace accessibility. She has also provided English captioning to the UK broadcasting industry since 2015.

Categories
Tags