Captions Stopped Being an Accommodation. Audio Description Has Not.

Audio description is closer to everyday than you think. See why it's catching up to captions, and four no-cost habits that close the gap this month.
What You'll Learn
Why captions made the curb-cut trip from accommodation to everyday setting, and audio description hasn't.
The three structural reasons audio description lags, none of which are about anyone caring less.
Why caption reading speed is the only on-screen-text rule anybody actually enforces.
What makes audio description cheap one way and brutal the other.
Four free habits your team can start this month.
Researchers asked 542 people why they turn captions on. Six possible reasons: sound quality, attention, academic use, foreign language, outside influence, and being hard of hearing.
Being hard of hearing came in last.
The least common reason people turn on captions is the exact reason captions were invented. That's not a failure. Captions were built for people who depended on them, then adopted by people who just liked them better, and somewhere in there they stopped being an accommodation and became a setting.
Audio description is sitting about where captions were fifteen years ago. Why it hasn't made that jump is more specific than "people haven't caught on," and a couple of the reasons are things you can act on in a few weeks.
Captions Already Made This Trip
You know the curb-cut story. Ramps built for wheelchair users now serve everything from wheelchairs to strollers to suitcases. It came up again in our conversation with Jay Wyant, Minnesota's Chief Information Accessibility Officer, on the Government Video Podcast last fall, where he made the case for accessibility as a shared value instead of a checkbox.
Here's the part worth noticing: captions actually finished the trip. Most accommodations don't.
Think about where you see captions now. The treadmill screen at the gym. The muted phone video in line at the pharmacy. A quiet apartment at 1 a.m. and a loud bar at 8 p.m. Nobody there is thinking about compliance. They're thinking about whether they can follow the show.
That's what finishing the trip looks like. The feature stops being for a group and starts being for a situation. Once that happens, nobody has to be convinced of anything.
Audio Description Hasn't, and the Reasons Are Structural
Audio description isn't lagging because anyone cares less about blind and low-vision viewers. It's lagging for three reasons that have almost nothing to do with intentions.
The mandate is a different shape. When the National Association of the Deaf sued Netflix, the consent decree covered the entire streaming library. Not a percentage of it, all of it. Audio description came up through the CVAA and FCC Part 79, which sets hour quotas instead: 87.5 hours a quarter for covered stations. Quotas get you some described content. Consent decrees get you all of it.
You can still see that difference in the catalogs. Netflix finished captioning its US library years ago under that 2012 decree. Audio description is nowhere close. One analysis of UK streaming catalogs found about a quarter of Netflix titles had description available, and the platforms doing better mostly get there by describing their own originals, not the licensed back catalog. The gap is closing, Netflix added more than 13,000 hours across 34 languages in 2025, but it started from a very different place than captions did.
There's no button. The CC symbol works worldwide without explanation. Audio description has no equivalent: no shared symbol, no agreed spot in the player, no reflex. Patent filings on AD detection have to work around this, proposing an "AD" mark by analogy to the CC symbol precisely because no such convention ever caught on. A feature nobody can find is a feature that doesn't exist.
Captions and audio share a screen. Descriptions and audio share a channel. Captions are additive. The text sits in the visual channel while the audio runs underneath, and you lose nothing by glancing down. Audio description is competitive. It's in the same channel as the dialogue it's describing, so something has to give.
That shows up in how the tools get built. Audio description needs dialogue-gap detection, because a description has to fit into a hole in the existing audio. Captions never needed it. Nobody had to solve that problem, because it doesn't exist for text.
Only one of those three is about difficulty. Captions didn't win because caption people tried harder. Captions won because captions were easier to make ambient.
The Text That Disappears Before You Read It
Your industry already solved this problem, for exactly one kind of on-screen text.
Caption reading speed has a real standard. Netflix caps adult subtitles at 17 characters per second, and every major broadcaster works to some version of the same limit. Go faster and most viewers can't keep up. Somebody measured how fast people read text on a screen, and an entire industry went along with it.
Now look at everything else on that screen. A lower third with a name and title. A budget chart with four data series. A slide with three bullets. A stat bug in the corner. None of it has a rule. It stays up for however long the editor felt like, and then it's gone.
Picture your own council presentation. The finance director puts up the budget breakdown, four categories, percentages next to each, on screen four seconds while she keeps talking. Nobody reads it. Not at home, not in the room. Sighted viewers miss it too. They just don't count it as missing, because the chart was technically right there.
We know exactly how long text has to stay up to be readable. We enforce that in one place and nowhere else. Audio description covers everywhere else.
4 Things Your Team Can Start Doing Now
The cost of audio description isn't production cost. It's retrofit cost.
Captions taught us the same lesson. Describing a meeting while you make it is a workflow question. Describing four hundred hours of archived meetings afterward is a budget question, and a painful one. Most teams we talk to don't have a spare developer sitting around waiting for that project. They've got one person who runs the meetings, edits, posts, and juggles three other jobs.
So the useful move isn't a big initiative. It's a few habits that cost nothing and start closing the gap now.
Ask presenters to say what's on the slide. "As you can see on this chart" becomes "Public Safety is 42 percent, Infrastructure is 28." That's most of what audio description does, delivered free by the person who already knows the numbers.
Ask speakers to identify themselves before they talk. Still the highest-value thing a meeting chair can do.
Hold your lower thirds for four seconds. Name-and-title graphics that flash for two don't register for anyone.
Before you publish, mute your own video and see what you lose. Then close your eyes and see what you lose. That second list is always longer than people expect. None of that needs a purchase order. All of it makes the eventual described version easier to produce.
Where This Goes
Public meetings belong to the people who can't make it there in person. That was the whole point of putting them online. A resident on the bus following the zoning discussion without staring at a screen. A commuter with the meeting playing in the car. Someone doing dishes with the council session on in the background. Roughly 7 million Americans live with vision impairment, about 1 million of them blind, and the audio economy already proved that plenty of sighted people will happily take in serious content without looking at it. So much for nobody would use this.
Here's what those four habits don't fix. Better habits improve the raw material. They don't produce a finished described track. Somebody still has to turn every council session, every hearing, every work session into described audio, every week, on top of the job they already have. That gap stays open no matter how good your lower thirds get.
That's what MediaScribe Narrate is for: automated audio description for prerecorded video. The AI drafts the narration, your staff reviews and approves, and nothing publishes until a person signs off. No scripting, no voice talent, no new headcount.
If you're a Title II public entity, you can apply to Access Granted, our program covering a year of automated audio description for qualifying organizations, with enough hours for a full year of public meetings. The application takes about two minutes.
FAQ
What is audio description, and how is it different from captions?
Captions put the audio into text on the screen, so a viewer can read what's being said. Audio description does the reverse: it puts what's on the screen into spoken audio, for viewers who can't see the visuals. Captions are additive, the text sits in the visual channel alongside the picture. Audio description is competitive, it shares the same audio channel as the dialogue, which is one reason it's harder to make ambient.
Why has audio description lagged when captions didn't?
Three structural reasons, none about anyone caring less. Its coverage came from hour quotas rather than a full-catalog consent decree, so it started from a much smaller base. It has no universal button or symbol the way captions have the CC mark, so viewers can't reliably find it. And it competes for the same audio channel as the dialogue, which makes the tooling harder to build. Captions won because they were easier to make ambient, not because anyone tried harder.
Do podcasts and smart speakers show there's real demand for audio-first content?
They show that plenty of people will happily take in serious content without looking at a screen. That retires the assumption that nobody would use audio description. It's a signal about listening habits, not a headcount of who needs description, the same way captions never won on the number of deaf viewers alone.
What can a small team do right now, at no cost?
Four habits. Ask presenters to read out what's on the slide instead of saying "as you can see here." Ask speakers to name themselves before they talk. Hold name-and-title graphics on screen for a full four seconds. And before publishing, watch your own video once muted and once with your eyes closed to notice what each version loses. None of it needs a purchase order, and all of it makes a future described version easier to produce.
How does MediaScribe fit in?
Habits improve the raw material, but someone still has to turn every meeting into a finished described track. That's what MediaScribe Narrate handles: automated audio description for prerecorded video, where AI drafts the narration and your staff reviews and approves before anything publishes. Qualifying Title II public entities can apply to Access Granted for a year of automated audio description. The application takes about two minutes.