Skilled lipreaders average just 50.7% of words from lips alone (PMC, 2023). Here's why pairing lipreading with in-view captions — one field of view, eyes on the face — beats either channel by itself.
By Madhav Lavakare · Published 2026-07-23 · 18 min read
Guides

Madhav Lavakare
·
July 23, 2026
·
18 min read

On this page
Table of Contents
▼
Editorial disclosure: AirCaps makes captioning smart glasses that display real-time captions inside your field of view, which is directly relevant to this article's argument that captions and lipreading work best together. We reference AirCaps specifications where they bear on the discussion. Every statistic is independently sourced and linked inline. Where lipreading, an audiologist, or a hearing aid remains the right tool, we say so. This is a guide to how two visual channels reinforce each other, not a claim that captions replace the work your eyes and ears already do.
Skilled lipreaders watching a naturalistic story catch an average of just 50.7% of the words, with individual scores ranging from 6% to 100% (PMC, 2023). Lipreading alone leaves half the message on the table. Captions alone, when they live on a phone in your lap, pull your eyes off the speaker's face and throw away a visual signal worth as much as 15 dB of clarity in noise (Sumby & Pollack, 1954). Combine both — read the lips and read the captions in the same glance — and each channel covers the other's blind spots.
That combination is the whole reason captioning glasses exist as a category distinct from phone apps. When captions render on a lens instead of a screen in your hand, you never look down. Your eyes stay on the face, the lips stay in play, and the text fills the gaps the lips can't. This guide explains the science behind why that pairing works, where each channel fails on its own, and what to look for in a device built to deliver both at once.
Key Takeaways
- Even skilled lipreaders average only 50.7% of words from a spoken narrative, with scores spanning 6% to 100% (PMC, 2023), because roughly two-thirds of English speech sounds are not distinguishable on the lips
- Seeing the talker's face is worth up to 15 dB of effective signal-to-noise improvement, with the largest gains in the noisiest rooms (Sumby & Pollack, 1954)
- At a punishing minus 16 dB signal-to-noise ratio, listeners jumped from 9% words correct with audio alone to 38% when they could also see the face (PMC, 2019)
- The brain fuses lips and sound automatically — the McGurk effect alters what up to 98% of adults perceive (McGurk & MacDonald, 1976) — so keeping the face in view is neurological, not optional
- Phone captions force a look-down that abandons the lips; captioning glasses keep both visual channels in one field of view
- AirCaps delivers 97% caption accuracy at 300ms latency using 4-microphone beamforming, on a 49-gram binocular MicroLED display, for $599 (HSA/FSA eligible, no required subscription)
Lipreading is a remarkable skill and a limited one. In a controlled study of adults lipreading a naturalistic narrative, the mean score was 50.7% of words correct, and the spread ran from 6% all the way to 100% (PMC, 2023). On isolated monosyllabic words the numbers are worse — human lipreaders scored around 32% (arXiv, 2018). So even a talented speechreader is guessing at a large fraction of what's said.
The reason is baked into the anatomy of speech. English has far more phonemes than visemes — the distinct mouth shapes you can actually see. Voiced and unvoiced pairs like "p," "b," and "m" are produced with nearly identical lip movements, so they look the same on the face. These homophenes make whole clusters of words visually indistinguishable. The figure commonly cited in hearing-health literature is that only about 30% of English speech sounds can be told apart by sight alone (Hearing Health & Technology Matters, 2019). Whether the exact number is 30% or a little higher, the mechanism is solid and well documented in the audiology literature (American Journal of Audiology, 2022).

Then there's the cost. Lipreading is cognitively expensive. In a study of 200 deaf and hard-of-hearing adults, heavier reliance on speechreading correlated with more listening-related fatigue (eScholarship, 2021). Anyone who has spent an evening straining to read a friend's lips across a loud table knows the feeling — you leave exhausted, and you still missed things. Lipreading is a real channel, but it was never meant to carry the whole message by itself.
Citation capsule: Skilled adults lipreading a naturalistic narrative averaged 50.7% of words correct, with a range of 6% to 100% (PMC, 2023). On isolated words, human lipreaders score around 32% (arXiv, 2018). Homophenes — sounds like "p," "b," and "m" that look identical on the lips — are why the visual channel alone leaves so much ambiguous.
Because the brain treats lips and sound as one signal, not two. In the foundational experiment, Sumby and Pollack found that letting listeners see the talker's face raised speech intelligibility in noise so much that the visual contribution was worth roughly 15 dB of signal-to-noise ratio, with the biggest gains exactly where the noise was worst (Sumby & Pollack, 1954). Seventy years of research has only reinforced it.
A modern replication makes the size of the effect concrete. At a brutal minus 16 dB signal-to-noise ratio — the kind of chaos you'd find in a packed bar — listeners understood 9% of words with audio alone. Let them see the face too, and they jumped to 38%, a gain of nearly 30 points (PMC, 2019). Meta-analytic estimates put the typical benefit at 10 to 15 dB of effective threshold improvement from adding visual speech cues (PMC, 2022).
The clincher is that this integration is automatic. In the McGurk effect, a listener hearing the sound "ba" while watching a face mouth "ga" perceives a third sound entirely — "da" — and this illusion reshapes perception in up to 98% of adults (McGurk & MacDonald, 1976). You can't turn it off. The face isn't a nice-to-have supplement; your auditory cortex is wired to weld what it sees to what it hears. Take the face away and you've amputated part of the speech signal.
Captions solve the exact problem lipreading can't: ambiguity. A caption disambiguates "pat" from "bat" from "mat" instantly, because it spells the word out. High-accuracy speech-to-text turns a 50% guess into near-certainty. So why isn't perfect captioning the entire answer? Because captions on their own throw away everything the face was carrying — tone, timing, who's about to speak, the prosody that tells you a question from a statement.
There's also a placement problem that has nothing to do with accuracy. When captions live on a separate screen, reading them means looking away from the person. You gain the text and lose the lips, the eye contact, and the automatic audiovisual fusion described above. In noisy rooms, that's a bad trade. The average mainstream restaurant runs at 78 dBA, with many exceeding 80 dBA, well past the roughly 70 dBA where speech understanding starts to break down for people with hearing loss (Noise & Health, 2014; NIDCD, 2024). Those are precisely the rooms where you need the visual channel most — and precisely where looking down at a phone costs you the most.

The mask era proved the point in reverse. When surgical masks hid people's mouths during COVID-19, speech intelligibility fell 12% to 16% and high-frequency speech was attenuated by 3 to 12 dB — an effect that compounded because listeners also lost the ability to read lips (PLOS One, 2021). Hiding the lips measurably degraded understanding. Captions that make you hide the lips yourself recreate that loss voluntarily.
Citation capsule: Covering the mouth with a surgical mask cut speech intelligibility by 12% to 16% and attenuated high-frequency speech by 3 to 12 dB, with the damage compounded by the loss of lipreading cues (PLOS One, 2021). Any caption system that forces the user to look away from the speaker's mouth reproduces the same visual loss on purpose.
Here's the insight that gets lost in most comparisons of caption apps: the cost of phone captioning isn't the reading, it's the looking down. Every glance at a screen in your hand is a glance away from the face — and the face was carrying a visual signal worth 3 to 15 dB of effective clarity in noise (Cognitive Research, 2022; Sumby & Pollack, 1954). You trade one visual channel for another instead of stacking them.
How much is that visual signal actually worth? The estimates converge on a large number. Studies of the face-visible benefit to the speech reception threshold cluster around 3 to 5 dB in everyday conditions, Summerfield's classic estimate put lipreading's contribution at roughly 4 to 6 dB, and the Sumby and Pollack ceiling reached up to 15 dB in the noisiest conditions (American Journal of Audiology, 2022). To put those numbers in perspective, each single decibel of signal-to-noise improvement is worth roughly 10% to 15% in speech recognition (American Journal of Audiology, 2022). Even the conservative end of the range is a meaningful chunk of a conversation.
Now flip it. Imagine a caption that appears next to the face instead of below the table — text and lips inside a single field of view, both in focus at once. You don't choose between them. The captions resolve the homophenes the lips can't distinguish, and the lips supply the timing, emotion, and turn-taking cues the text can't. That's not a hypothetical. It's the design premise of captions rendered on the lens rather than on a screen you have to hold.
Captioning glasses solve the look-down problem by geometry. The text appears on a transparent display roughly where the speaker's face already is, so you read the captions and watch the lips in the same glance. AirCaps runs this on a binocular MicroLED display — one image per eye — which reduces the eye strain that plagues single-display alternatives and matters a great deal when you're wearing the device through a long dinner or a full workday.
Latency is the spec that decides whether the two channels feel like one. If captions lag too far behind the lips, the brain notices the mismatch and the illusion of a single synchronized signal collapses. Human viewers begin to detect audiovisual desync when the offset stretches past roughly 125 to 200 milliseconds (arXiv, 2022). AirCaps renders captions at 300 milliseconds end-to-end — close enough that the text arrives while the mouth is still moving, so the caption reads as part of the same event rather than a delayed subtitle. A 4-microphone beamforming array isolates whoever you're facing and filters the 78 dBA restaurant around you, delivering 97% caption accuracy in exactly the noise where lipreading and hearing both struggle.

The privacy dimension matters too. AirCaps' binocular display leaks less than 2% of light from the front, so the person across from you doesn't see text glowing in your lenses — you keep eye contact, they see your eyes, and nobody's conversation turns into a demonstration. The glasses weigh 49 grams, lighter than many regular eyeglasses, and fit any prescription from minus 16 to plus 16 diopters through any optician. For multilingual settings, automatic language detection across 60-plus languages means the captions appear in whatever language the speaker uses, without you touching a menu.
Citation capsule: Viewers start to notice audiovisual desynchronization once the offset passes roughly 125 to 200 milliseconds (arXiv, 2022). AirCaps captions render at 300 milliseconds end-to-end with 97% accuracy, using a 4-microphone beamforming array to isolate the speaker in an average 78 dBA restaurant — close enough to lip movement that the text and the lips read as a single synchronized signal.
The clearest way to see the argument is side by side. Each channel has a distinct failure mode, and those failure modes barely overlap — which is exactly why stacking them works. Lipreading fails on ambiguity; captions fail on context and placement. Together they cover both.
| Dimension | Lipreading Alone | Captions Alone (on a phone) | Both Together (in-view captions) |
|---|---|---|---|
| Word-level accuracy | ~50% of words; worse on isolated words | High, up to 97% with good speech-to-text | Highest; text resolves what lips can't |
| Homophenes ("p" vs "b" vs "m") | Indistinguishable | Fully disambiguated | Disambiguated by the text |
| Tone, emotion, turn-taking | Strong; the face carries it | Lost; text is flat | Preserved; eyes stay on the face |
| Eyes on the speaker | Yes, required | No; you look down | Yes; captions sit near the face |
| Performance in 78 dBA noise | Helps, but incomplete | Helps, if you sacrifice the lips | Both visual channels stack |
| Cognitive load / fatigue | High; effortful guessing | Moderate; constant look-down | Lower; redundancy reduces strain |
| Audiovisual fusion (McGurk) | Engaged | Broken by looking away | Engaged; face stays in view |
Read down the last column and the pattern is obvious. The combined approach doesn't win on any single row by a landslide — it wins because it never loses a row badly. Lipreading gives up on ambiguous sounds. Phone captions give up on the face. In-view captions are the only option that keeps every row in the "works" column at the same time.
If you already lipread, you are not doing it wrong — you're doing heroic work with an incomplete signal. Roughly 1.5 billion people live with some degree of hearing loss worldwide, a figure projected to reach 2.5 billion by 2050 (WHO, 2025), and in the US about 48 million people, or 1 in 7, have some hearing loss (HLAA, 2023). Many of them lean on lipreading daily, filling the gaps by hand and paying for it in fatigue.
The practical takeaway is to stop asking your lips to do all the work. Add a second visual channel that resolves the ambiguity, and keep it in the same field of view so you don't sacrifice the lips to get it. That's the case for captioning glasses over phone-based apps for anyone whose primary strategy is watching faces. In professional settings the argument compounds: speaker identification labels who's talking in a group, and a searchable transcript means you don't have to hold the whole meeting in working memory while you lipread it in real time.
None of this replaces an audiologist, a hearing aid, or a cochlear implant where those are the right tools. Hearing devices restore audibility; captions and lipreading restore intelligibility when audibility isn't enough. For a lot of people, the honest answer is all of the above — the hearing aid for the sound, the lips for the tone, and the captions for the words the first two can't quite pin down. The point of this article is narrower and, I think, under-appreciated: of the two visual channels, you should never have to pick just one.
Even skilled lipreaders average about 50.7% of words when reading a naturalistic spoken narrative, with individual scores ranging from 6% to 100% (PMC, 2023). On isolated words, humans score closer to 32% (arXiv, 2018). Homophenes — sounds like "p," "b," and "m" that look identical on the lips — are why so much stays ambiguous without a second channel like captions.
Yes, dramatically. Seeing the talker's face is worth up to 15 dB of effective signal-to-noise improvement, with the largest gains in the noisiest conditions (Sumby & Pollack, 1954). In one modern test, listeners went from 9% words correct with audio alone to 38% when they could also see the face at a minus 16 dB signal-to-noise ratio (PMC, 2019).
Phone captions work, but reading them means looking down and away from the speaker's face. That sacrifices the visual speech signal worth 3 to 15 dB of clarity in noise (American Journal of Audiology, 2022). Captioning glasses put the text in your field of view near the face, so you keep the lips and add the captions instead of trading one for the other.
The McGurk effect is proof that your brain fuses lips and sound automatically. Hearing "ba" while watching a face mouth "ga" makes most people perceive "da" — an illusion that alters perception in up to 98% of adults (McGurk & MacDonald, 1976). You can't switch it off, which is why keeping the speaker's face in view is neurologically valuable, not just convenient.
Viewers start noticing audiovisual desynchronization once the offset passes roughly 125 to 200 milliseconds (arXiv, 2022). AirCaps renders captions at 300 milliseconds end-to-end, close enough that the text arrives while the mouth is still moving. That keeps the caption and the lips reading as one synchronized event rather than a lagging subtitle.
No. They supplement both. A hearing aid restores audibility, lipreading carries tone and timing, and captions resolve the words the other two leave ambiguous. AirCaps captioning glasses deliver 97% accuracy at 300ms latency using 4-microphone beamforming, weigh 49 grams, and cost $599 (HSA/FSA eligible, no required subscription). For most people the strongest setup combines all of the tools rather than choosing one.
Sources: PMC — Lipreading a Naturalistic Narrative, 2023. arXiv — Lipreading Human Baseline, 2018. Sumby & Pollack — Visual Contribution to Speech Intelligibility in Noise, JASA, 1954. PMC — Visual Speech Benefit in Clear and Degraded Speech, 2019. PMC — Multisensory Benefits for Speech Recognition in Noisy Environments, 2022. American Journal of Audiology — Lipreading: A Review, 2022. McGurk & MacDonald — Hearing Lips and Seeing Voices, Nature, 1976. Cognitive Research — Communication With Face Masks During COVID-19, 2022. PLOS One — Influence of Surgical and N95 Masks on Speech Perception, 2021. Hearing Health & Technology Matters — Giving Good Lip for Better Speechreading, 2019. Noise & Health — Noise in Restaurants, 2014. NIDCD — Noise-Induced Hearing Loss, 2024. eScholarship — Listening-Related Fatigue and Cognitive Effort in Deaf and HoH Bilinguals, 2021. arXiv — Audiovisual Desynchronization Perception, 2022. WHO — Deafness and Hearing Loss Fact Sheet, 2025. HLAA — Hearing Loss by the Numbers, 2023. Image credits: Pexels (royalty-free).
On this page
Table of Contents
▼
Written by

Madhav Lavakare
Co-founder & CEO, AirCaps
Co-founder of AirCaps. Building AI-powered smart glasses for conversation since 2013. Yale graduate, Y Combinator alum. Built his first Google Glass apps at age 13 and has spent 11+ years in speech AI and wearable computing.
Related Articles

Guides
Captioning Glasses for Cochlear Implant Users: A Practical Companion Guide
More than 118,000 US adults wear a cochlear implant (NIDCD, 2024). How captioning smart glasses complement an implant in restaurants, meetings, and family dinners — and where they don't.

Madhav Lavakare
·
Jun 22, 2026
·
24 min read

Guides
Auditory Processing Disorder (APD) and Captioning Glasses: Hearing Without Hearing Loss
APD affects 0.2-6.2% of school-age children who have normal audiograms but can't process speech in noise. See how captioning glasses bypass the ear entirely.

Nirbhay Narang
·
Jul 17, 2026
·
20 min read

Guides
Captioning Glasses for Tinnitus: Why Reducing Listening Effort Quiets the Ringing
14.4% of adults worldwide live with tinnitus (JAMA Neurology, 2022). Why captioning glasses that cut listening effort can lower tinnitus distress when straining to hear makes the ringing louder.

Nirbhay Narang
·
Jul 4, 2026
·
19 min read
© 2025 AirCaps. All rights reserved.