WCAG 1.2.2 Captions (Prerecorded)
Success Criterion 1.2.2 requires captions for all prerecorded audio content within synchronized media (a video with sound). It's a Level A criterion under Guideline 1.2 (Time-based Media).
What counts as a caption, specifically
Captions are a synchronized, on-screen text version of everything spoken and every meaningful sound effect (a doorbell, a phone ringing, dramatic music) — not just dialogue. This is what distinguishes captions from subtitles: subtitles typically only translate dialogue and assume the viewer can already hear other audio cues, while captions are built for viewers who can't hear any of it.
Auto-generated captions aren't automatically sufficient
Platform auto-captioning (YouTube's automatic captions, for instance) is a reasonable starting point but frequently contains real errors — misheard words, missed punctuation, garbled proper nouns and technical terms. WCAG doesn't specify a required accuracy threshold in the text of 1.2.2 itself, but industry practice and most accessibility audits treat unreviewed auto-captions as a genuine risk area: a caption that says the wrong thing can be worse than no caption, since a deaf viewer has no way to tell it's wrong without independently verifying against another source. Reviewing and correcting auto-generated captions before publishing is the practical bar this criterion is understood to require.
Where this fits alongside other media criteria
1.2.2 is one of several time-based-media criteria that often get satisfied together as part of the same production step: captioning software or a manual transcript pass typically produces the caption file needed here, the transcript content needed for 1.2.1-style alternatives, and much of what's needed for audio description under 1.2.3/1.2.5 (AA). Treating video accessibility as one combined production step — captions, transcript, and audio description together — is usually more efficient than tackling each criterion in isolation after the fact.
Official references
Common questions
- What is WCAG 1.2.2?
- A Level A criterion requiring captions for prerecorded video with audio, covering spoken dialogue and significant sound effects synchronized with the video.
- How do captions and subtitles differ?
- Captions include non-speech audio (sound effects, speaker identification) for accessibility; subtitles typically assume the viewer can hear and only translate dialogue. WCAG requires captions.
- Is a transcript enough for 1.2.2?
- No — 1.2.2 requires synchronized captions on the video itself; a separate transcript alone doesn't satisfy it, though it may help meet other criteria.
Want to see how your own site scores?
Run a free accessibility scan