Compliance

Captions and Accessibility in Australia: WCAG, the DDA, and What You Actually Have to Do

Most Australian organisations know captions matter and are hazy on what “done properly” actually means. There is a published standard that answers that — the Web Content Accessibility Guidelines — and it is more specific and more achievable than people expect. This is a plain guide to what it asks for.

A note on what this covers

This describes WCAG, a published technical standard, and where it commonly shows up in Australian policy and procurement. We build transcription software and are not in a position to tell you what the law requires of your organisation — for that, talk to your own legal adviser.

Why WCAG is the standard people point to

WCAG is published by the World Wide Web Consortium and is the reference point almost everyone in Australia ends up using, because it is specific where broader accessibility principles are not. It sets out testable criteria at three conformance levels — A, AA and AAA — rather than leaving “accessible” open to interpretation.

The Australian Human Rights Commission has published guidance that points organisations to WCAG as the practical reference for accessible online content, which is a large part of why it has become the common benchmark here.

For a lot of organisations the question is settled by policy or contract rather than by first principles. Commonwealth government bodies have long worked to WCAG 2.0 AA, with WCAG 2.1 AA now the common baseline across government and increasingly written into procurement. If you sell to government, the standard often arrives through the contract.

The criteria that apply to audio and video

You do not need to read the whole of WCAG. For media, a handful of success criteria carry almost all of the weight.

CriterionLevelWhat it requires
1.2.1 Audio-only and Video-onlyAA transcript for audio-only content such as a podcast
1.2.2 Captions (Prerecorded)ACaptions for all prerecorded video with audio
1.2.3 Audio Description or AlternativeAA description or full text alternative for visual information
1.2.4 Captions (Live)AACaptions for live video with audio
1.2.5 Audio DescriptionAAAudio description of important visual content

For most organisations the practical reading is simple. Prerecorded video needs captions. Audio-only content needs a transcript. Both are Level A — the minimum conformance level, not an advanced target.

Live captioning at 1.2.4 is where cost rises sharply, and it is a genuine consideration for organisations that stream events. That is a different product from what we do; we are asynchronous only.

Captions and transcripts are not the same thing

These get used interchangeably and WCAG treats them as distinct.

Captions are time-synchronised text displayed with the video. They serve someone watching who cannot hear the audio, which means they need to convey not just dialogue but relevant non-speech sound — a doorbell, laughter, a phone ringing — where it matters to understanding.

Transcripts are a static text document of everything said. They serve someone who wants to read rather than watch, someone using a screen reader, and anyone who wants to search or quote the content. A transcript alone does not satisfy the captions criterion for video, and captions alone do not give you the searchable document.

The practical answer for most video is to produce both, because you get the transcript for free once you have transcribed the audio, and the caption file is a formatting step away.

Are auto-generated captions good enough?

This is the question that matters most in practice, and it deserves a careful answer rather than a marketing one.

WCAG does not set a numerical accuracy threshold. But captions that misrepresent what was said do not provide an equivalent experience, and the accessibility community's long-standing objection to unreviewed automatic captions — memorably nicknamed 'craptions' — is well founded. Captions full of errors can be worse than none, because they look like provision while leaving the viewer with a garbled account.

The realistic position is that modern AI transcription produces captions that are close, and a human skim closes the gap cheaply. Errors cluster predictably in proper nouns, technical terms, and passages with crosstalk or poor audio — which means a reviewer knows where to look.

A defensible workflow

Transcribe with AI, seed the prompt with names and terms that appear in the video, then have a person read the captions against the audio once before publishing. That is minutes of work per video rather than hours, and it is the difference between captions that genuinely serve someone and captions that only look like provision.

Where the standard usually shows up

Government. Commonwealth agencies and most state and territory bodies work to WCAG through published policy, and accessibility conformance is a routine line item in government procurement. If you sell to government, expect the standard to reach you through the contract.

Education. Universities and training providers generally hold themselves to WCAG for teaching material, and uncaptioned lecture recordings are a recurring source of student complaints. Institutions typically have their own accessibility policy that sets the expectation internally.

Health and essential services. Where a video conveys something a person needs in order to use a service or make a decision, treating captions as optional is a hard position to defend to your own users, whatever the formal position.

Everyone else. Nothing about WCAG is scoped to large organisations, and the criteria that matter for video are the cheapest ones to meet. Smaller teams tend to skip captions because nobody has asked, rather than because the standard treats them differently.

The part that is not about compliance

About one in six Australians has some degree of hearing loss, so the audience captions serve is large. But the majority of caption users are not deaf. They are watching without sound on public transport, in an open-plan office, or in bed. Social platforms autoplay muted, and captioned video demonstrably outperforms uncaptioned video on engagement as a result.

Captions also give search engines something to index. Video is otherwise opaque to search; a transcript on the page is the only way the content of a video becomes findable.

At $0.02 per audio minute, captioning a library of a hundred one-hour videos costs $120. The compliance argument is real, but the business case usually gets there on its own.

Where the audio goes

One point specific to Australian organisations: video containing identifiable people is personal information, and often sensitive information. Recorded lectures, client testimonials, internal training featuring staff, and consultations all qualify.

Sending that video overseas for captioning is a cross-border disclosure engaging APP 8. It is a slightly absurd outcome to create a privacy problem in the course of solving an accessibility one, but it happens routinely because captioning is treated as a production task rather than a data-handling one.

Frequently asked questions

Do I need to caption video in Australia?

Whether you are required to is a question for your own legal adviser, not for us. What we can tell you is the standard people measure against: WCAG treats captions for prerecorded video as a Level A criterion, which is its minimum conformance level. The Australian Human Rights Commission points organisations to WCAG for accessible online content, and government bodies generally work to it by policy, so it is the benchmark you will most often be asked about.

What WCAG level do I need for captions?

Captions for prerecorded video (1.2.2) and transcripts for audio-only content (1.2.1) are both Level A, the minimum conformance level rather than an advanced target. Live captioning (1.2.4) and audio description (1.2.5) are Level AA. Australian government policy commonly sets WCAG 2.1 AA as the baseline.

Is a transcript enough, or do I need captions too?

For video, you need captions — a transcript alone does not satisfy the captions criterion, because captions are time-synchronised with what is on screen. For audio-only content like a podcast, a transcript is what is required. Producing both for video is usually the right answer, since the transcript comes free with transcription and the caption file is a formatting step away.

Are auto-generated captions acceptable for accessibility?

WCAG sets no numerical accuracy threshold, but captions that misrepresent what was said do not provide an equivalent experience, and unreviewed automatic captions have a poor reputation for good reason. Modern AI output plus a human skim before publishing is a defensible workflow — errors cluster in proper nouns and difficult audio, so a reviewer knows where to look.

Does the standard treat small businesses differently?

No — WCAG is not scoped by organisation size, and the criteria that matter for video are among the cheapest to meet. Smaller teams usually skip captions because nobody has asked rather than because the standard says they can. At cents per video, it is rarely worth the saving.

Does sending video overseas for captioning create a privacy issue?

It can. Video containing identifiable people is personal information, and recorded consultations, lectures and internal training often contain sensitive information. Sending it to an overseas captioning service is a cross-border disclosure engaging APP 8 under the Privacy Act. Processing in Australia avoids creating a privacy problem while solving an accessibility one.

Caption your library without sending it overseas

90 minutes free, no credit card required. Timestamped, speaker-labelled output processed entirely in AWS Sydney.