The two words describe text that looks the same on screen. The difference is the reader the text was written for, and that decides what goes into it.
The difference between captions and subtitles is the audience
Captions assume the viewer can't hear the soundtrack. That could be someone deaf or hard of hearing, or someone watching on a train with the sound off. Everything the audio tells a hearing person has to be on screen: the words, who's saying them, and the sounds that matter to the story.
Subtitles assume the viewer hears everything but doesn't understand the language. They can hear the door slam and the music swell, and they can tell voices apart. So subtitles translate the dialogue and leave the rest to the ears.
Whether the text is burned in (open) or switchable (closed) is a separate question, covered in open vs closed captions. A caption can be open or closed, and so can a subtitle.
What captions include that subtitles leave out
A proper caption track carries three things a translation subtitle usually doesn't:
- Sound cues in square brackets, such as [phone buzzing], [laughter] or [door slams].
- Speaker identification, like "MAYA:" before a line, or a name when someone speaks off screen.
- Music notes: a ♪ symbol around sung lyrics, or a description like [tense music] when the score carries the mood.
Captions also stay close to what's said. They're in the same language as the audio and aim to be near verbatim. Subtitles can shorten and paraphrase, because the translator is fitting meaning into a line a reader can finish in a couple of seconds.
The same scene, written both ways
Here's one short moment from a made-up kitchen scene. The audio is in English. The subtitles are for a Spanish-speaking viewer.
| What happens | Closed captions (English) | Subtitles (Spanish) |
|---|---|---|
| A phone buzzes on the counter | [phone buzzing] | (nothing) |
| Maya, off screen, calls out | MAYA: Don't answer that! | ¡No contestes! |
| Sam answers anyway | SAM: Hello? ... Yeah, she's here. | ¿Hola? Sí, está aquí. |
| Music changes | [ominous music] | (nothing) |
| A glass breaks | [glass shatters] | (nothing) |
The Spanish viewer hears the phone, the music and the glass, so the subtitles skip them. A deaf viewer would miss the whole point of the scene without them.
What are SDH subtitles?
SDH stands for subtitles for the deaf and hard of hearing. They carry what captions carry (speaker names, sound cues, music notes) but are delivered as a subtitle track. You'll see "English (SDH)" in the language menu on many streaming services and Blu-ray discs.
In practice SDH and closed captions serve the same viewer. The difference is mostly technical and visual. Traditional US TV captions used a fixed broadcast format, often white text on a black box, while SDH looks like normal subtitles and can use the player's fonts and positioning. Netflix's English (USA) timed text style guide is a good public example of SDH rules, including square brackets for sound labels and speaker IDs.
Why the words get mixed up: US vs UK usage
In the US and Canada, "captions" means text for deaf and hard-of-hearing viewers and "subtitles" usually means translation. In the UK and much of Europe, "subtitles" covers both. The UK regulator Ofcom, for example, lists subtitles next to signing and audio description as TV access services for deaf and hard-of-hearing viewers, and its access services code defines subtitling as on-screen text representing speech and sound effects. An American would call that captioning. BBC iPlayer works the same way: its "Subtitles" switch turns on text written for deaf viewers. In British TV, "caption" more often means an on-screen label, like a name banner under an interviewee.
So if a British colleague asks for subtitles on an English video, they probably mean captions. Ask who it's for before you start.
Which one should creators make?
Start with captions in the language you speak in the video. That serves deaf and hard-of-hearing viewers, everyone watching on mute, and people who follow speech better when they can read along. For most YouTube, course and social videos, that's the job.
Add sound cues where a sound carries meaning: a notification ping in a tutorial, laughter in an interview, music in a vlog intro. You don't need to label every breath. Add speaker names when two or more people talk and it isn't obvious who's speaking.
Make translated subtitles only when your analytics show a real audience in another language. Translation is a second project, and it helps to have a clean caption file to translate from.
To get the caption file, run your video through our subtitle generator or the homepage auto captions tool. Both write a timed transcript of the speech in 50+ languages. They don't add sound cues, speaker names or translations for you, so type the [bracketed] cues and names into the editor, then export SRT or VTT. If you're unsure which file to pick, our explainer on SRT files covers the formats.