Two YouTube videos, both with transcripts available. One reads like an article. The other is a single sentence four thousand words long with no capital letters after the first one. The difference is not the tool you used to open them — it is who wrote the captions.
This matters more than it sounds, because the transcript is increasingly the way people consume long video: skimmed, searched, quoted, read aloud on a walk. A transcript that falls apart is a video you quietly give up on.
The two kinds of caption track
Human-written captions are uploaded by the creator, or by whoever they pay. The text has punctuation. Sentences end. When the speaker changes, the caption says so. Numbers are written the way the speaker meant them, and proper nouns are spelled correctly because a person who understood the content typed them.
Automatic captions are produced by speech recognition after the video is published. They are free, they cover an enormous amount of content that would otherwise have nothing, and they are genuinely impressive as engineering. They also have two structural gaps that no amount of accuracy improvement closes:
- No punctuation. Not "poor punctuation" — none. The recogniser emits a stream of words.
- No speaker labels. Three people in a conversation produce one undifferentiated stream.
Everything that follows comes from those two facts.
Why a lecture survives and a panel does not
A single presenter talking through slides is the friendly case. One voice, so the missing speaker labels cost nothing. Long complete sentences with clear pauses, so even without punctuation your ear inserts the breaks. Technical vocabulary is usually repeated often enough that a misrecognition in one sentence is corrected by context in the next. An automatic transcript of a conference talk is genuinely usable.
A four-person panel is the hostile case, and it fails on both counts at once. Nobody is labelled, so when the transcript says "well I disagree with that entirely" you have no idea who said it — which is exactly the information a panel exists to convey. And people interrupt, so the stream interleaves half-sentences from two speakers into something that parses as neither.
Between those extremes, a rough ordering of what to expect:
| Video type | Automatic captions |
|---|---|
| Lecture, single presenter | Usually fine |
| Tutorial with screen recording | Fine, but references to "this" and "here" lose their meaning without the picture |
| Two-person interview | Workable if they take turns; confusing when they overlap |
| Panel discussion | Rough |
| Podcast with music beds and banter | Rough |
| Anything with heavy accents plus jargon | Unpredictable — check a paragraph before committing an hour |
What this means if you are listening rather than reading
Reading an unpunctuated transcript is annoying. Listening to one read aloud is worse, and it is worth understanding why: a text-to-speech voice uses punctuation to decide where to pause and how to shape intonation. Feed it a stream with no full stops and it produces a flat, breathless delivery that never resolves. Your brain gets no cues about where one thought ends and the next begins.
So the same transcript that is merely tiring to read can be genuinely unlistenable. If you plan to listen rather than skim, the human-versus-automatic question matters roughly twice as much.
This is why CastReader always prefers a human-written caption track when the video has one — in your reading language first, then English — and only falls back to automatic captions when there is no alternative. It is not a quality setting you toggle; it is the order the track gets chosen in, because the difference in outcome is that large.
Checking before you commit
The transcript panel does not label tracks as human or automatic in an obvious way, but you rarely need it to. Open the transcript and read three lines.
Punctuation present, sentences ending, occasional speaker names — a human wrote it. Proceed.
A stream of lowercase words with no full stops — automatic. Now judge it against the table above: if it is one person talking, carry on; if you can see the conversation jumping between people with nothing marking the change, consider whether you actually need this particular video.
Ten seconds of checking saves an hour of fighting a transcript that was never going to work.
When neither exists
Some videos have no caption track at all — the uploader disabled captions, the video is too recent for automatic generation to have finished, the spoken language is outside YouTube's automatic coverage, or the audio defeated the recogniser. No caption-based tool can help there, and it is better to know that in ten seconds than after ten minutes of trying different apps. We wrote up how to confirm it and what is left to try separately.
The short version
Automatic captions are a gift for search and a mixed bag for reading. Their limits are structural rather than incidental, so they will not improve their way out of the panel-discussion problem. Check three lines before you invest an hour, prefer human tracks when they exist, and be more careful when you intend to listen than when you intend to skim.