Formats
Subtitle formats and video basics: SRT, VTT, ASS, containers, codecs and burn-in
Updated 5 October 2026 · ~10 min read
Most subtitle frustration is not about transcription accuracy. It is about not knowing that an MP4 is not a codec, that a “subtitle” is not part of the picture until you make it so, and that making it so costs you a generation of video quality. This is the vocabulary, with the gotchas attached.
Containers versus codecs: the distinction everything else rests on
A codec is a compression scheme for one kind of data. H.264/AVC and H.265/HEVC compress pictures; AAC and Opus compress audio.[1] A codec defines how to turn samples into a compact bitstream and back.
A container is the file format that holds those bitstreams together and records how they relate. MP4 — formally ISO/IEC 14496-14, derived from the ISO base media file format — stores video, audio, subtitle and metadata tracks as a tree of boxes, with tables saying which chunk belongs to which track and at what timestamp each sample should be presented.[2] MKV, WebM and MOV are alternative containers solving the same problem with different trade-offs; MKV is the most permissive about what it will carry.
This is why “my MP4 doesn’t play” is almost never about the MP4. The container parses fine; the codec inside is one the player cannot decode. It is also why remuxing (repackaging streams into a different container) is instant and lossless, while transcoding (decoding and re-compressing) is slow and lossy. Knowing which operation you are asking for predicts both the time it takes and the quality you lose.
Two more terms worth having:
- Muxing (multiplexing) interleaves separate streams into one container so a player can read them in step without seeking back and forth. Demuxing is the reverse, and it is the first thing any subtitle pipeline does — see stage 1 of the AI subtitling pipeline.
- Keyframes (I-frames) are frames coded without reference to any other. The frames between them store only differences, which is where most of the compression comes from, and it is why you can only cut cleanly at a keyframe without re-encoding.
Frame rates, time bases and the 23.976 problem
Video is a sequence of frames shown at a frame rate. Common ones are 24 fps (cinema), 25 fps (PAL regions), 30 and 60 fps (screen capture and sport), and the two odd ones: 23.976 and 29.97 fps.
Those two exist for a historical reason. When colour was added to the NTSC television standard, the frame rate was pulled down by a factor of 1000/1001 to keep the new colour subcarrier from interfering with the audio carrier.[3] Seventy years later, every file in the North American production chain still carries the consequence.
The 0.1% difference between 24 and 23.976 sounds negligible. Over a two-hour feature it accumulates to about 4.3 seconds — which is the single most common cause of subtitles that start in sync and end hopelessly late. The diagnostic is simple and worth memorising:
- Constant offset (wrong from the first cue, equally wrong at the end) → shift every timestamp by the same amount.
- Growing drift (fine at the start, progressively worse) → frame-rate mismatch. Multiply every timestamp by the ratio of the two rates. Nudging will never fix this.
Subtitle formats themselves store wall-clock times, not frame numbers, so they are nominally frame-rate agnostic. The mismatch comes from the timings having been authored against a differently-rated copy of the video. Separately, a container has its own time base — the unit its timestamps are counted in — and rounding when converting between time bases is why a cue can land a frame off even when everything else is right.
Soft, muxed and hard subtitles
There are three places subtitle text can live, and they behave very differently:
| Form | Where the text lives | Properties |
|---|---|---|
| Sidecar | A separate .srt/.vtt file next to the video |
Editable in any text editor, switchable, zero quality cost, kilobytes. Can be lost or renamed apart from the video; must be delivered alongside it. |
| Muxed (soft) track | A subtitle track inside the container | Travels with the file, multiple languages selectable, still lossless. Player support for the specific subtitle codec varies; some platforms discard the track on re-upload. |
| Hardsubs (burned in) | Drawn into the pixels; the video is re-encoded | Displays on literally anything, survives any platform, works in silent autoplay feeds. Irreversible, not switchable, not selectable as text, and costs one lossy generation. |
The practical rule: keep the sidecar as your master and burn in only as a final export for a specific destination. Hardsubs are a rendering decision, not a storage format. Once you have thrown away the text, correcting a typo means re-encoding the whole video again — a second lossy generation on top of the first.
SubRip (.srt): the lingua franca
SRT came out of the SubRip DVD-ripping program in the late 1990s and became ubiquitous by being trivially simple. There is no formal specification; the format is whatever the ecosystem agrees to parse.[4]
1
00:00:01,000 --> 00:00:03,500
The first line of dialogue.
2
00:00:03,600 --> 00:00:06,120
A second cue, split
across two lines.
A cue is: a sequence number, a timecode line with -->, one or more text lines, a blank line.
Notable details:
- The decimal separator in the timecode is a comma. This is the main thing distinguishing SRT timecodes from WebVTT ones, and the usual cause of a “corrupt file” error after a careless rename.
- Timecodes are
HH:MM:SS,mmm— hours are mandatory, milliseconds are three digits. - Styling is limited to a handful of HTML-ish tags (
<i>,<b>,<u>, sometimes<font>) that players support inconsistently. There is no positioning. - Encoding is the classic trap. SRT carries no encoding declaration, so a file saved as Windows-1252 and read as UTF-8 produces mojibake on every accented character. Always save SRT as UTF-8.
- Cue numbers are conventionally sequential from 1, and most parsers tolerate gaps — but not all, so renumber after editing if a player complains.
Choose SRT when you want the file to work in the largest number of places with the least thought.
WebVTT (.vtt): the web standard
WebVTT is a W3C specification, and it is the format HTML natively understands: a
<track> element inside a <video> points at a .vtt file and the
browser renders the cues itself, exposing them to scripts through the TextTrack API.[5][6]
WEBVTT
intro
00:00:01.000 --> 00:00:03.500
The first line of dialogue.
00:00:03.600 --> 00:00:06.120 line:90% align:center
<i>A styled, positioned cue.</i>
Differences from SRT that matter in practice:
- The file must begin with
WEBVTT. Omit it and conforming parsers reject the whole file. - The decimal separator is a period, not a comma. Converting SRT to VTT is mostly adding the header and swapping those commas.
- Cue identifiers are optional and may be any string, not just a number.
- Cue settings control position, alignment, line and size — genuinely useful for moving a caption off burned-in on-screen text.
- Cues can be styled with CSS via the
::cuepseudo-element, which is how a web player restyles captions without touching the file. - Encoding is fixed: WebVTT files are UTF-8. No ambiguity, by specification.
- Comments (
NOTE), chapter tracks and metadata tracks are part of the format.
Choose WebVTT when the target is a browser, an HTML5 player, or HLS streaming.
ASS, and the broadcast formats
Advanced SubStation Alpha (.ass) is a different class of format: it carries named
styles, fonts, colours, arbitrary positioning, rotation, animated transforms and karaoke timing, with a
scripting syntax inside the cue text.[7] It grew up in anime fansubbing, where typesetting
signage and song lyrics mattered, and it is the format libass renders — the library most players and
ffmpeg use for subtitle drawing, including when burning in.[8] Use it when you need
typographic control; accept that fewer platforms ingest it.
Beyond these sit the professional formats you will meet only at delivery: TTML and its profiles (IMSC, SMPTE-TT), XML-based and the basis of most modern broadcast and streaming exchange;[9] EBU-STL, the long-standing European broadcast interchange format; and CEA-608/708, closed captions encoded into the video signal itself rather than carried as a separate text track. If a distributor asks for “IMSC 1.1”, that is a TTML profile, and you convert rather than author.
Which one to export
- Uploading to a video platform, or unsure →
.srt, UTF-8. - Your own website or HTML5 player →
.vtt. - Archiving alongside a master →
.srtor.vtt, kept as a separate file. Never only hardsubs. - Social feeds, silent autoplay, or a platform that strips subtitle tracks → burn in, but keep the sidecar too.
- Heavy typesetting or karaoke →
.ass. - A broadcaster or streaming service asked for a specific format → give them exactly that; convert from your SRT/VTT master.
Reading-speed rules that make subtitles legible
Correct text is not the same as readable subtitles. The numbers below come from published professional guidelines and are worth applying even to casual work — they are the difference between subtitles people can follow and subtitles people switch off.
| Constraint | Typical rule | Why |
|---|---|---|
| Reading speed | Up to ~20 characters/second for adults; ~17 CPS for children[10]; the BBC works to around 160–180 wpm[11] | Above this, viewers stop watching the picture and still miss text |
| Line length | ≤ 42 characters per line[10] | Long lines force horizontal eye travel and cover the frame |
| Lines per cue | Maximum 2 | Three lines obscure too much picture and invert reading order |
| Minimum duration | ~5/6 second (20 frames at 24 fps)[10] | Shorter and the cue flashes before it can be fixated |
| Maximum duration | ~7 seconds[10] | A cue lingering past its dialogue reads as a fault |
| Gap between cues | At least 2 frames | Without a gap, consecutive cues look like one flickering block |
| Line breaks | At clause boundaries; never between article and noun[11] | Syntactic breaks are read in one saccade; arbitrary ones are not |
| Shot changes | Avoid cues spanning a cut[11] | A cut makes the eye re-read the text it had already started |
The consequence professionals accept and amateurs resist: at these limits you often cannot fit every word. Condensing dense speech — cutting filler, compressing phrasing, keeping meaning — is a core subtitling skill, and it is the one thing no automatic system currently does well.[12] Automatic transcription is verbatim by nature, so the single highest-value edit you can make to generated subtitles is usually deletion.
What burn-in actually costs
Burning in means decoding every frame, drawing text onto it, and encoding the result. Three consequences follow, and all three surprise people:
1. A lossy generation
The source video was already compressed with loss. Re-encoding compresses the decoded result again, and the artefacts of the first pass are now treated as detail worth preserving, which wastes bitrate. With H.264 this is controlled by the constant rate factor: lower CRF means higher quality and a larger file, and the FFmpeg project’s guidance is that CRF 18–23 is the sane range for x264, with roughly ±6 halving or doubling the bitrate.[13] Picking a conservative CRF keeps the second generation visually close to the first; picking a careless one is where “why does it look soft now?” comes from.
2. Real compute time
Transcoding is frame-by-frame work on every pixel, so burn-in time scales with resolution and length, not with how much text there is. In the browser, where ffmpeg runs compiled to WebAssembly,[14] expect it to be slower than a native encoder — a reasonable trade for never uploading the file.
3. Irreversibility
Pixels are pixels. There is no un-burn. The text can no longer be searched, selected, translated, restyled or switched off, and a screen reader cannot reach it. For accessibility specifically, soft subtitles are the better artefact in every respect.
The workflow that avoids all three: transcribe, edit, export the sidecar, keep it. Burn in last, from the original video rather than from a previously burned copy, once per destination that needs it. This tool does exactly that — subtitle export and burn-in are separate buttons, and burn-in always works from your original file.
Subtitles versus captions — not a synonym
In North American usage the distinction is substantive, not pedantic. Subtitles render dialogue for someone who can hear the audio but does not understand the language. Closed captions serve viewers who cannot hear it, so they additionally convey speaker identification and non-speech audio — [door slams], [ominous music] — because that information is otherwise lost entirely.[15]
Automatic speech recognition produces subtitles, not captions: it transcribes speech and is blind to everything else, and it does not know who is speaking (see what Whisper cannot do). If the deliverable is an accessibility caption track, the sound-effect cues and speaker labels are human additions — which is fine, because both are easy to add to an SRT by hand.
Troubleshooting, in order of likelihood
- Garbled accented characters → encoding. Re-save as UTF-8.
- Player ignores the file entirely → missing
WEBVTTheader in a.vtt, or an.srtrenamed to.vttwith commas still in the timecodes. - Constant offset → shift all timestamps equally.
- Growing drift → frame-rate mismatch; rescale all timestamps by the frame-rate ratio.
- Subtitles missing after upload → the platform dropped the muxed track. Supply the sidecar through its caption-upload field, or burn in.
- Only the first cue shows → a malformed timecode line or missing blank line between cues stopped the parser. Validate the file.
- Burned-in text clipped at the edge → no margin set in the render style, or a cue longer than the safe area. Shorten the line or set a margin.
- Video looks softer after burn-in → the CRF was too high. Re-encode from the original with a lower one.
References
- Wikipedia. Advanced Video Coding (H.264), HEVC (H.265) and Advanced Audio Coding.
- Wikipedia. MP4 file format (ISO/IEC 14496-14) and ISO base media file format; see also Digital container format.
- Wikipedia. NTSC: colour encoding — the 1000/1001 frame-rate adjustment behind 29.97 and 23.976 fps; see also Frame rate.
- Wikipedia. SubRip — the de facto
description of the
.srtformat, which has no formal specification. - W3C. WebVTT: The Web Video Text Tracks Format
— the normative specification, including the file structure, cue settings and the
::cuestyling model. - MDN. WebVTT API and the <track> element.
- Wikipedia. SubStation Alpha / Advanced SubStation Alpha.
- libass. libass/libass — the ASS/SSA renderer
used by most players and by FFmpeg’s
subtitlesfilter when burning subtitles into video. - W3C. Timed Text Markup Language 2 (TTML2); see also Wikipedia: TTML.
- Netflix Partner Help Center. English Timed Text Style Guide — 20 CPS adult / 17 CPS children, 42 characters per line, 5/6 s minimum and 7 s maximum duration.
- BBC. Subtitle Guidelines — reading rates, line-break grammar, shot changes, positioning.
- Díaz Cintas, J. & Remael, A. Subtitling: Concepts and Practices (Routledge, 2021) — the standard academic treatment, including the condensation strategies this section alludes to.
- FFmpeg Wiki. Encode/H.264 — CRF ranges, presets and the bitrate/quality relationship. See also the ffmpeg manual for stream mapping and remuxing.
- ffmpeg.wasm. Project site and documentation — FFmpeg compiled to WebAssembly, which is what performs the burn-in in this browser tool.
- Wikipedia. Closed captioning and Subtitles; for the legal and accessibility framing see W3C’s WAI guidance on captions.
Related: how AI translates video covers the stages that produce these files, and how Whisper works explains where the timestamps in them come from. Or export an SRT, a VTT or a burned-in MP4 from your own video — entirely in your browser.