Formats

Subtitle formats and video basics: SRT, VTT, ASS, containers, codecs and burn-in

Updated 5 October 2026 · ~10 min read

Most subtitle frustration is not about transcription accuracy. It is about not knowing that an MP4 is not a codec, that a “subtitle” is not part of the picture until you make it so, and that making it so costs you a generation of video quality. This is the vocabulary, with the gotchas attached.

Containers versus codecs: the distinction everything else rests on

A codec is a compression scheme for one kind of data. H.264/AVC and H.265/HEVC compress pictures; AAC and Opus compress audio.[1] A codec defines how to turn samples into a compact bitstream and back.

A container is the file format that holds those bitstreams together and records how they relate. MP4 — formally ISO/IEC 14496-14, derived from the ISO base media file format — stores video, audio, subtitle and metadata tracks as a tree of boxes, with tables saying which chunk belongs to which track and at what timestamp each sample should be presented.[2] MKV, WebM and MOV are alternative containers solving the same problem with different trade-offs; MKV is the most permissive about what it will carry.

This is why “my MP4 doesn’t play” is almost never about the MP4. The container parses fine; the codec inside is one the player cannot decode. It is also why remuxing (repackaging streams into a different container) is instant and lossless, while transcoding (decoding and re-compressing) is slow and lossy. Knowing which operation you are asking for predicts both the time it takes and the quality you lose.

Two more terms worth having:

Frame rates, time bases and the 23.976 problem

Video is a sequence of frames shown at a frame rate. Common ones are 24 fps (cinema), 25 fps (PAL regions), 30 and 60 fps (screen capture and sport), and the two odd ones: 23.976 and 29.97 fps.

Those two exist for a historical reason. When colour was added to the NTSC television standard, the frame rate was pulled down by a factor of 1000/1001 to keep the new colour subcarrier from interfering with the audio carrier.[3] Seventy years later, every file in the North American production chain still carries the consequence.

The 0.1% difference between 24 and 23.976 sounds negligible. Over a two-hour feature it accumulates to about 4.3 seconds — which is the single most common cause of subtitles that start in sync and end hopelessly late. The diagnostic is simple and worth memorising:

Subtitle formats themselves store wall-clock times, not frame numbers, so they are nominally frame-rate agnostic. The mismatch comes from the timings having been authored against a differently-rated copy of the video. Separately, a container has its own time base — the unit its timestamps are counted in — and rounding when converting between time bases is why a cue can land a frame off even when everything else is right.

Soft, muxed and hard subtitles

There are three places subtitle text can live, and they behave very differently:

FormWhere the text livesProperties
Sidecar A separate .srt/.vtt file next to the video Editable in any text editor, switchable, zero quality cost, kilobytes. Can be lost or renamed apart from the video; must be delivered alongside it.
Muxed (soft) track A subtitle track inside the container Travels with the file, multiple languages selectable, still lossless. Player support for the specific subtitle codec varies; some platforms discard the track on re-upload.
Hardsubs (burned in) Drawn into the pixels; the video is re-encoded Displays on literally anything, survives any platform, works in silent autoplay feeds. Irreversible, not switchable, not selectable as text, and costs one lossy generation.

The practical rule: keep the sidecar as your master and burn in only as a final export for a specific destination. Hardsubs are a rendering decision, not a storage format. Once you have thrown away the text, correcting a typo means re-encoding the whole video again — a second lossy generation on top of the first.

SubRip (.srt): the lingua franca

SRT came out of the SubRip DVD-ripping program in the late 1990s and became ubiquitous by being trivially simple. There is no formal specification; the format is whatever the ecosystem agrees to parse.[4]

1
00:00:01,000 --> 00:00:03,500
The first line of dialogue.

2
00:00:03,600 --> 00:00:06,120
A second cue, split
across two lines.

A cue is: a sequence number, a timecode line with -->, one or more text lines, a blank line. Notable details:

Choose SRT when you want the file to work in the largest number of places with the least thought.

WebVTT (.vtt): the web standard

WebVTT is a W3C specification, and it is the format HTML natively understands: a <track> element inside a <video> points at a .vtt file and the browser renders the cues itself, exposing them to scripts through the TextTrack API.[5][6]

WEBVTT

intro
00:00:01.000 --> 00:00:03.500
The first line of dialogue.

00:00:03.600 --> 00:00:06.120 line:90% align:center
<i>A styled, positioned cue.</i>

Differences from SRT that matter in practice:

Choose WebVTT when the target is a browser, an HTML5 player, or HLS streaming.

ASS, and the broadcast formats

Advanced SubStation Alpha (.ass) is a different class of format: it carries named styles, fonts, colours, arbitrary positioning, rotation, animated transforms and karaoke timing, with a scripting syntax inside the cue text.[7] It grew up in anime fansubbing, where typesetting signage and song lyrics mattered, and it is the format libass renders — the library most players and ffmpeg use for subtitle drawing, including when burning in.[8] Use it when you need typographic control; accept that fewer platforms ingest it.

Beyond these sit the professional formats you will meet only at delivery: TTML and its profiles (IMSC, SMPTE-TT), XML-based and the basis of most modern broadcast and streaming exchange;[9] EBU-STL, the long-standing European broadcast interchange format; and CEA-608/708, closed captions encoded into the video signal itself rather than carried as a separate text track. If a distributor asks for “IMSC 1.1”, that is a TTML profile, and you convert rather than author.

Which one to export

Reading-speed rules that make subtitles legible

Correct text is not the same as readable subtitles. The numbers below come from published professional guidelines and are worth applying even to casual work — they are the difference between subtitles people can follow and subtitles people switch off.

ConstraintTypical ruleWhy
Reading speedUp to ~20 characters/second for adults; ~17 CPS for children[10]; the BBC works to around 160–180 wpm[11]Above this, viewers stop watching the picture and still miss text
Line length≤ 42 characters per line[10]Long lines force horizontal eye travel and cover the frame
Lines per cueMaximum 2Three lines obscure too much picture and invert reading order
Minimum duration~5/6 second (20 frames at 24 fps)[10]Shorter and the cue flashes before it can be fixated
Maximum duration~7 seconds[10]A cue lingering past its dialogue reads as a fault
Gap between cuesAt least 2 framesWithout a gap, consecutive cues look like one flickering block
Line breaksAt clause boundaries; never between article and noun[11]Syntactic breaks are read in one saccade; arbitrary ones are not
Shot changesAvoid cues spanning a cut[11]A cut makes the eye re-read the text it had already started

The consequence professionals accept and amateurs resist: at these limits you often cannot fit every word. Condensing dense speech — cutting filler, compressing phrasing, keeping meaning — is a core subtitling skill, and it is the one thing no automatic system currently does well.[12] Automatic transcription is verbatim by nature, so the single highest-value edit you can make to generated subtitles is usually deletion.

What burn-in actually costs

Burning in means decoding every frame, drawing text onto it, and encoding the result. Three consequences follow, and all three surprise people:

1. A lossy generation

The source video was already compressed with loss. Re-encoding compresses the decoded result again, and the artefacts of the first pass are now treated as detail worth preserving, which wastes bitrate. With H.264 this is controlled by the constant rate factor: lower CRF means higher quality and a larger file, and the FFmpeg project’s guidance is that CRF 18–23 is the sane range for x264, with roughly ±6 halving or doubling the bitrate.[13] Picking a conservative CRF keeps the second generation visually close to the first; picking a careless one is where “why does it look soft now?” comes from.

2. Real compute time

Transcoding is frame-by-frame work on every pixel, so burn-in time scales with resolution and length, not with how much text there is. In the browser, where ffmpeg runs compiled to WebAssembly,[14] expect it to be slower than a native encoder — a reasonable trade for never uploading the file.

3. Irreversibility

Pixels are pixels. There is no un-burn. The text can no longer be searched, selected, translated, restyled or switched off, and a screen reader cannot reach it. For accessibility specifically, soft subtitles are the better artefact in every respect.

The workflow that avoids all three: transcribe, edit, export the sidecar, keep it. Burn in last, from the original video rather than from a previously burned copy, once per destination that needs it. This tool does exactly that — subtitle export and burn-in are separate buttons, and burn-in always works from your original file.

Subtitles versus captions — not a synonym

In North American usage the distinction is substantive, not pedantic. Subtitles render dialogue for someone who can hear the audio but does not understand the language. Closed captions serve viewers who cannot hear it, so they additionally convey speaker identification and non-speech audio — [door slams], [ominous music] — because that information is otherwise lost entirely.[15]

Automatic speech recognition produces subtitles, not captions: it transcribes speech and is blind to everything else, and it does not know who is speaking (see what Whisper cannot do). If the deliverable is an accessibility caption track, the sound-effect cues and speaker labels are human additions — which is fine, because both are easy to add to an SRT by hand.

Troubleshooting, in order of likelihood

References

  1. Wikipedia. Advanced Video Coding (H.264), HEVC (H.265) and Advanced Audio Coding.
  2. Wikipedia. MP4 file format (ISO/IEC 14496-14) and ISO base media file format; see also Digital container format.
  3. Wikipedia. NTSC: colour encoding — the 1000/1001 frame-rate adjustment behind 29.97 and 23.976 fps; see also Frame rate.
  4. Wikipedia. SubRip — the de facto description of the .srt format, which has no formal specification.
  5. W3C. WebVTT: The Web Video Text Tracks Format — the normative specification, including the file structure, cue settings and the ::cue styling model.
  6. MDN. WebVTT API and the <track> element.
  7. Wikipedia. SubStation Alpha / Advanced SubStation Alpha.
  8. libass. libass/libass — the ASS/SSA renderer used by most players and by FFmpeg’s subtitles filter when burning subtitles into video.
  9. W3C. Timed Text Markup Language 2 (TTML2); see also Wikipedia: TTML.
  10. Netflix Partner Help Center. English Timed Text Style Guide — 20 CPS adult / 17 CPS children, 42 characters per line, 5/6 s minimum and 7 s maximum duration.
  11. BBC. Subtitle Guidelines — reading rates, line-break grammar, shot changes, positioning.
  12. Díaz Cintas, J. & Remael, A. Subtitling: Concepts and Practices (Routledge, 2021) — the standard academic treatment, including the condensation strategies this section alludes to.
  13. FFmpeg Wiki. Encode/H.264 — CRF ranges, presets and the bitrate/quality relationship. See also the ffmpeg manual for stream mapping and remuxing.
  14. ffmpeg.wasm. Project site and documentation — FFmpeg compiled to WebAssembly, which is what performs the burn-in in this browser tool.
  15. Wikipedia. Closed captioning and Subtitles; for the legal and accessibility framing see W3C’s WAI guidance on captions.

Related: how AI translates video covers the stages that produce these files, and how Whisper works explains where the timestamps in them come from. Or export an SRT, a VTT or a burned-in MP4 from your own video — entirely in your browser.