Pipeline

How AI translates video: the whole pipeline, from waveform to subtitle

Updated 5 October 2026 · ~11 min read

“AI translated my video” is shorthand for six distinct things happening in order. Knowing which stage produced a given mistake is the difference between fixing it in ten seconds and re-running the whole job pointlessly. This is a tour of all six, with the failure mode of each.

1. Getting the audio out of the video

No speech model looks at video. The first step is demuxing: opening the container — an MP4, MKV or WebM file — and pulling out the audio elementary stream, discarding the picture entirely.[1] That stream is compressed, usually with AAC in an MP4, so it is then decoded to raw samples.

Those samples almost never match what the model wants. Speech models are trained on a fixed sample rate — 16 kHz mono is the near-universal convention, Whisper included[2] — while your video probably carries 48 kHz stereo. So the audio is downmixed to one channel and resampled. 16 kHz is not an arbitrary number: by the Nyquist–Shannon sampling theorem it preserves frequencies up to 8 kHz,[3] which covers essentially all the information that distinguishes one speech sound from another, while discarding two thirds of the data a music-grade rate would carry.

Failure mode: this stage is where multi-language audio goes wrong. If a film carries a separate dubbed track per language, a naive extractor takes the first audio stream, which may not be the one you meant. Loud background music also gets downmixed on top of the dialogue, and no later stage can undo that — which is why source separation is sometimes run here, before recognition.

2. Speech recognition: audio to timed text

Automatic speech recognition (ASR) converts the waveform into words with timestamps. Classical systems built this from three separately-trained parts: an acoustic model mapping sound to phonemes, a pronunciation dictionary, and a language model scoring word sequences.[4] Modern systems replace all three with one neural network trained end to end — and because it has learned the statistics of language as part of the same objective, it will happily correct “their” to “there” from context.

The model used in this tool is OpenAI’s Whisper, trained on 680,000 hours of weakly supervised multilingual audio scraped from the web.[5] The interesting design decision in Whisper is that it does not treat transcription, translation, language identification and timestamping as separate tasks: all four are expressed as special tokens in a single sequence the decoder predicts. Ask it to transcribe and it writes the source language; flip one token to <|translate|> and the same weights produce English directly from the foreign audio. That mechanism is covered in detail in how Whisper works.

Failure mode: proper nouns, domain jargon, overlapping speakers and heavy accents. ASR accuracy is also strongly language-dependent — Whisper’s own paper reports error rates varying by more than an order of magnitude across its supported languages, tracking how much data each one contributed to training.[5] A model that is excellent in English can be unusable in a low-resource language.

3. Segmentation: text to readable cues

ASR output is a stream of words and times. Subtitles are cues: discrete blocks of one or two lines, each with a start and end time. Turning the former into the latter is a genuine editorial task, and it is the stage most often done badly by automated tools.

The constraints that good segmentation respects are well documented by broadcasters:

Failure mode: cues that are technically correct and physically unreadable — a 14-word block on screen for 1.2 seconds. If generated subtitles feel exhausting to watch even though every word is right, this is the stage at fault, and it is fixable by hand in a text editor without re-running any model.

4. Translation: cues to another language

Neural machine translation (NMT) treats translation as sequence-to-sequence prediction: an encoder reads the source sentence into vectors, a decoder emits target tokens one at a time, attending back to the encoded source at every step.[8] Since 2017 the dominant architecture for both halves has been the transformer,[9] and the current frontier of open multilingual translation is massively multilingual models — Meta’s NLLB-200 covers 200 languages in a single set of weights, deliberately targeting the long tail that bilingual systems ignore.[10]

Two properties of this stage matter enormously for subtitling, and both are about context.

Context starvation

A translation model can only disambiguate using what you give it. Feed it the lone cue “It’s fine.” and it cannot know whether “it” is a car or a decision, whether the speaker is being addressed formally (German Sie vs. du, French vous vs. tu), or what gender agreement any adjective needs. Human subtitlers resolve this from the picture; the model has neither. The practical fix is to join cues back into sentences, translate the sentences, then re-split — which is why translating a finished, readable subtitle file sometimes produces worse output than translating the raw transcript and segmenting afterwards.

Length expansion

Translations change length. English to German or Finnish commonly expands by 20–35%; English to Chinese contracts sharply in character count. A cue that was comfortable at 15 characters per second in the source can breach every reading-speed limit in the target without a single word being mistranslated. Subtitle-aware translation research therefore treats length as an explicit constraint rather than an afterthought.

Failure mode: fluent nonsense. NMT output is grammatical by construction, so translation errors do not look like errors — they look like confident sentences that happen to say the wrong thing. This is categorically harder to spot than an ASR mistake, which usually produces something visibly garbled.

5. Re-timing for the new language

Because of length expansion, the timings carried over from the source language are a starting point, not an answer. A competent pipeline recomputes cue boundaries after translation: redistributing time between adjacent cues, merging two short cues whose combined translation is one sentence, or splitting one cue whose translation no longer fits two lines. The hard constraint is that a cue must not start before its speech does — subtitles that pre-empt the dialogue spoil jokes and reveal plot.

Note that timestamps are ultimately tied to the video’s time base, and a frame-rate mismatch will drift an otherwise perfect subtitle file out of sync over the length of a feature. That is a video-engineering problem rather than an AI one, and it is covered in subtitle formats and video basics.

6. Rendering: sidecar file or burned in

Finally the cues become an artefact. There are two routes, and the choice is consequential:

ApproachWhat it isTrade-off
Sidecar / soft subtitles A separate .srt or .vtt file, or a subtitle track multiplexed into the container; the player draws the text. Editable, switchable, selectable as text, zero quality loss, tiny. But the player must support it, and some platforms silently ignore it.
Burned-in / hard subtitles The text is drawn into the video frames and the video is re-encoded. Displays everywhere, including platforms that strip subtitle tracks and silent autoplay feeds. But it is irreversible, costs one generation of lossy re-encoding, and cannot be turned off or translated again.

For anything you might revise, export the sidecar file and keep it. Burn in only at the last step, for the specific platform that needs it.

Cascaded versus end-to-end

Everything above describes a cascaded pipeline: recognise, then translate. The alternative, end-to-end speech translation, maps source-language audio directly to target-language text in one model — which is exactly what Whisper’s <|translate|> task token does for English output.[5]

End-to-end avoids error propagation: a cascade that mishears a word is guaranteed to mistranslate it, whereas a direct model may still recover from the acoustics. It also preserves prosodic information that a transcript throws away. The costs are real, though: you need training data for that specific direction, you cannot inspect or correct an intermediate transcript, and you get no source-language subtitles as a by-product. For a workflow where a human reviews the output — which is every workflow that cares about quality — the cascade’s inspectability usually wins. You can fix a transcript; you cannot fix a translation whose source you never saw.

How anyone knows whether it worked

Word error rate is the standard ASR metric: the minimum number of insertions, deletions and substitutions needed to turn the hypothesis into the reference, divided by the number of reference words.[11] It is blunt — it weights a missing “the” the same as a negated clause, and can exceed 100% — but it is comparable across systems.

For translation, BLEU compares n-gram overlap with reference translations[12] and remains the reporting default, though character-level chrF and learned neural metrics such as COMET correlate considerably better with human judgement. None of these metrics measures whether a subtitle is readable; that is what the presentation rules in stage 3 are for. A caption track can score well on BLEU and still be unwatchable.

Why running it on your own device matters

Every stage above is now small enough to run in a browser tab. Whisper’s smaller checkpoints are tens to hundreds of megabytes, Transformers.js executes them via WebAssembly and WebGPU,[13] and ffmpeg compiled to WebAssembly handles the demuxing in stage 1 and the burn-in in stage 6.[14]

That is not only a convenience. A video file is unusually revealing — faces, voices, locations, whatever is on the desk in the background — and medical, legal, HR and pre-release material routinely cannot lawfully be handed to a third-party processor at all. A pipeline that never transmits the file removes that question rather than answering it, and it has no per-minute price, no quota and no account. This tool is built that way on purpose.

The practical takeaway

When generated subtitles disappoint, diagnose before re-running:

The pipeline is a chain, and the output is bounded by its worst link — but each link is separately inspectable, and the cheap fix is usually not the one that involves downloading a bigger model.

References

  1. FFmpeg Project. ffmpeg documentation — stream selection, demuxing and the -map option.
  2. OpenAI. openai/whisper, model card and audio.py (16 kHz mono input, 30-second windows).
  3. Wikipedia. Nyquist–Shannon sampling theorem.
  4. Jurafsky, D. & Martin, J. H. Speech and Language Processing, 3rd ed. draft — chapters on automatic speech recognition and machine translation. The standard free textbook on both halves of this pipeline.
  5. Radford, A. et al. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356 — the Whisper paper; see the per-language error-rate tables and the multitask token format.
  6. BBC. Subtitle Guidelines — reading rates, line breaking, shot changes. The most detailed public guidance on subtitle presentation.
  7. Netflix Partner Help Center. English Timed Text Style Guide — 20 CPS, 42 characters per line, duration limits.
  8. Wikipedia. Neural machine translation.
  9. Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762 — the transformer, introduced as a machine-translation architecture.
  10. NLLB Team et al. (2022). No Language Left Behind: Scaling Human-Centered Machine Translation. arXiv:2207.04672.
  11. Wikipedia. Word error rate.
  12. Wikipedia. BLEU; see also Papineni et al. (2002), the originating paper.
  13. Hugging Face. Transformers.js documentation — running transformer models in the browser via ONNX Runtime Web.
  14. ffmpeg.wasm. Project site and documentation.
  15. Further reading: Díaz Cintas, J. & Remael, A. Subtitling: Concepts and Practices (Routledge, 2021) — the standard academic treatment of subtitling as a translation discipline, including condensation strategies no automated system currently matches.

Next: how Whisper works goes inside stage 2, and subtitle formats and video basics covers stage 6. Or run the pipeline on your own video — it never leaves your browser.