Long-form, referenced explainers about the machinery behind this free browser-based subtitle generator: how speech recognition turns air pressure into text, how translation models move that text between languages, and what the subtitle and video formats on the export button actually are. Every substantive claim links to a primary source — a paper, a specification, a standards body or a published style guide.
Pipeline
Six stages sit between a video file and a translated caption track: demuxing and resampling the audio, recognising the speech, splitting it into readable cues, translating it, re-timing it, and rendering it. What each stage does, why errors compound, and which problems are actually translation problems rather than recognition problems.
~11 min read · ASR, NMT, cue segmentation, evaluation metrics
Model internals
Whisper is a plain encoder–decoder transformer taught to treat transcription, translation, language
identification and timestamping as one next-token prediction problem. Log-Mel spectrograms, 30-second
windows, the special-token prompt format, why it sometimes invents a sentence of text over silence, and
what you trade away when you pick tiny over
small.
~12 min read · spectrograms, transformers, quantisation, WER
Formats
An MP4 is not a codec, a subtitle is not part of the picture until you make it so, and “hardcoding” captions costs you a generation of video quality. The practical differences between SubRip, WebVTT and ASS; soft versus hard subtitles; frame rates and the 23.976 problem; and the reading-speed limits professional subtitlers actually work to.
~10 min read · SRT/VTT/ASS, muxing, H.264, CPS limits
Everything discussed here runs locally in
Local Whisper Subtitles:
drop in an MP4, transcribe it with Whisper, edit the text, translate it, then export
.srt/.vtt
or burn the captions into the video. No upload, no account, no quota.