Model internals
How Whisper works: inside the speech model that writes your subtitles
Updated 5 October 2026 · ~12 min read
Whisper is not an exotic piece of engineering. It is the same encoder–decoder transformer that powers machine translation, pointed at a picture of sound, and trained on an enormous pile of imperfectly-labelled audio from the web. Almost everything surprising about its behaviour — including the strange habit of writing subtitles over total silence — falls out of those three facts.
Step 1: sound becomes a picture
Audio arrives as a long list of amplitude samples — 16,000 numbers per second of mono audio, in Whisper’s case.[1] Raw samples are a terrible input for a recognition model: two recordings of the same word, shifted by a few milliseconds, look numerically unrelated.
So the audio is converted into a log-Mel spectrogram, which is built in four moves:
- Window it. Chop the signal into short overlapping frames (Whisper uses a 25 ms window stepped every 10 ms), because speech is only approximately stationary over tens of milliseconds.
- Transform each frame. A short-time Fourier transform converts each frame from amplitude over time into energy per frequency.[2]
- Warp the frequency axis. Group those frequencies into Mel bands — 80 of them for Whisper. The Mel scale is approximately linear below 1 kHz and logarithmic above it, modelling the fact that human hearing discriminates low frequencies far more finely than high ones.[3] A 100 Hz difference is obvious at 200 Hz and inaudible at 8 kHz.
- Take the logarithm. Perceived loudness is roughly logarithmic in intensity, so a log compresses the enormous dynamic range into something a neural network can train on stably.
The result is a 2-D array: 80 frequency bands × 3,000 time steps for a 30-second window. Visually it is a heat map where formants show up as horizontal bands and consonant bursts as vertical streaks. This representation, and the pipeline that produces it, long predates deep learning — Mel-frequency features were the backbone of speech recognition for decades.[3] Whisper kept the front end and replaced everything after it.
Step 2: the encoder reads the picture
Two strided 1-D convolutions first compress the 3,000 time steps to 1,500 and project each into a vector. A fixed sinusoidal positional encoding is added so the model knows frame order, and the sequence goes through a stack of transformer encoder blocks.[1]
Each block does two things. Self-attention lets every time step look at every other time step and decide which ones matter: that is how the model uses the end of a word to disambiguate its beginning, or a speaker’s pitch range several seconds earlier to interpret a vowel now.[4] A position-wise feed-forward network then transforms each step independently. Residual connections and layer normalisation around both keep gradients healthy through dozens of layers.[5]
The crucial property is that attention is global and unmasked in the encoder. The model sees the whole 30-second window at once, both directions in time. This is why Whisper handles disfluency and accent better than streaming systems that must commit to a word before hearing what follows.
The encoder’s output is 1,500 vectors describing what is happening acoustically at each moment. No text yet.
Step 3: the decoder writes text
The decoder is an autoregressive language model. It predicts one token at a time, where a token is a subword fragment from a byte-pair-encoded vocabulary — common words are single tokens, rare ones split into pieces, so any string in any language is representable without an unbounded vocabulary.[6]
Each decoder block contains three sublayers:
- Masked self-attention over the tokens generated so far — masked so position i cannot see the future, which is what makes generation causal.
- Cross-attention onto the encoder output — this is the link between sound and text, and the only place the audio enters. The decoder effectively asks “given what I have written, which part of the audio should I be listening to now?”
- A feed-forward network.
At each step the model produces a probability distribution over the whole vocabulary. Picking the single most likely token every time (greedy decoding) is fast but gets trapped by locally-attractive choices; keeping the k best partial sequences and extending them in parallel — beam search — finds higher-probability complete sequences at proportionally more compute.[7] Whisper’s reference implementation also uses a temperature fallback: if the output looks degenerate (compression ratio too high, average log-probability too low) it retries that window with increasing randomness.[8]
The multitask trick: tasks as tokens
This is Whisper’s defining design choice. Rather than train one model for transcription, one for language identification, one for translation and a separate aligner for timestamps, all four are folded into the decoder’s output sequence as special tokens.[1] A sequence looks roughly like:
<|startoftranscript|> <|de|> <|transcribe|> <|0.00|>
Guten Abend, meine Damen und Herren.
<|3.44|> ... <|endoftext|>
The implications of that one format are large:
- Language identification is free. The token right after
<|startoftranscript|>is the language. Whisper’s auto-detect is literally reading off the probability distribution for that single position — which is also why it detects best when it has a few seconds of clear speech, and can latch onto the wrong language if the window opens on noise. - Translation is a flipped token. Emit
<|translate|>instead of<|transcribe|>and the same weights produce English text from non-English audio, with no separate translation model. - Timestamps are vocabulary items. Times are tokens quantised to 20 ms, predicted alongside words.
- Voice activity detection is free too. A
<|nospeech|>token gives the model a way to say “nothing was said here,” and its probability is a usable silence detector.
The underlying idea — recast every task as text-to-text prediction so one model and one objective cover them all — is the same one behind text models like T5,[9] and it is why the interface to Whisper is so uniform.
30-second windows and long audio
Whisper is architecturally fixed to 30 seconds of audio. Shorter clips are zero-padded to 30; longer ones are not simply fed in.[1] Your 40-minute recording is processed as a sequence of windows, and the obvious naive approach — cut every 30 seconds exactly — slices words in half at every boundary.
The reference implementation instead advances the window using the model’s own last reliable timestamp, so each window starts where the previous one genuinely finished mid-silence rather than mid-syllable. It also optionally feeds the previous window’s text back in as a prompt, giving the decoder textual context across the boundary: useful for consistent spelling of names, but a liability if the previous window went wrong, because the error gets conditioned on and can propagate for minutes.
An alternative is chunked batch inference: overlapping windows transcribed independently and stitched by matching the overlap. That loses cross-window context but is embarrassingly parallel and cannot propagate an error forward — which is why it is the common choice in browser and server implementations where throughput matters.
Where the timestamps come from — and why they are approximate
Because timestamps are predicted tokens rather than measured alignments, they are a model opinion about when speech happened. They are usually good to a few hundred milliseconds at segment level and distinctly unreliable at word level.
For subtitles this matters at the edges: a cue that starts 300 ms early is unobjectionable, but one that drifts a second late reads as broken. Systems needing tighter sync run forced alignment afterwards — taking the transcript as known and finding the most probable time alignment against the audio with a separate phoneme-level model. If you only need readable subtitles, Whisper’s native timestamps plus manual nudging is enough; if you need frame-accurate karaoke, you need an aligner.
Why weak supervision made it robust
The architecture is ordinary. The data is what made Whisper notable: 680,000 hours of audio paired with transcripts harvested from the internet, 117,000 hours of it non-English, filtered only heuristically — machine-generated transcripts were detected and removed, to avoid training a model to imitate another ASR system’s mistakes.[1]
Academic ASR had long been trained on clean, carefully-annotated corpora, and models trained that way score superbly on their own test set and degrade badly on real-world audio. Whisper’s training set is instead full of the mess that real audio contains: room reverb, telephone bandwidth, music beds, overlapping speakers, every accent. The paper’s central result is about zero-shot generalisation — competitive accuracy on benchmarks it was never fine-tuned on, approaching human robustness across varied conditions.[1] Scale and diversity of weakly-labelled data bought generalisation that clean data at smaller scale did not.
The same data explains the uneven language coverage: performance per language tracks that language’s share of the 680,000 hours. English dominates; a language contributing a few hundred hours is served far worse. There is no architectural fix for that, only more data.
Hallucination, and why it happens
Whisper will sometimes produce fluent text for audio that contains no speech — a plausible sentence over silence, a stretch of music transcribed as dialogue, or a phrase like a subtitle-site credit line that was frequent in its training data. Users report it as a bug. It is a direct consequence of the design.
The decoder is a language model that must emit tokens for every window. Over clear speech, cross-attention constrains it tightly. Over silence or noise there is no acoustic evidence to constrain anything, so the language-model prior takes over and generates what is linguistically likely, which over silence means whatever boilerplate was common in the training transcripts. The model has no mechanism for “I am not confident”; its output distribution is always normalised to sum to one.
Looping is the same phenomenon: having emitted a phrase, a repetition of it becomes the highest-probability continuation, and the model falls into a cycle. The standard mitigations all amount to detecting the condition from outside the model:
- Run voice activity detection first and skip windows with no speech.
- Threshold on the
<|nospeech|>probability and on average log-probability. - Watch the compression ratio of the output — highly repetitive text compresses suspiciously well — and retry at a different temperature.[8]
- Read the result. Hallucinations tend to cluster at the start and end of recordings, where lead-in silence and trailing room tone live — worth a specific look.
This is the strongest argument for keeping subtitles editable rather than treating ASR output as finished: the failures are not random noise, they are confident prose, and a human spots them in seconds.
tiny vs. base vs. small: what you actually trade
Whisper ships in a family of sizes sharing one architecture, differing in layer count, width and attention heads.[1] Three fit comfortably in a browser:
| Model | Download | Best for | Watch out for |
|---|---|---|---|
tiny | ~75 MB | Fast drafts of clean English speech; low-powered devices | Weak on accents, noise, proper nouns; leans English |
base | ~145 MB | The sensible default; handles non-English audio competently | Still struggles with jargon and heavy overlap |
small | ~480 MB | Accents, background noise, technical vocabulary, non-English audio | Several times slower; a real download |
Two rules of thumb hold up well in practice. First, the accuracy gain from going up a size is much
larger for non-English audio than for clean English — if you are transcribing English podcast audio,
base is often within a point or two of small; if you are transcribing Hindi or
Turkish, the step up is substantial. Second, if auto-detect picks the wrong language, selecting the language
explicitly helps more than a bigger model does, and costs nothing.
How a transformer runs in a browser tab
Four pieces of infrastructure make this practical without a server:
- WebAssembly gives near-native-speed numeric execution in the browser sandbox,[10] which is what the ONNX Runtime backend behind Transformers.js compiles to.[11]
- WebGPU exposes the GPU for general-purpose compute with modern explicit APIs,[12] and matrix multiplication is exactly the workload GPUs exist for. Where it is available, transcription is dramatically faster than on the CPU path.
- Web Workers move inference onto a background thread so the UI does not freeze for the
minutes a transcription takes.[13] (This is also why these pages need to be served over
HTTP: browsers block workers on
file://origins.) - Quantisation shrinks the weights. Storing them as 8-bit integers instead of 32-bit floats cuts the download roughly fourfold and speeds up inference, at a small and usually imperceptible accuracy cost.[14] This is the single technique that makes a 150 MB browser download of a real speech model possible at all.
Once fetched, the weights sit in the browser’s cache, so the second transcription starts immediately. The tool on this site deliberately downloads nothing until you pick a model.
What Whisper cannot do
- Speaker diarisation. It transcribes words, not who said them. No speaker labels, no notion of distinct speakers. That needs a separate model aligned onto Whisper’s timestamps.
- Reliable word-level timing. See above — segment-level is fine, word-level needs a forced aligner.
- Streaming. The encoder is bidirectional over a full 30-second window, so true low-latency live captioning requires a different architecture or awkward overlapping-window tricks.
- Translating into arbitrary languages. The
<|translate|>task was trained for X → English. Other target languages need a separate translation model — which is what the translate step in this tool uses. - Knowing that it is wrong. Its confidence is not calibrated to truth. Treat the output as a first draft, always.
References
- Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C. & Sutskever, I. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356. The primary source for everything structural here: the 680k-hour dataset, the multitask token format, the 30-second window, the model-size table and the per-language results. (PDF)
- Wikipedia. Short-time Fourier transform and Spectrogram.
- Wikipedia. Mel scale and Mel-frequency cepstrum.
- Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762 — scaled dot-product attention, multi-head attention, the encoder–decoder stack and sinusoidal positional encoding, all of which Whisper uses essentially unmodified. See also Wikipedia: Transformer.
- Goodfellow, I., Bengio, Y. & Courville, A. Deep Learning (MIT Press, 2016), free online — background on residual connections, normalisation and sequence models.
- Sennrich, R., Haddow, B. & Birch, A. (2015). Neural Machine Translation of Rare Words with Subword Units. arXiv:1508.07909 — byte-pair encoding. See also Wikipedia: Byte pair encoding.
- Wikipedia. Beam search.
- OpenAI. openai/whisper — the reference
implementation;
transcribe.pycontains the sliding-window logic, temperature fallback, and the compression-ratio and log-probability thresholds used to detect degenerate output. - Raffel, C. et al. (2019). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 — the “every task is text-to-text” framing Whisper applies to speech.
- W3C. WebAssembly — overview and specifications.
- Hugging Face. Transformers.js documentation and ONNX Runtime Web.
- W3C. WebGPU specification; see also MDN: WebGPU API.
- MDN. Using Web Workers.
- Wikipedia. Quantization; for the neural-network case see Jacob, B. et al. (2017), Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, arXiv:1712.05877.
- Also useful: Jurafsky, D. & Martin, J. H., Speech and Language Processing, 3rd ed. draft — the ASR and feature-extraction chapters cover the spectrogram front end and evaluation in textbook depth.
Related: how AI translates video puts this model in the context of a full subtitling pipeline, and subtitle formats and video basics covers what happens to its output afterwards. Or try all three model sizes on your own video — nothing is uploaded.