Speech recognition

Transcribe English speech to text on your own machine, with the words appearing as the model decodes them.

Referencewhisper-base
Footprint80 MB
Pipelineautomatic-speech-recognition
RuntimeWebGPU / WASM

Local runtime test

Checking cache
Data input

Pick a clip or upload your own

TranscriptAwaiting input

Implementation

The same pipeline can be called from either Python or JavaScript.

inference_pipeline.py
from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="openai/whisper-base")

# Telling it the language matters. Left to guess on a short clip it often
# guesses wrong, and then transcribes confidently into the wrong one.
print(asr("clip.wav", generate_kwargs={"language": "english", "task": "transcribe"}))

What this does

Play a clip and read the transcript. The audio is decoded in this tab, the model runs in this tab, and the words appear as they are decoded rather than all at once at the end.

This page transcribes English only. The model is multilingual on paper, and the section below explains why the page does not pretend it is.

It is also the only experiment here that takes sound, which means it is the one where “nothing leaves your browser” stops being an abstract claim. Upload a voice memo and the recording stays on your machine — the usual way to get a transcript is to send your voice to somebody else’s computer.

The model

Whisper is OpenAI’s speech recognition model, trained on 680,000 hours of audio collected from the web in 96 languages. whisper-base is the second smallest of the family: 74 million parameters, 80 MB quantized.

It is an encoder-decoder, like the translation model. The encoder reads a spectrogram of the audio; the decoder writes text one token at a time, which is why the transcript fills in as it works.

How it works

  1. The audio is resampled to 16 kHz mono, the only rate Whisper understands. The page does this with an AudioContext constructed at that rate, so the browser’s own resampler does the work.
  2. Those samples become a log-Mel spectrogram: a picture of which frequencies are present over time. That picture is what the encoder actually reads.
  3. The decoder is given the language and the task, then writes tokens until it emits an end-of-transcript marker.

English only, and told so

Whisper can detect the spoken language itself, but on a short clip it often guesses wrong — and then transcribes confidently into the wrong language. So the page always declares the language rather than letting the model guess.

It declares English, and only English, because that is where this size of Whisper spends its capacity. Most of its training audio is English: accented English transcribes well, anything else degrades quickly into confident nonsense. A language selector would offer a setting with one good value, so there isn’t one.

Limitations

It needs clean speech. Close-mic, quiet recordings work; noisy, far-field or archival audio degrades fast. Famous historical recordings on noisy 1960s radio, for example, sit far from anything in the training distribution.

Silence and noise produce invented text, not silence. The decoder always emits something, so a clip with little or no speech tends to come back as a plausible-sounding sentence. There is no “I don’t know” output.

Capacity shows up as accent robustness. Smaller checkpoints handle studio-clear native speech but mangle ordinary accented voices that the larger ones get right — which is why this page ships whisper-base.