Speech recognition
Transcribe English speech to text on your own machine, with the words appearing as the model decodes them.
Local runtime test
Checking cacheImplementation
The same pipeline can be called from either Python or JavaScript.
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="openai/whisper-base")
# Telling it the language matters. Left to guess on a short clip it often
# guesses wrong, and then transcribes confidently into the wrong one.
print(asr("clip.wav", generate_kwargs={"language": "english", "task": "transcribe"}))import { pipeline } from "@huggingface/transformers";
const asr = await pipeline(
"automatic-speech-recognition",
"Xenova/whisper-base",
);
// samples: mono Float32Array at 16 kHz, decoded in the page with an
// AudioContext constructed at that rate.
console.log(await asr(samples, { language: "english", task: "transcribe" }));What this does
Play a clip and read the transcript. The audio is decoded in this tab, the model runs in this tab, and the words appear as they are decoded rather than all at once at the end.
This page transcribes English only. The model is multilingual on paper, and the section below explains why the page does not pretend it is.
It is also the only experiment here that takes sound, which means it is the one where “nothing leaves your browser” stops being an abstract claim. Upload a voice memo and the recording stays on your machine — the usual way to get a transcript is to send your voice to somebody else’s computer.
The model
Whisper is OpenAI’s speech recognition model, trained on 680,000 hours of
audio collected from the web in 96 languages. whisper-base is the second
smallest of the family: 74 million parameters, 80 MB quantized.
It is an encoder-decoder, like the translation model. The encoder reads a spectrogram of the audio; the decoder writes text one token at a time, which is why the transcript fills in as it works.
How it works
- The audio is resampled to 16 kHz mono, the only rate Whisper
understands. The page does this with an
AudioContextconstructed at that rate, so the browser’s own resampler does the work. - Those samples become a log-Mel spectrogram: a picture of which frequencies are present over time. That picture is what the encoder actually reads.
- The decoder is given the language and the task, then writes tokens until it emits an end-of-transcript marker.
English only, and told so
Whisper can detect the spoken language itself, but on a short clip it often guesses wrong — and then transcribes confidently into the wrong language. So the page always declares the language rather than letting the model guess.
It declares English, and only English, because that is where this size of Whisper spends its capacity. Most of its training audio is English: accented English transcribes well, anything else degrades quickly into confident nonsense. A language selector would offer a setting with one good value, so there isn’t one.
Limitations
It needs clean speech. Close-mic, quiet recordings work; noisy, far-field or archival audio degrades fast. Famous historical recordings on noisy 1960s radio, for example, sit far from anything in the training distribution.
Silence and noise produce invented text, not silence. The decoder always emits something, so a clip with little or no speech tends to come back as a plausible-sounding sentence. There is no “I don’t know” output.
Capacity shows up as accent robustness. Smaller checkpoints handle
studio-clear native speech but mangle ordinary accented voices that the larger
ones get right — which is why this page ships whisper-base.