Translation

Translate English into Spanish word by word, with the whole model running on your machine.

Referenceopus-mt-en-es
Footprint119 MB
Pipelinetranslation
RuntimeWebGPU / WASM

Local runtime test

Checking cache
Data inputEnglish
SpanishAwaiting input

Implementation

The same pipeline can be called from either Python or JavaScript.

inference_pipeline.py
from transformers import pipeline

translator = pipeline("translation", model="Helsinki-NLP/opus-mt-en-es")

print(translator("The cat is sleeping on the sofa."))
# [{'translation_text': 'El gato está durmiendo en el sofá.'}]

What this does

Type English, read Spanish. The model is a full translation system running in this tab, and the Spanish arrives a word at a time rather than appearing all at once, because that is the order in which it is produced.

This is the only demo on the site you can check yourself. Everywhere else the page hands you a label and a number and asks you to take its word for it — is that a zebra? is that ranking right? Here both sides are in front of you. If the Spanish is wrong, you will know before the page tells you.

The model

opus-mt-en-es comes from the Helsinki-NLP group’s OPUS-MT project, which published hundreds of translation models trained on the OPUS corpus — a collection of texts that already existed in several languages, such as film subtitles, software localisations and European Parliament proceedings.

It is a Marian model: a small, fast encoder-decoder architecture built for translation specifically, from before general-purpose language models took the job over. It has no instructions, no prompt and no chat. It does one thing.

The quantized ONNX build here is 119 MB — a 53 MB encoder, a 60 MB decoder and a 6 MB tokenizer — which makes it the largest download on this site.

How it works

Translation is generative, and that is what separates this demo from the other five. Sentiment classification computes one answer in a single pass. This model writes its output one token at a time, and each token it writes becomes part of the input for the next one.

  1. The encoder reads the whole English sentence and turns it into a sequence of vectors. This happens once.
  2. The decoder starts with a begin-of-sequence token and predicts the first Spanish token, attending both to the encoder’s output and to what it has written so far.
  3. That token is appended, and the decoder runs again. And again.
  4. It stops when it emits an end-of-sequence token.

The page subscribes to that loop rather than waiting for it. Each token is decoded and posted out of the worker as it is produced, which is why the Spanish fills in across the panel instead of appearing complete. That is not a typing effect: it is the actual rate at which the model is working.

It also explains why a longer sentence takes proportionally longer. The encoder runs once; the decoder runs once per word.

Limitations

It translates one sentence at a time, with no memory of the surrounding text. A pronoun whose referent lives in the previous sentence is a guess, so a paragraph can drift in ways its individual sentences don’t show.

Formality is guessed, not chosen. English “you” maps to both informal tú and formal usted, so the output can switch between them mid-paragraph with nothing in the input to be consistent about.

Fluency is not correctness. Negations and long-range dependencies occasionally drop out while the Spanish stays smooth and grammatical — the output looks no less confident when it is wrong, so check anything that matters.