Image classification
Drop in a photo and get the top five ImageNet labels, computed on your own machine.
Local runtime test
Checking cacheImplementation
The same pipeline can be called from either Python or JavaScript.
from transformers import pipeline
classifier = pipeline(
"image-classification",
model="google/vit-base-patch16-224",
)
result = classifier("photo.jpg", top_k=5)
print(result)
# [{'label': 'golden retriever', 'score': 0.81}, ...]import { pipeline } from "@huggingface/transformers";
const classifier = await pipeline(
"image-classification",
"Xenova/vit-base-patch16-224",
);
const result = await classifier("photo.jpg", { top_k: 5 });
console.log(result);
// [{ label: "golden retriever", score: 0.81 }, ...]What this does
Pick one of the sample photos or upload your own, and this page returns the top five labels the model thinks describe the image, each with a confidence score. The model download and every classification run happen entirely inside your browser tab — no image you upload is ever sent to a server, and the demo keeps working offline once the model has loaded.
The model
This demo runs vit-base-patch16-224,
a Vision Transformer (ViT) trained on ImageNet-21k and fine-tuned for
classification over the standard 1000-category ImageNet-1k label set. ViT
applies the transformer architecture — originally built for text — directly
to images, treating an image as a sequence of patches the same way a
sentence is treated as a sequence of tokens. The build used here is the
Xenova/vit-base-patch16-224
ONNX export, quantized to around 88 MB, prepared for the transformers.js
runtime.
How it works
- The page loads transformers.js,
a JavaScript port of the Hugging Face
transformerslibrary that runs models with ONNX Runtime Web. - On first use, the browser downloads the quantized model weights and caches them, so later visits skip the download entirely.
- The chosen image is split into a grid of 16×16 pixel patches. Each patch is flattened and linearly projected into an embedding, exactly the way a word is turned into a token embedding in a text transformer.
- Those patch embeddings, plus a position embedding for each one, are fed through the transformer encoder locally, using WebGPU when it’s available and falling back to WebAssembly otherwise. A classification head on top produces a probability over all 1000 ImageNet categories, and the page renders the top five immediately.
Because classification is a search over a fixed list of 1000 categories, the model can only ever answer with one of those labels — it has no way to say “I don’t recognize this.” A photo of something outside ImageNet’s vocabulary, like a specific person or a niche object, still gets forced into the closest matching category, often with a surprisingly high confidence score. That’s worth keeping in mind when reading the results: a high score means “this is the best match among 1000 fixed options,” not “this is definitely correct.”
Limitations
Classification is a search over a fixed list of 1000 ImageNet categories: the model cannot say “I don’t recognize this,” so anything outside that vocabulary gets forced into the closest match, sometimes with high confidence. It also sees the image at 224×224 pixels and has no notion of location — a small object in a corner counts the same as background, and two dogs count the same as one.