Browser-local inference · No API keys
Local AI Lab
A collection of experiments exploring local AI inference in the browser using open-source models from Hugging Face.
How it runs
The purpose of this repository is to document what I learn while experimenting with AI models that can run locally, without relying on external inference APIs, API keys, or sending data to a server. Each experiment is both a short technical write-up and a working demo.
Each page focuses on a specific AI experiment. The first interaction shows the model's download size, then downloads quantized ONNX weights from the Hugging Face CDN and caches them in the browser. Inference happens in a Web Worker: WebGPU when the browser can initialize it, with WASM as a CPU fallback.
Pick an experiment, load it once, and try the samples. The Python and JavaScript snippets under each demo show the same pipeline outside this site.
Experiments
Review sentiment analysis
Classify a restaurant review as positive or negative, without sending a single character to a server.
Image classification
Drop in a photo and get the top five ImageNet labels, computed on your own machine.
Zero-shot classification
Invent your own categories and classify text into them, with no training data and no fine-tuning.
Semantic search
Search a handful of documents by meaning rather than by keyword, with the whole index built in your tab.
Object detection
Find every object in a photo and draw a box around it, with the boxes computed on your own machine.
Translation
Translate English into Spanish word by word, with the whole model running on your machine.
Speech recognition
Transcribe English speech to text on your own machine, with the words appearing as the model decodes them.