Semantic search
Search a handful of documents by meaning rather than by keyword, with the whole index built in your tab.
Local runtime test
Checking cacheImplementation
The same pipeline can be called from either Python or JavaScript.
from sentence_transformers import SentenceTransformer
# sentence-transformers rather than plain transformers: the pooling and the
# normalization below are what turn token vectors into one comparable vector.
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
query = "A pet taking a nap at home"
documents = [
"The cat is sleeping on the sofa.",
"The stock market fell sharply after the announcement.",
]
vectors = model.encode([query, *documents], normalize_embeddings=True)
scores = vectors[0] @ vectors[1:].T
for document, score in sorted(zip(documents, scores), key=lambda p: -p[1]):
print(f"{score:.2f} {document}")
# ~0.35 The cat is sleeping on the sofa.
# ~0.00 The stock market fell sharply after the announcement.
# Approximate: the quantized build used in the browser shifts them slightly.import { pipeline } from "@huggingface/transformers";
const extractor = await pipeline(
"feature-extraction",
"Xenova/all-MiniLM-L6-v2",
);
const query = "A pet taking a nap at home";
const documents = [
"The cat is sleeping on the sofa.",
"The stock market fell sharply after the announcement.",
];
// One pass over the query and the corpus: mean-pooled and normalized, so a
// dot product between two rows is their cosine similarity.
const output = await extractor([query, ...documents], {
pooling: "mean",
normalize: true,
});
const [width] = [output.dims[1]];
const row = (i) => Array.from(output.data.slice(i * width, (i + 1) * width));
const q = row(0);
documents
.map((document, i) => ({
document,
score: row(i + 1).reduce((sum, value, j) => sum + value * q[j], 0),
}))
.sort((a, b) => b.score - a.score)
.forEach((r) => console.log(r.score.toFixed(2), r.document));What this does
Type a question and a handful of documents, and the page ranks the documents by how close their meaning is to the question — not by which words they share.
Search the five sentences above for “a pet taking a nap at home” and the top result is “The cat is sleeping on the sofa.”
Now look at what those two have in common: not one word. Not “pet”, not “nap”, not “home” — not even an article, since one says “a” and the other says “the”. A keyword search returns nothing at all here. For this to work the model has to know that a cat is a pet, that a nap is sleeping, and that a sofa is somewhere in a home.
The margin says the rest: the cat scores 0.35 and the next sentence, about a train timetable, scores 0.08. Two of the five score 0.00. The model is not hedging between five vaguely similar sentences; it found the one that means what you asked and dismissed the others.
This is the retrieval half of what people mean by “RAG”. There is no index server, no vector database and no API: the embeddings are computed in this tab and compared with arithmetic.
The model
all-MiniLM-L6-v2 is a six-layer sentence encoder distilled from a larger
MiniLM, trained on about a billion sentence pairs to place sentences that
mean the same thing near each other. It maps any text to a 384-number vector.
At roughly 24 MB quantized it is by far the smallest model on this site — a quarter of the image classifier — because it never has to produce language, only place it.
How it works
The word “semantic” does a lot of work in the phrase “semantic search”, so it is worth spelling out what actually happens:
- Every document and the query are passed through the encoder, which produces one vector per token.
- Those token vectors are mean-pooled into a single vector per text. This is the step that turns a sentence of any length into a fixed 384-number address.
- Each vector is normalized to unit length, which makes the dot product between two of them equal to the cosine of the angle between them.
- The query’s vector is compared against each document’s. A score of 1 means the two point the same way; 0 means they are unrelated.
Nothing here is a keyword match. “A pet taking a nap at home” and “the cat is sleeping on the sofa” share no word at all, and they still land close together, because the training pushed sentences with the same meaning into the same region of the space.
Note that the query and the documents are embedded in a single pass. That is not an optimization — comparing vectors only makes sense when they came from the same model with the same pooling.
Read the order, not the number
The scores above are worth a warning. The right answer wins with 0.35 — unambiguous against a runner-up at 0.08, and still nowhere near the 1.0 that the phrase “semantic similarity” tends to suggest. Rewrite the query to name the cat and the sofa outright and the same answer jumps past 0.7 without having become any more correct.
This is normal for MiniLM. Its scores for short English sentences cluster in a narrow band well below 1, because almost any two sentences share some structure. The ranking is the signal; the absolute value is not a probability and should never be shown to a user as a percentage of anything. Systems that threshold on a raw cosine (“show results above 0.8”) tend to return nothing.
Limitations
The comparison is a plain scan over every document, which is exactly right for the handful here and exactly wrong for a million. A real index uses an approximate nearest-neighbour structure so the search does not grow with the corpus.
The model has a 256-word-piece window and truncates beyond it, so pasting a long article gives you an embedding of its opening rather than of the whole thing. Real systems split documents into chunks first, and choosing that chunk size is most of the work.
It was also trained overwhelmingly on English. It will return an ordering for Spanish text, and that ordering will be worse than it looks.