THE ORIGINAL SOURCE

Exploring Static Embedding Retrieval

LlamaIndex Blog ·

Logan Markewich reports on the team’s experiments with static embedding retrieval. He explains where MaxSim scoring and small convolutional adapters fell short, despite the speed of static embeddings.

At a glance

Key passages3

Attributed passages with the context to verify them. Open the original text to check the source.

model distillation

Teacher quality stops mattering when the student is too small

Original excerpt

A better teacher is not a better student. When the student is too small to absorb what the teacher knows (i.e. a small conv model), teacher quality stops mattering.

Additional context is not included in this excerpt. Read the original conversation.

retrieval methodology

Raw static token vectors are a poor fit for MaxSim

Original excerpt

Don't score raw static token vectors with MaxSim. It's worse than the standard pooling, the scoring method is technically changing the training target. MaxSim assumes contextualized vectors and static embeddings are not contextual.

Additional context is not included in this excerpt. Read the original conversation.

static embedding performance

Static embeddings can be fast on CPU

Original excerpt

Static embedding models can seem like a cheat code at first glance. Embedding a line of text is just a table lookup plus an average calculation. This translates to about 0.05 ms per text line on CPU, no transformer forward pass, it's WASM-friendly, and ~100× faster than even a small dense model.

Additional context is not included in this excerpt. Read the original conversation.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue