LA SOURCE ORIGINALE

Exploring Static Embedding Retrieval

LlamaIndex Blog ·

Logan Markewich reports on the team’s experiments with static embedding retrieval. He explains where MaxSim scoring and small convolutional adapters fell short, despite the speed of static embeddings.

En un coup d’œil

Passages clés3

Passages attribués et accompagnés du contexte nécessaire à leur vérification. Ouvrez le texte original pour vérifier la source.

model distillation

Teacher quality stops mattering when the student is too small

Extrait original

A better teacher is not a better student. When the student is too small to absorb what the teacher knows (i.e. a small conv model), teacher quality stops mattering.

Cet extrait ne contient pas de contexte supplémentaire. Consultez la conversation originale.

retrieval methodology

Raw static token vectors are a poor fit for MaxSim

Extrait original

Don't score raw static token vectors with MaxSim. It's worse than the standard pooling, the scoring method is technically changing the training target. MaxSim assumes contextualized vectors and static embeddings are not contextual.

Cet extrait ne contient pas de contexte supplémentaire. Consultez la conversation originale.

static embedding performance

Static embeddings can be fast on CPU

Extrait original

Static embedding models can seem like a cheat code at first glance. Embedding a line of text is just a table lookup plus an average calculation. This translates to about 0.05 ms per text line on CPU, no transformer forward pass, it's WASM-friendly, and ~100× faster than even a small dense model.

Cet extrait ne contient pas de contexte supplémentaire. Consultez la conversation originale.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème