PUBLIC VIEWPOINTS

Logan Markewich

Source author · LlamaIndex

1 interviews · 3 viewpoints · 3 topics

Content updated:

Explore connections ↗

Translations are for reading; original excerpts remain the evidence.

Viewpoints by topic

Attributed viewpoints, ordered by interview date. A snapshot of these conversations, not a definitive statement of someone’s beliefs.

static embedding performance

View this topic

Static embeddings can be fast on CPU

Markewich reports that static embeddings take about 0.05 ms per text line on CPU, need no transformer forward pass, are WASM-friendly, and run about 100 times faster than a small dense model.

Supporting evidence

Original excerpt

Static embedding models can seem like a cheat code at first glance. Embedding a line of text is just a table lookup plus an average calculation. This translates to about 0.05 ms per text line on CPU, no transformer forward pass, it's WASM-friendly, and ~100× faster than even a small dense model.
Exploring Static Embedding Retrieval

Additional context is not included in this excerpt. Read the original conversation.

Text location from the supplied original; confirm wording and attribution in context.

model distillation

View this topic
model distillation

Teacher quality stops mattering when the student is too small

From the team’s experiments, Markewich concludes that a better teacher does not help when a small convolutional student lacks the capacity to absorb what the teacher knows.

Supporting evidence

Original excerpt

A better teacher is not a better student. When the student is too small to absorb what the teacher knows (i.e. a small conv model), teacher quality stops mattering.
Exploring Static Embedding Retrieval

Additional context is not included in this excerpt. Read the original conversation.

Text location from the supplied original; confirm wording and attribution in context.

retrieval methodology

View this topic

Raw static token vectors are a poor fit for MaxSim

Markewich advises against scoring raw static token vectors with MaxSim: it performed worse than standard pooling in the team’s exploration, changes the training target, and assumes contextualized vectors that static embeddings do not provide.

Supporting evidence

Original excerpt

Don't score raw static token vectors with MaxSim. It's worse than the standard pooling, the scoring method is technically changing the training target. MaxSim assumes contextualized vectors and static embeddings are not contextual.
Exploring Static Embedding Retrieval

Additional context is not included in this excerpt. Read the original conversation.

Text location from the supplied original; confirm wording and attribution in context.

Statements by interview date1

Statements are ordered by the original interview publication date; differences in wording do not establish a change of position.