Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Berkeley BAIR Blog ·

ABBEL uses belief states as an agent’s working context in place of the full interaction history. The authors describe a reconstruction-based belief grader and a grader that uses domain knowledge. They report comparisons with full-context models and summarization without belief grading. Their CombinationLock result concerns ABBEL with the grader that uses domain knowledge. Lee 3 puntos de vista con sus evidencias y enlaces a las fuentes.

Jakob Bjorner, Aly Lidayan, Satvik Golechha, Kartik Goyal, Alane Suhr

De un vistazo

  • Beliefs replace full interaction history as working context

    Belief states replace the full interaction history as the agent’s working context; belief grading supervises the content of each belief state to improve performance.

    Ver el momento de apoyo · Párrafo 1
  • Reconstruction-based belief grading improves reported training comparisons

    The authors report that a general reconstruction-based belief grader reduces the performance gap with full-context models by about 50%. It trains in 50% fewer steps than models trained to summarize without belief grading (no BG). After training, memory use is significantly lower than in the full-context setting, measured by peak context token length (Peak Tokens).

    Ver el momento de apoyo · Párrafo 18
  • Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency

    In CombinationLock, ABBEL with a domain-knowledge belief grader—using statistics over history and checking reconstructability from the belief state—achieves higher learning efficiency than full-context models.

    Ver el momento de apoyo · Párrafo 19

Pasajes clave3

Pasajes atribuidos con contexto para verificarlos. Abra el texto original para comprobar la fuente.

belief grading efficacy

Reconstruction-based belief grading improves reported training comparisons

Extracto original

We see that with the general reconstruction-based belief grading function we reduce the performance gap from full context models by about 50%, and train in 50% fewer steps compared to training models to summarize without belief grading (no BG). After training, ABBEL still uses significantly less memory than the full context setting, as measured by the peak context token length (Peak Tokens).
LLM context management

Beliefs replace full interaction history as working context

Extracto original

Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.
Contexto

Overview of ABBEL compared to traditional recursive summarization. .

domain-knowledge belief grading

Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency

Extracto original

Additionally, in CombinationLock, we demonstrate that ABBEL with a belief grader which leverages domain knowledge (by computing useful statistics over the history and checking that they can be reconstructed from the belief state), enables even higher learning efficiency than full context (FULL CTX) models.

Fuente y metodología

Estas perspectivas enlazan a sus fuentes originales. Las paráfrasis están identificadas y no son citas textuales.

Abrir transcripción o material de origen (se abre en una pestaña nueva)Reportar un problema