Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Berkeley BAIR Blog ·

ABBEL uses belief states as an agent’s working context in place of the full interaction history. The authors describe a reconstruction-based belief grader and a grader that uses domain knowledge. They report comparisons with full-context models and summarization without belief grading. Their CombinationLock result concerns ABBEL with the grader that uses domain knowledge. Lies 3 Standpunkte mit Belegen und Links zu den Originalquellen.

Jakob Bjorner, Aly Lidayan, Satvik Golechha, Kartik Goyal, Alane Suhr

Auf einen Blick

  • Beliefs replace full interaction history as working context

    Belief states replace the full interaction history as the agent’s working context; belief grading supervises the content of each belief state to improve performance.

    Unterstützendes Moment lesen · Absatz 1
  • Reconstruction-based belief grading improves reported training comparisons

    The authors report that a general reconstruction-based belief grader reduces the performance gap with full-context models by about 50%. It trains in 50% fewer steps than models trained to summarize without belief grading (no BG). After training, memory use is significantly lower than in the full-context setting, measured by peak context token length (Peak Tokens).

    Unterstützendes Moment lesen · Absatz 18
  • Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency

    In CombinationLock, ABBEL with a domain-knowledge belief grader—using statistics over history and checking reconstructability from the belief state—achieves higher learning efficiency than full-context models.

    Unterstützendes Moment lesen · Absatz 19

Wichtige Passagen3

Zugeordnete Passagen mit dem Kontext zur Überprüfung. Öffnen Sie den Originaltext, um die Quelle zu prüfen.

belief grading efficacy

Reconstruction-based belief grading improves reported training comparisons

Originalauszug

We see that with the general reconstruction-based belief grading function we reduce the performance gap from full context models by about 50%, and train in 50% fewer steps compared to training models to summarize without belief grading (no BG). After training, ABBEL still uses significantly less memory than the full context setting, as measured by the peak context token length (Peak Tokens).
LLM context management

Beliefs replace full interaction history as working context

Originalauszug

Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.
Kontext

Overview of ABBEL compared to traditional recursive summarization. .

domain-knowledge belief grading

Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency

Originalauszug

Additionally, in CombinationLock, we demonstrate that ABBEL with a belief grader which leverages domain knowledge (by computing useful statistics over the history and checking that they can be reconstructed from the belief state), enables even higher learning efficiency than full context (FULL CTX) models.

Quelle & Methodik

Diese Standpunkte sind mit ihren Originalquellen verknüpft. Paraphrasen sind gekennzeichnet und keine wörtlichen Zitate.

Transkript oder Quellenmaterial öffnen (wird in einem neuen Tab geöffnet)Ein Problem melden