Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Berkeley BAIR Blog ·

ABBEL uses belief states as an agent’s working context in place of the full interaction history. The authors describe a reconstruction-based belief grader and a grader that uses domain knowledge. They report comparisons with full-context models and summarization without belief grading. Their CombinationLock result concerns ABBEL with the grader that uses domain knowledge. Lisez 3 points de vue avec leurs éléments à l’appui et les liens vers les sources.

Jakob Bjorner, Aly Lidayan, Satvik Golechha, Kartik Goyal, Alane Suhr

En un coup d’œil

  • Beliefs replace full interaction history as working context

    Belief states replace the full interaction history as the agent’s working context; belief grading supervises the content of each belief state to improve performance.

    Lire le moment probant · Paragraphe 1
  • Reconstruction-based belief grading improves reported training comparisons

    The authors report that a general reconstruction-based belief grader reduces the performance gap with full-context models by about 50%. It trains in 50% fewer steps than models trained to summarize without belief grading (no BG). After training, memory use is significantly lower than in the full-context setting, measured by peak context token length (Peak Tokens).

    Lire le moment probant · Paragraphe 18
  • Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency

    In CombinationLock, ABBEL with a domain-knowledge belief grader—using statistics over history and checking reconstructability from the belief state—achieves higher learning efficiency than full-context models.

    Lire le moment probant · Paragraphe 19

Passages clés3

Passages attribués et accompagnés du contexte nécessaire à leur vérification. Ouvrez le texte original pour vérifier la source.

belief grading efficacy

Reconstruction-based belief grading improves reported training comparisons

Extrait original

We see that with the general reconstruction-based belief grading function we reduce the performance gap from full context models by about 50%, and train in 50% fewer steps compared to training models to summarize without belief grading (no BG). After training, ABBEL still uses significantly less memory than the full context setting, as measured by the peak context token length (Peak Tokens).
LLM context management

Beliefs replace full interaction history as working context

Extrait original

Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.
Contexte

Overview of ABBEL compared to traditional recursive summarization. .

domain-knowledge belief grading

Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency

Extrait original

Additionally, in CombinationLock, we demonstrate that ABBEL with a belief grader which leverages domain knowledge (by computing useful statistics over the history and checking that they can be reconstructed from the belief state), enables even higher learning efficiency than full context (FULL CTX) models.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème