Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency
In CombinationLock, ABBEL with a domain-knowledge belief grader—using statistics over history and checking reconstructability from the belief state—achieves higher learning efficiency than full-context models.
Stützende Belege
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Originalauszug
Additionally, in CombinationLock, we demonstrate that ABBEL with a belief grader which leverages domain knowledge (by computing useful statistics over the history and checking that they can be reconstructed from the belief state), enables even higher learning efficiency than full context (FULL CTX) models.