Domain-knowledge belief grading enables ABBEL to match or exceed full-context learning efficiency
In CombinationLock, ABBEL with a domain-knowledge belief grader—using statistics over history and checking reconstructability from the belief state—achieves higher learning efficiency than full-context models.
Supporting evidence
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Original excerpt
Additionally, in CombinationLock, we demonstrate that ABBEL with a belief grader which leverages domain knowledge (by computing useful statistics over the history and checking that they can be reconstructed from the belief state), enables even higher learning efficiency than full context (FULL CTX) models.