Reconstruction-based belief grading improves reported training comparisons
The authors report that a general reconstruction-based belief grader reduces the performance gap with full-context models by about 50%. It trains in 50% fewer steps than models trained to summarize without belief grading (no BG). After training, memory use is significantly lower than in the full-context setting, measured by peak context token length (Peak Tokens).
Supporting evidence
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Original excerpt
We see that with the general reconstruction-based belief grading function we reduce the performance gap from full context models by about 50%, and train in 50% fewer steps compared to training models to summarize without belief grading (no BG). After training, ABBEL still uses significantly less memory than the full context setting, as measured by the peak context token length (Peak Tokens).