基于重建的信念分级改善了所报告的训练对比结果
我们看到,借助通用的基于重建的信念分级函数,我们将与完整上下文模型的性能差距缩小了约 50%,并且与训练模型进行无信念分级(no BG)摘要相比,训练步数减少了 50%。训练完成后,以峰值上下文字元长度(Peak Tokens)衡量,ABBEL 的内存占用仍显著低于完整上下文设置。
原始摘录
We see that with the general reconstruction-based belief grading function we reduce the performance gap from full context models by about 50%, and train in 50% fewer steps compared to training models to summarize without belief grading (no BG). After training, ABBEL still uses significantly less memory than the full context setting, as measured by the peak context token length (Peak Tokens).