用于自我改进的 Harness 工程
MLE-bench:在离线 Kaggle 竞赛上评测机器学习工程智能体。该基准包含从 Kaggle 平台精选的 75 场机器学习工程类竞赛。
原始摘录
MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
上下文
评测内容涵盖模型训练、数据集准备、实验运行及向评分脚本提交预测结果等环节。采用 Kaggle 公开排行榜作为人类基线。论文中表现最佳的配置(o1-preview 模型搭配 AIDE 框架)在 16.9% 的竞赛中达到 Kaggle 铜牌水平。该基准同时包含资源扩展性分析与污染分析。
原始上下文
Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.