Lilian Weng describes MLE-bench as a benchmark of 75 curated Kaggle competitions used to evaluate ML engineering agents, with Kaggle public leaderboards serving as human baselines.
Lilian Weng ·
Supporting evidence
Original excerpt
MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Context
Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.
About this interpretation
The object-specific interpretation and Chinese translation were checked independently against the source. This is an AI semantic review, not playback verification. Reviewed Oct 5, 2026 · qwen3.8-max-0902
Report an issue