CONNECTED THINKING

Knowledge atlas

Follow people, viewpoints and their original evidence.

1 people · 1 sources · 1 viewpoints

IN CONTEXT

MLE-bench offline Kaggle benchmark

Choose a viewpoint. Follow it back to the conversation.

MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

IN CONTEXTMLE-bench offline Kaggle benchmark1 viewpoints
2026-07-04

Equal sectors are reading positions, not rankings.

Showing 1–1 of 1 viewpoints · Newest sources first

1 / 1

Selected viewpoint

MLE-bench uses 75 Kaggle competitions as offline ML engineering benchmarks

MLE-bench evaluates ML engineering agents on 75 curated offline Kaggle competitions, testing model training, dataset preparation, experiment execution, and submission to grading scripts. Kaggle public leaderboards serve as human baselines. The best-performing setup—o1-preview with AIDE scaffolding—reached at least Kaggle bronze-medal level in 16.9% of competitions.

These are individual perspectives, not a measure of consensus. Source material stays in its original language.

Supporting evidence

Harness Engineering for Self-Improvement

Original excerpt

MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Context

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

Publication dates describe the sources, not changes in belief.