Tools & products

MLE-bench

1 sources · 1 mentions

Views belong to a particular speaker and passage. Counts describe this reviewed collection, not product ratings or market consensus.

Mention only

Lilian Weng describes MLE-bench as a benchmark of 75 curated Kaggle competitions used to evaluate ML engineering agents, with Kaggle public leaderboards serving as human baselines.

Lilian Weng ·

Supporting evidence

Harness Engineering for Self-Improvement

Original excerpt

MLE-bench : evaluate ML engineering agents on offline Kaggle competitions. Contains 75 ML-engineering competitions curated from Kaggle.
Context

Tests training models, preparing datasets, running experiments, and submitting predictions to grading scripts. Uses Kaggle public leaderboards as human baselines. Best setup in the paper, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level in 16.9% of competitions. Includes resource-scaling and contamination analyses.

About this interpretation

The object-specific interpretation and Chinese translation were checked independently against the source. This is an AI semantic review, not playback verification. Reviewed Oct 5, 2026 · qwen3.8-max-0902

Report an issue
← All mentioned things