Tools & products

RE-Bench

1 sources · 1 mentions

Views belong to a particular speaker and passage. Counts describe this reviewed collection, not product ratings or market consensus.

Mention only

Lilian Weng introduces RE-Bench as a benchmark evaluating frontier AI agents across 7 open-ended ML research-engineering environments, with human expert performance data included for comparison.

Lilian Weng ·

Supporting evidence

Harness Engineering for Self-Improvement

Original excerpt

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
Context

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

About this interpretation

The object-specific interpretation and Chinese translation were checked independently against the source. This is an AI semantic review, not playback verification. Reviewed Oct 5, 2026 · qwen3.8-max-0902

Report an issue
← All mentioned things