UN THÈME, DANS SON CONTEXTE

RE-Bench AI agent evaluation methodology

Judgments in this source concerning RE-Bench AI agent evaluation methodology. Explorez 1 point de vue avec des éléments tirés de 1 source.

1 personnes · 1 sources · 1 opinions exprimées

Contenu mis à jour:

Explorer les liens ↗

Carte des points de vue

Explorer par personne. Sélectionnez deux ou trois personnes pour les comparer.

1 personnes · 1 sources · 1 opinions exprimées

Lilian Weng

RE-Bench evaluates AI agents on realistic ML engineering tasks, using human performance as a benchmark

RE-Bench evaluates frontier AI agents on 7 open-ended ML research-engineering environments, each defined by a scoring function, starting solution, and reference solution, runnable with 8 or fewer H100 GPUs. It includes data from 71 eight-hour attempts by 61 distinct human experts, where humans achieved non-zero scores in 82% of attempts and matched or exceeded strong reference solutions in 24%.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments.
Contexte

Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions. Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.

Partager un aperçuVérifier cette affirmation

Il s’agit de points de vue individuels, non d’une mesure du consensus. Le matériel source reste dans sa langue d’origine.