UN THÈME, DANS SON CONTEXTE

Human–AI performance tradeoffs in RE-Bench

Judgments in this source concerning Human–AI performance tradeoffs in RE-Bench. Explorez 1 point de vue avec des éléments tirés de 1 source.

1 personnes · 1 sources · 1 opinions exprimées

Contenu mis à jour:

Explorer les liens ↗

Carte des points de vue

Explorer par personne. Sélectionnez deux ou trois personnes pour les comparer.

1 personnes · 1 sources · 1 opinions exprimées

Lilian Weng

AI agents outperform humans at short time budgets but not longer ones

In RE-Bench, the best AI agents scored 4× higher than human experts at a 2-hour budget, yet humans demonstrated better returns to extended effort and surpassed agents at both 8-hour and 32-hour time budgets.

Éléments favorables

Harness Engineering for Self-Improvement

Extrait original

Best AI agents scored 4× higher than humans at a 2-hour budget, but humans had better returns to longer budgets and exceeded agents at 8-hour and 32-hour settings.
Contexte

RE-Bench : evaluate frontier AI agents on realistic ML research-engineering envs against human experts. 7 challenging, open-ended ML research-engineering environments. Each environment = (scoring function, starting solution, reference solution); each can be run with 8 or fewer H100 GPUs. Examples: optimize a kernel, run a scaling-law experiment, fix an embedding, fine-tune GPT-2 for QA, etc. Includes data from 71 eight-hour attempts by 61 distinct human experts. Human experts achieved non-zero score in 82% of 8-hour attempts; 24% matched or exceeded strong reference solutions.

Partager un aperçuVérifier cette affirmation

Il s’agit de points de vue individuels, non d’une mesure du consensus. Le matériel source reste dans sa langue d’origine.