UN THÈME, DANS SON CONTEXTE

human-AI collaboration performance

Judgments in this source concerning human-AI collaboration performance. Explorez 1 point de vue avec des éléments tirés de 1 source.

0 personnes · 1 sources · 1 opinions exprimées

Contenu mis à jour:

Explorer les liens ↗

Carte des points de vue

0 personnes · 1 sources · 1 opinions exprimées

Reported success with engineer context

The authors report 23.9% success on held-out tasks for their best standalone configuration, Claude Opus 5 with max reasoning in Claude Code. Paired with an engineer with deep context, the same class of model reaches 82.2% on those tasks.

Éléments favorables

Hyper-𝜏-bench: Evaluating agents that build agents

Extrait original

Working alone, our best configuration — Claude Opus 5 (max reasoning) running in Claude Code — passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.

Ces résultats reflètent les sources disponibles, sans constituer une vue exhaustive ou à jour.

Conversations originales1

ARTICLE

Hyper-𝜏-bench: Evaluating agents that build agents

Sierra Blog