UN TEMA, EN CONTEXTO

human-AI collaboration performance

Judgments in this source concerning human-AI collaboration performance. Explora 1 punto de vista con evidencias de 1 fuente.

0 personas · 1 fuentes · 1 opiniones expresadas

Contenido actualizado:

Explorar conexiones ↗

Mapa de perspectivas

0 personas · 1 fuentes · 1 opiniones expresadas

Reported success with engineer context

The authors report 23.9% success on held-out tasks for their best standalone configuration, Claude Opus 5 with max reasoning in Claude Code. Paired with an engineer with deep context, the same class of model reaches 82.2% on those tasks.

Evidencia a favor

Hyper-𝜏-bench: Evaluating agents that build agents

Extracto original

Working alone, our best configuration — Claude Opus 5 (max reasoning) running in Claude Code — passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.

Estos resultados reflejan las fuentes disponibles, no una visión exhaustiva ni necesariamente actual.

Conversaciones originales1

ARTÍCULO

Hyper-𝜏-bench: Evaluating agents that build agents

Sierra Blog