A TOPIC, IN CONTEXT

human-AI collaboration performance

Judgments in this source concerning human-AI collaboration performance. Explore 1 viewpoint with evidence from 1 source.

0 people · 1 sources · 1 viewpoints

Content updated:

Explore connections ↗

Viewpoint map

0 people · 1 sources · 1 viewpoints

Reported success with engineer context

The authors report 23.9% success on held-out tasks for their best standalone configuration, Claude Opus 5 with max reasoning in Claude Code. Paired with an engineer with deep context, the same class of model reaches 82.2% on those tasks.

Supporting evidence

Hyper-𝜏-bench: Evaluating agents that build agents

Original excerpt

Working alone, our best configuration — Claude Opus 5 (max reasoning) running in Claude Code — passes just 23.9% of the held-out evaluation tasks. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.

These findings reflect the available sources, not an exhaustive or current view.

Original conversations1

ARTICLE

Hyper-𝜏-bench: Evaluating agents that build agents

Sierra Blog