Agentic Harness Benchmarks
Properties
created
26.04.2026, 21:17
modified
27.04.2026, 08:01
published
Empty
sources
Empty
topics
Agent Design , Agent Harness
authors
Jakub
ai-assisted
No
- This note aggregates information about benchmarks targeted specifically at agentic harnesses
# List
- AlphaEval
- Production-grounded benchmark evaluating agents as complete products (tool use, orchestration, UI) across 94 tasks from 7 companies; best result is Claude Code + Opus 4.6 at 64.41/100
- Benchmarks the whole harness
- TODO: Read the paper!