Search

Search IconIcon to open search

Agentic Harness Benchmarks

Last updatedUpdated: by Jakub Žovák · 1 min read

Properties
created 26.04.2026, 21:17
modified 27.04.2026, 08:01
published Empty
sources Empty
topics Agent Design , Agent Harness
authors Jakub
ai-assisted No
  • This note aggregates information about benchmarks targeted specifically at agentic harnesses

# List

  • AlphaEval
    • Production-grounded benchmark evaluating agents as complete products (tool use, orchestration, UI) across 94 tasks from 7 companies; best result is Claude Code + Opus 4.6 at 64.41/100
    • Benchmarks the whole harness
    • TODO: Read the paper!