Agent Benchmarks
Properties
created
12.03.2025, 16:18
modified
26.07.2026, 13:33
published
Empty
sources
Empty
topics
Agent Benchmark
authors
Jakub
ai-assisted
No
This page aggregates benchmarks specifically designed for evaluating agents based on the LLMs. It differs from LLM Benchmarks page which focuses on benchmarks designed specifically for LLMs, even though there might be some overlap between these benchmarks.
# Sub-Hubs
- Agent Memory Benchmarks
- LoCoMo, LongMemEval, BEAM - multi-session memory evaluation distinct from long-context attention
# List of Benchmarks
- ARC-AGI-3
- Interactive agentic benchmark requiring goal inference, environment modeling, and action planning without explicit instructions; frontier AI scores <1% vs 100% human
- Paper:
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
- Submitted on 2026-03-24
- A Benchmark for Conversational Data Retrieval
- WebArena
- Agentic benchmark based on operating in a web environment
- Paper:
WebArena: A Realistic Web Environment for Building Autonomous Agents
- 317 citations
- Submitted on 2023-7-25
- AgentBench
- A multi-dimensional evolving benchmark that currently consists of 8 distinct environments to assess LLM-as-Agent’s reasoning and decision-making abilities in a multi-turn open-ended generation setting
- Paper:
AgentBench: Evaluating LLMs as Agents
- 146 citations
- Submitted on 7 Aug 2023
- GAIA2
- We introduce Gaia2, a benchmark for evaluating large language model agents in realistic, asynchronous environments. Unlike prior static or synchronous evaluations, Gaia2 introduces scenarios where environments evolve independently of agent actions, requiring agents to operate under temporal constraints, adapt to noisy and dynamic events, resolve ambiguity, and collaborate with other agents.
- Paper:
GAIA2: BENCHMARKING LLM AGENTS ON DYNAMIC AND ASYNCHRONOUS ENVIRONMENTS
- 2026-02-12
- 7 citations
- GAIA
- A benchmark for General AI Assistants that, if solved, would represent a milestone in AI research. GAIA proposes real-world questions that require a set of fundamental abilities such as reasoning, multi-modality handling, web browsing, and generally tool-use proficiency.
- Public Leaderboard
- Paper:
GAIA: a benchmark for General AI Assistants
- 100 citations
- Submitted on 21 Nov 2023
- τ2-bench
- Paper:
τ2 -Bench: Evaluating Conversational Agents in a Dual-Control Environment
- 93 citations
- Submitted on 9 Jun 2025
- Paper:
τ2 -Bench: Evaluating Conversational Agents in a Dual-Control Environment
- τ-bench
- A benchmark emulating dynamic conversations between a user (simulated by language models) and a language agent provided with domain-specific API tools and policy guidelines
- Paper:
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- 20 citations
- Submitted on 17 Jun 2024
- AppWorld
- Controllable simulation of 9 day-to-day apps (457 APIs, ~100 simulated users) with 750 interactive coding tasks; ACL'24 Best Resource Paper
- Paper:
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- 210 citations
- Submitted on 2024-07-26