Search

Search IconIcon to open search

Agent Benchmarks

Last updatedUpdated: by Jakub Žovák · 2 min read

Properties
created 12.03.2025, 16:18
modified 26.07.2026, 13:33
published Empty
sources Empty
topics Agent Benchmark
authors Jakub
ai-assisted No

This page aggregates benchmarks specifically designed for evaluating agents based on the LLMs. It differs from LLM Benchmarks page which focuses on benchmarks designed specifically for LLMs, even though there might be some overlap between these benchmarks.

# Sub-Hubs

  • Agent Memory Benchmarks
    • LoCoMo, LongMemEval, BEAM - multi-session memory evaluation distinct from long-context attention

# List of Benchmarks