Evaluation
Properties
created
Empty
modified
25.06.2026, 09:54
published
Empty
sources
Empty
topics
Evaluation
authors
Empty
ai-assisted
No
A note dedicated to all the aspects of the evaluation of the AI applications
# Topics
- Agent Evaluation
- Evaluation Libraries
- Libraries that contain metrics definitions that can be used to evaluate LLM outputs
- User Simulation
- Experiment Runners
- LLM Metrics - metrics use to evaluate the output of the LLM
- IR Metrics - metrics to evaluate retrieval
- Datasets for AI evaluation:
- RAG-Instruct-Benchmark-Tester - https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester
- LLM Benchmark that uses LLM-as-a-Judge
- Companies with AI evaluation products
- Galileo - https://www.rungalileo.io/
- Humanloop - https://humanloop.com/
- ComposoAI - https://www.composo.ai/s
# Resources
- Evaluation of AI Applications In Practice
- 2025-08-06 — Practical guide to evaluating LLM-backed applications: dataset construction, metrics selection, and infrastructure for iterative tuning
- How to evaluate AI applications
- Gaia2 and ARE: Empowering the Community to Evaluate Agents
- Understanding the 4 Main Approaches to LLM Evaluation
- Article about evaluating LLMs themselves and not the applications built on top of them