Top Agent Evaluation & Benchmarking Tools (2026)
You cannot improve an agent you cannot measure. This list splits into two kinds. Benchmarks are fixed datasets and leaderboards, like SWE-bench. Eval frameworks are tools to run your own evals, like DeepEval or Promptfoo. Durable facts only. Benchmark scores are left off on purpose, since they move all the time.
An open-source (MIT) benchmark. It scores models on fixing real GitHub issues. The model writes a patch that must pass the repo's tests. Runs happen in Docker so they can be repeated.
- Type:
- Benchmark
- License:
- MIT
An open-source (Apache-2.0) benchmark. It tests agents on end-to-end terminal tasks. It pairs a task dataset with a harness. The harness connects a model to a sandboxed terminal.
- Type:
- Benchmark
- License:
- Apache-2.0
An open-source (MIT) benchmark and eval harness for code generation. It checks that generated code passes unit tests. It uses the pass@k metric.
- Type:
- Benchmark
- License:
- MIT
An open-source (Apache-2.0) eval framework, pitched as Pytest for LLM apps. Metrics run locally. An optional Confident AI cloud backs it.
- Type:
- Framework
- License:
- Apache-2.0
An open-source (MIT) CLI and library for testing and red-teaming LLM apps. It runs fully on your machine, so prompts never leave it.
- Type:
- Framework
- License:
- MIT
An open-source (MIT) framework for testing LLMs and systems built on them. It ships an open registry of benchmarks. It supports custom evals.
- Type:
- Framework
- License:
- MIT
An open-source (Apache-2.0) eval framework for LLM apps. It focuses on RAG systems. It gives you metrics and can build test sets for you.
- Type:
- Framework
- License:
- Apache-2.0
An open-source (MIT) framework for LLM evals from the UK AI Security Institute. It has built-in parts for prompting, tool use, and multi-turn dialog. It also does model-graded scoring.
- Type:
- Framework
- License:
- MIT
Reviewed quarterly. This is the fastest-rotting list here. Leaderboards and top scores change weekly. So no scores appear on this page at all. The durable fields (type, license, repo) are stable. Each row carries its last-checked date and source.
Common questions
- What is the difference between a benchmark and an eval framework?
- A benchmark is a fixed dataset with a scoring rule, and usually a public leaderboard. SWE-bench and Terminal-Bench are examples. It is used to compare models on the same task. An eval framework, like DeepEval or Promptfoo, is a tool. You use it to write and run your own evals on your own app and data.
- How do you evaluate a coding agent specifically?
- For coding, the strongest signal is task resolution. Does the agent's change pass the repo's tests? SWE-bench and Terminal-Bench encode that. For your own app, use an eval framework. It lets you define pass/fail checks on real tasks instead of leaning on a public benchmark's mix.