Agent Evaluation and Benchmarks🌱 growing
AgentBenchTHUDM/AgentBench
Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Python#Benchmark
GitHub ↗Benchmark for evaluating LLMs on real-world software engineering tasks from GitHub issues.
Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Frontier benchmark for measuring general intelligence capabilities in AI agents beyond pattern matching.
Evaluates web agents on 283 real-world tasks across 163 live websites with interception and trace-based scoring.