Agent Evaluation and Benchmarks🌱 growing
AgentBenchTHUDM/AgentBench
Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Python#Benchmark
GitHub ↗Framework for evaluating large language models with composable tasks and scoring.
Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Frontier benchmark for measuring general intelligence capabilities in AI agents beyond pattern matching.
Evaluates web agents on 283 real-world tasks across 163 live websites with interception and trace-based scoring.