Agent Evaluation and Benchmarks🌱 growing
AgentBenchTHUDM/AgentBench
Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Python#Benchmark
GitHub ↗Evaluates web agents on 283 real-world tasks across 163 live websites with interception and trace-based scoring.
Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Frontier benchmark for measuring general intelligence capabilities in AI agents beyond pattern matching.
Benchmark for General AI Assistants measuring real-world reasoning and tool use.