Agent Evaluation and Benchmarks🌱 growing
AgentBenchTHUDM/AgentBench
Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Python#Benchmark
GitHub ↗Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.
Frontier benchmark for measuring general intelligence capabilities in AI agents beyond pattern matching.
Evaluates web agents on 283 real-world tasks across 163 live websites with interception and trace-based scoring.
Benchmark for General AI Assistants measuring real-world reasoning and tool use.
Framework for evaluating large language models with composable tasks and scoring.
Benchmark for evaluating LLMs on real-world software engineering tasks from GitHub issues.
Benchmark for web agent evaluation using real websites with realistic task completion metrics.