← Back to Explore
Agent Evaluation and Benchmarks🌱 growing

AgentBench

THUDM/AgentBench
GitHub ↗

Comprehensive benchmark for evaluating LLMs as agents across 8 distinct environments.

LanguagePython
SourceREADME.md (Line 379)
#Python#Benchmark

Related Resources in Agent Evaluation and Benchmarks