← Back to Explore
Agent Evaluation and Benchmarks🌱 growing

GAIA Benchmark

Website ↗

Benchmark for General AI Assistants measuring real-world reasoning and tool use.

LanguagePython
SourceREADME.md (Line 382)
#Python#Benchmark

Related Resources in Agent Evaluation and Benchmarks