← Back to Explore
Agent Evaluation and Benchmarks🔬 emerging

ClawBench

TIGER-AI-Lab/ClawBench
GitHub ↗

Evaluates web agents on 283 real-world tasks across 163 live websites with interception and trace-based scoring.

LanguagePython
SourceREADME.md (Line 381)
#Python#Benchmark

Related Resources in Agent Evaluation and Benchmarks