← Back to Explore
Agent Evaluation and Benchmarks🌱 growing

Inspect AI

UKGovernmentBEIS/inspect_ai
GitHub ↗

Framework for evaluating large language models with composable tasks and scoring.

LanguagePython
SourceREADME.md (Line 383)
#Python#Evaluation

Related Resources in Agent Evaluation and Benchmarks