GAIA
General AI Assistants
Benchmark for multi-hop reasoning and real-world tool use. View Guide
AgentV provides first-class support for the world’s most rigorous academic and industrial benchmarks. By using native URI schemes, researchers can execute standardized evaluations without manual data munging or environment setup.
GAIA
General AI Assistants
Benchmark for multi-hop reasoning and real-world tool use. View Guide
AssistantBench
Enterprise Intelligence
Specialized for business reasoning and industrial accuracy. View Guide
OpenCompass
Coming Soon
Native support for OpenCompass integration is currently in the v1.5.x roadmap.
AgentV abstracts benchmark loading behind the Unified Benchmark URI Schema. This allows researchers to target specific datasets, versions, and difficultly tiers directly from the CLI.
# General syntaxagentv evaluate --path <benchmark_scheme>://<version> --agent <agent_url>
# Examplesagentv evaluate --path gaia://2025_validation --agent http://localhost:5001/execute_taskagentv evaluate --path assistantbench://v1_standard --agent http://localhost:5001/execute_taskCorrect Answer) to high-fidelity AES Success Criteria.gaia://2023), researchers ensure they are always evaluating against a specific, immutable dataset snapshot.