BENCHMARK CATALOG

Find benchmark fit, not benchmark fame.

Search a living registry across reasoning, coding, multimodal, agents and generative media. Inspect tasks, metrics, versions and runnable examples before you evaluate.

256benchmarks in registry143
LLM
75
VLM
32
Agent
6
AIGC

Search the benchmark catalog.

Find · filter · compare · evaluate

Showing24/ 256 benchmarks

A-OKVQAVLM · Visual Question Answering with Knowledge ReasoningAnswer accuracy

A-OKVQA (Augmented OK-VQA) is a benchmark designed to evaluate commonsense reasoning and external world knowledge in visual question answering. It extends beyond basic VQA tasks that rely solely on image content, requiring models to leverage a broad spectru…

Open documentation
AA-LCRLLM · Long-Context Question AnsweringAnswer accuracy

AA-LCR (Artificial Analysis Long Context Retrieval) is a benchmark for evaluating long-context retrieval and reasoning capabilities of language models. It requires models to find and synthesize information across multiple documents.

Open documentation
ACEBenchAgent · Function calling and agentic tool useAnswer accuracy · Process accuracy

ACEBench evaluates whether large language models can use tools in realistic settings: picking the right API, filling its arguments, pushing back on requests that cannot be satisfied, and driving multi-step agent tasks against a simulated environment. Data i…

Open documentation
AGIEvalLLM · Mixed (Multiple-Choice QA + Open-ended Math)Answer accuracy

AGIEval is a human-centric benchmark designed to evaluate foundation models in the context of human cognition and problem-solving. It uses official, standard, and authoritative admission and qualification exams intended for general human test-takers, such a…

Open documentation
AI2DVLM · Diagram Understanding and Visual ReasoningAnswer accuracy

AI2D (AI2 Diagrams) is a benchmark dataset for evaluating AI systems' ability to understand and reason about scientific diagrams. It contains over 5,000 diverse diagrams from science textbooks covering topics like the water cycle, food webs, and biological…

Open documentation
AIME-2024LLM · Competition Mathematics Problem SolvingAnswer accuracy

AIME 2024 (American Invitational Mathematics Examination 2024) is a benchmark based on problems from the prestigious AIME competition. These problems represent some of the most challenging high school mathematics problems, requiring creative problem-solving…

Open documentation
AIME-2025LLM · Competition Mathematics Problem SolvingAnswer accuracy

AIME 2025 (American Invitational Mathematics Examination 2025) is a benchmark based on problems from the prestigious AIME competition, one of the most challenging high school mathematics contests in the United States. It tests advanced mathematical reasonin…

Open documentation
AIME-2026LLM · Competition Mathematics Problem SolvingAnswer accuracy

AIME 2026 (American Invitational Mathematics Examination 2026) is a benchmark based on problems from the prestigious AIME competition, one of the most challenging high school mathematics contests in the United States. It tests advanced mathematical reasonin…

Open documentation
AIR-Bench-ChatVLM · Open-ended audio question answering.judge score · win rate

AIR-Bench Chat is the generative half of AIR-Bench (Audio InstRuction Benchmark, ACL 2024 main conference) — the first instruction-following benchmark for large audio-language models (LALMs), covering human speech, natural sounds and music. It contains roug…

Open documentation
AIR-Bench-FoundationVLM · Single-choice question answering grounded on an audio clip.Answer accuracy

AIR-Bench Foundation is the discriminative half of AIR-Bench (Audio InstRuction Benchmark, ACL 2024 main conference) — the first instruction-following benchmark for large audio-language models (LALMs), covering human speech, natural sounds and music. The Fo…

Open documentation
AlpacaEval2.0LLM · Instruction-Following Evaluation (Pairwise Comparison)win rate

AlpacaEval 2.0 is an evaluation framework for instruction-following language models that uses an LLM judge to compare model outputs against a strong baseline. It provides win-rate metrics reflecting human preferences.

Open documentation
AMCLLM · Competition Mathematics (Multiple Choice)Answer accuracy

AMC (American Mathematics Competitions) is a benchmark based on problems from the AMC 10/12 competitions from 2022-2024. These multiple-choice problems test mathematical problem-solving skills at the high school level and serve as qualifiers for the AIME co…

Open documentation
AnatEMLLM · Biomedical Named Entity Recognition (NER)precision · recall

The AnatEM corpus is an extensive resource for anatomical entity recognition, created by extending and combining previous corpora. It includes over 13,000 annotations across 1,212 biomedical documents, focusing on identifying anatomical structures from subc…

Open documentation
ARCLLM · Multiple-Choice Science Question AnsweringAnswer accuracy

ARC (AI2 Reasoning Challenge) is a benchmark designed to evaluate science question answering capabilities of AI models. It consists of multiple-choice science questions from grade 3 to grade 9, divided into an Easy set and a Challenge set based on difficulty.

Open documentation
ARC-AGI-2LLM · Abstract Reasoning / Pattern RecognitionAnswer accuracy

ARC-AGI-2 (Abstraction and Reasoning Corpus for Artificial General Intelligence 2) is a benchmark designed to measure an AI system's ability to efficiently acquire new skills on-the-fly, using only a handful of demonstrations. It evaluates abstract reasonin…

Open documentation
ARC-Challenge-IndicLLM · Multilingual Multiple-Choice Science Question AnsweringAnswer accuracy

ARC-Challenge-Indic is a translation of the AI2 Reasoning Challenge (ARC-Challenge) science question-answering benchmark into 10 Indic languages, plus the original English set, for evaluating multilingual scientific reasoning.

Open documentation
ArenaHardLLM · Competitive Model Evaluation (Arena-style)win rate

ArenaHard is a challenging benchmark that evaluates language models through competitive pairwise comparison. Models are judged against a GPT-4 baseline on difficult tasks requiring reasoning, understanding, and generation capabilities.

Open documentation
ArXiv-MathLLM · Research-Level Mathematics Problem SolvingAnswer accuracy

ArXiv-Math is a benchmark of 103 research-level mathematics problems extracted from arXiv preprints. These problems represent cutting-edge mathematical research and test the ability of language models to reason about advanced mathematical concepts at the fr…

Open documentation
ArxivRollBenchLLM · Multiple-choice scientific text reasoningAnswer accuracy

ArxivRollBench is a rolling benchmark built from recent arXiv papers. It evaluates whether large language models can reason over fresh scientific text through three task formats: sequencing, cloze, and next-fragment prediction.

Open documentation
ArxivRollBench-FullLLM · Multiple-choice scientific text reasoningAnswer accuracy

ArxivRollBench is a rolling benchmark built from recent arXiv papers. It evaluates whether large language models can reason over fresh scientific text through three task formats: sequencing, cloze, and next-fragment prediction.

Open documentation
AutomationBenchAgent · Stateful business workflows across 47 simulated SaaS tools.Task pass rate · partial credit

AutomationBench evaluates agents on realistic business workflows across sales, marketing, operations, support, finance, and HR. EvalScope runs the public tasks, simulated SaaS services, and assertion-based scoring provided by Zapier's official Python package.

Open documentation
BabyVisionVLM · Visual Perception (Choice + Fill-in-the-blank)Answer accuracy

BabyVision is a visual perception benchmark that evaluates the fundamental visual abilities of multimodal large language models through tasks inspired by infant and early childhood visual development. It focuses on fine-grained discrimination, spatial perce…

Open documentation
BBHLLM · Mixed (Multiple-Choice and Free-Form)Answer accuracy

BBH (BIG-Bench Hard) is a subset of 23 challenging tasks from the BIG-Bench benchmark that are specifically selected because language models initially struggled with them. These tasks require complex reasoning abilities that benefit from Chain-of-Thought (C…

Open documentation
BC2GMLLM · Biomedical Named Entity Recognition (NER)precision · recall

The BC2GM (BioCreative II Gene Mention) dataset is a widely used corpus for gene mention recognition, consisting of 20,000 sentences from MEDLINE abstracts where gene and protein names have been manually annotated by domain experts.

Open documentation

Metadata is evidence

Each catalog entry routes to generated documentation and carries task type, modality and metrics. Missing metadata causes the site check to fail.

Contribute without ceremony

Add an adapter, declare BenchmarkMeta, run its narrow smoke test, then generate docs. Ask an AI agent to follow the project skill for the surrounding workflow.

Open full setup guide

Built to grow without losing trust.

Open contributions · quality first

1

Adapter

Implement a standard interface.

2

Smoke test

Verify a narrow runnable path.

3

Metadata

Declare version and metrics.

4

Docs

Generate bilingual details.

5

Catalog

Make it discoverable.

Light EvalScope Dashboard with model, benchmark and metric history
A contribution becomes useful when model, benchmark and metric history stay inspectable.