BENCHMARK 目录

寻找合适的 benchmark,而非追逐名气。

搜索覆盖推理、代码、多模态、Agent 和生成式媒体的动态目录。在评测前检查任务、指标、版本与可运行示例。

256条 registry benchmark143
LLM
75
VLM
32
Agent
6
AIGC

搜索 benchmark 目录。

检索 · 筛选 · 对比 · 评测

已显示24/ 256 个评测集

A-OKVQAVLM · Visual Question Answering with Knowledge ReasoningAnswer accuracy

A-OKVQA (Augmented OK-VQA) is a benchmark designed to evaluate commonsense reasoning and external world knowledge in visual question answering. It extends beyond basic VQA tasks that rely solely on image content, requiring models to leverage a broad spectru…

查看文档
AA-LCRLLM · Long-Context Question AnsweringAnswer accuracy

AA-LCR (Artificial Analysis Long Context Retrieval) is a benchmark for evaluating long-context retrieval and reasoning capabilities of language models. It requires models to find and synthesize information across multiple documents.

查看文档
ACEBenchAgent · Function calling and agentic tool useAnswer accuracy · Process accuracy

ACEBench evaluates whether large language models can use tools in realistic settings: picking the right API, filling its arguments, pushing back on requests that cannot be satisfied, and driving multi-step agent tasks against a simulated environment. Data i…

查看文档
AGIEvalLLM · Mixed (Multiple-Choice QA + Open-ended Math)Answer accuracy

AGIEval is a human-centric benchmark designed to evaluate foundation models in the context of human cognition and problem-solving. It uses official, standard, and authoritative admission and qualification exams intended for general human test-takers, such a…

查看文档
AI2DVLM · Diagram Understanding and Visual ReasoningAnswer accuracy

AI2D (AI2 Diagrams) is a benchmark dataset for evaluating AI systems' ability to understand and reason about scientific diagrams. It contains over 5,000 diverse diagrams from science textbooks covering topics like the water cycle, food webs, and biological…

查看文档
AIME-2024LLM · Competition Mathematics Problem SolvingAnswer accuracy

AIME 2024 (American Invitational Mathematics Examination 2024) is a benchmark based on problems from the prestigious AIME competition. These problems represent some of the most challenging high school mathematics problems, requiring creative problem-solving…

查看文档
AIME-2025LLM · Competition Mathematics Problem SolvingAnswer accuracy

AIME 2025 (American Invitational Mathematics Examination 2025) is a benchmark based on problems from the prestigious AIME competition, one of the most challenging high school mathematics contests in the United States. It tests advanced mathematical reasonin…

查看文档
AIME-2026LLM · Competition Mathematics Problem SolvingAnswer accuracy

AIME 2026 (American Invitational Mathematics Examination 2026) is a benchmark based on problems from the prestigious AIME competition, one of the most challenging high school mathematics contests in the United States. It tests advanced mathematical reasonin…

查看文档
AIR-Bench-ChatVLM · Open-ended audio question answering.judge score · win rate

AIR-Bench Chat is the generative half of AIR-Bench (Audio InstRuction Benchmark, ACL 2024 main conference) — the first instruction-following benchmark for large audio-language models (LALMs), covering human speech, natural sounds and music. It contains roug…

查看文档
AIR-Bench-FoundationVLM · Single-choice question answering grounded on an audio clip.Answer accuracy

AIR-Bench Foundation is the discriminative half of AIR-Bench (Audio InstRuction Benchmark, ACL 2024 main conference) — the first instruction-following benchmark for large audio-language models (LALMs), covering human speech, natural sounds and music. The Fo…

查看文档
AlpacaEval2.0LLM · Instruction-Following Evaluation (Pairwise Comparison)win rate

AlpacaEval 2.0 is an evaluation framework for instruction-following language models that uses an LLM judge to compare model outputs against a strong baseline. It provides win-rate metrics reflecting human preferences.

查看文档
AMCLLM · Competition Mathematics (Multiple Choice)Answer accuracy

AMC (American Mathematics Competitions) is a benchmark based on problems from the AMC 10/12 competitions from 2022-2024. These multiple-choice problems test mathematical problem-solving skills at the high school level and serve as qualifiers for the AIME co…

查看文档
AnatEMLLM · Biomedical Named Entity Recognition (NER)precision · recall

The AnatEM corpus is an extensive resource for anatomical entity recognition, created by extending and combining previous corpora. It includes over 13,000 annotations across 1,212 biomedical documents, focusing on identifying anatomical structures from subc…

查看文档
ARCLLM · Multiple-Choice Science Question AnsweringAnswer accuracy

ARC (AI2 Reasoning Challenge) is a benchmark designed to evaluate science question answering capabilities of AI models. It consists of multiple-choice science questions from grade 3 to grade 9, divided into an Easy set and a Challenge set based on difficulty.

查看文档
ARC-AGI-2LLM · Abstract Reasoning / Pattern RecognitionAnswer accuracy

ARC-AGI-2 (Abstraction and Reasoning Corpus for Artificial General Intelligence 2) is a benchmark designed to measure an AI system's ability to efficiently acquire new skills on-the-fly, using only a handful of demonstrations. It evaluates abstract reasonin…

查看文档
ARC-Challenge-IndicLLM · Multilingual Multiple-Choice Science Question AnsweringAnswer accuracy

ARC-Challenge-Indic is a translation of the AI2 Reasoning Challenge (ARC-Challenge) science question-answering benchmark into 10 Indic languages, plus the original English set, for evaluating multilingual scientific reasoning.

查看文档
ArenaHardLLM · Competitive Model Evaluation (Arena-style)win rate

ArenaHard is a challenging benchmark that evaluates language models through competitive pairwise comparison. Models are judged against a GPT-4 baseline on difficult tasks requiring reasoning, understanding, and generation capabilities.

查看文档
ArXiv-MathLLM · Research-Level Mathematics Problem SolvingAnswer accuracy

ArXiv-Math is a benchmark of 103 research-level mathematics problems extracted from arXiv preprints. These problems represent cutting-edge mathematical research and test the ability of language models to reason about advanced mathematical concepts at the fr…

查看文档
ArxivRollBenchLLM · Multiple-choice scientific text reasoningAnswer accuracy

ArxivRollBench is a rolling benchmark built from recent arXiv papers. It evaluates whether large language models can reason over fresh scientific text through three task formats: sequencing, cloze, and next-fragment prediction.

查看文档
ArxivRollBench-FullLLM · Multiple-choice scientific text reasoningAnswer accuracy

ArxivRollBench is a rolling benchmark built from recent arXiv papers. It evaluates whether large language models can reason over fresh scientific text through three task formats: sequencing, cloze, and next-fragment prediction.

查看文档
AutomationBenchAgent · Stateful business workflows across 47 simulated SaaS tools.Task pass rate · partial credit

AutomationBench evaluates agents on realistic business workflows across sales, marketing, operations, support, finance, and HR. EvalScope runs the public tasks, simulated SaaS services, and assertion-based scoring provided by Zapier's official Python package.

查看文档
BabyVisionVLM · Visual Perception (Choice + Fill-in-the-blank)Answer accuracy

BabyVision is a visual perception benchmark that evaluates the fundamental visual abilities of multimodal large language models through tasks inspired by infant and early childhood visual development. It focuses on fine-grained discrimination, spatial perce…

查看文档
BBHLLM · Mixed (Multiple-Choice and Free-Form)Answer accuracy

BBH (BIG-Bench Hard) is a subset of 23 challenging tasks from the BIG-Bench benchmark that are specifically selected because language models initially struggled with them. These tasks require complex reasoning abilities that benefit from Chain-of-Thought (C…

查看文档
BC2GMLLM · Biomedical Named Entity Recognition (NER)precision · recall

The BC2GM (BioCreative II Gene Mention) dataset is a widely used corpus for gene mention recognition, consisting of 20,000 sentences from MEDLINE abstracts where gene and protein names have been manually annotated by domain experts.

查看文档

Metadata 即证据

每条目录项链接到生成的文档,包含任务类型、模态和指标。缺少元数据文件会使网站检查失败。

简洁地参与贡献

新增适配器,声明 BenchmarkMeta,运行小范围冒烟测试,再生成文档。也可以让 AI 编程助手按项目 Skill 完成周边流程。

打开完整使用指南

持续扩展,也不丢失可信度。

开放贡献 · 质量优先

1

Adapter

实现标准接口。

2

Smoke test

验证小范围可运行路径。

3

Metadata

声明版本与指标。

4

Docs

生成双语详情。

5

Catalog

使其可发现。

浅色 EvalScope Dashboard:模型、基准与指标历史
当模型、基准与指标历史持续可审查时,贡献才真正产生价值。