Metadata is evidence
Each catalog entry routes to generated documentation and carries task type, modality and metrics. Missing metadata causes the site check to fail.
BENCHMARK CATALOG
Search a living registry across reasoning, coding, multimodal, agents and generative media. Inspect tasks, metrics, versions and runnable examples before you evaluate.
Find · filter · compare · evaluate
Showing24/ 256 benchmarks
A-OKVQA (Augmented OK-VQA) is a benchmark designed to evaluate commonsense reasoning and external world knowledge in visual question answering. It extends beyond basic VQA tasks that rely solely on image content, requiring models to leverage a broad spectru…
Open documentationAA-LCR (Artificial Analysis Long Context Retrieval) is a benchmark for evaluating long-context retrieval and reasoning capabilities of language models. It requires models to find and synthesize information across multiple documents.
Open documentationACEBench evaluates whether large language models can use tools in realistic settings: picking the right API, filling its arguments, pushing back on requests that cannot be satisfied, and driving multi-step agent tasks against a simulated environment. Data i…
Open documentationAGIEval is a human-centric benchmark designed to evaluate foundation models in the context of human cognition and problem-solving. It uses official, standard, and authoritative admission and qualification exams intended for general human test-takers, such a…
Open documentationAI2D (AI2 Diagrams) is a benchmark dataset for evaluating AI systems' ability to understand and reason about scientific diagrams. It contains over 5,000 diverse diagrams from science textbooks covering topics like the water cycle, food webs, and biological…
Open documentationAIME 2024 (American Invitational Mathematics Examination 2024) is a benchmark based on problems from the prestigious AIME competition. These problems represent some of the most challenging high school mathematics problems, requiring creative problem-solving…
Open documentationAIME 2025 (American Invitational Mathematics Examination 2025) is a benchmark based on problems from the prestigious AIME competition, one of the most challenging high school mathematics contests in the United States. It tests advanced mathematical reasonin…
Open documentationAIME 2026 (American Invitational Mathematics Examination 2026) is a benchmark based on problems from the prestigious AIME competition, one of the most challenging high school mathematics contests in the United States. It tests advanced mathematical reasonin…
Open documentationAIR-Bench Chat is the generative half of AIR-Bench (Audio InstRuction Benchmark, ACL 2024 main conference) — the first instruction-following benchmark for large audio-language models (LALMs), covering human speech, natural sounds and music. It contains roug…
Open documentationAIR-Bench Foundation is the discriminative half of AIR-Bench (Audio InstRuction Benchmark, ACL 2024 main conference) — the first instruction-following benchmark for large audio-language models (LALMs), covering human speech, natural sounds and music. The Fo…
Open documentationAlpacaEval 2.0 is an evaluation framework for instruction-following language models that uses an LLM judge to compare model outputs against a strong baseline. It provides win-rate metrics reflecting human preferences.
Open documentationAMC (American Mathematics Competitions) is a benchmark based on problems from the AMC 10/12 competitions from 2022-2024. These multiple-choice problems test mathematical problem-solving skills at the high school level and serve as qualifiers for the AIME co…
Open documentationThe AnatEM corpus is an extensive resource for anatomical entity recognition, created by extending and combining previous corpora. It includes over 13,000 annotations across 1,212 biomedical documents, focusing on identifying anatomical structures from subc…
Open documentationARC (AI2 Reasoning Challenge) is a benchmark designed to evaluate science question answering capabilities of AI models. It consists of multiple-choice science questions from grade 3 to grade 9, divided into an Easy set and a Challenge set based on difficulty.
Open documentationARC-AGI-2 (Abstraction and Reasoning Corpus for Artificial General Intelligence 2) is a benchmark designed to measure an AI system's ability to efficiently acquire new skills on-the-fly, using only a handful of demonstrations. It evaluates abstract reasonin…
Open documentationARC-Challenge-Indic is a translation of the AI2 Reasoning Challenge (ARC-Challenge) science question-answering benchmark into 10 Indic languages, plus the original English set, for evaluating multilingual scientific reasoning.
Open documentationArenaHard is a challenging benchmark that evaluates language models through competitive pairwise comparison. Models are judged against a GPT-4 baseline on difficult tasks requiring reasoning, understanding, and generation capabilities.
Open documentationArXiv-Math is a benchmark of 103 research-level mathematics problems extracted from arXiv preprints. These problems represent cutting-edge mathematical research and test the ability of language models to reason about advanced mathematical concepts at the fr…
Open documentationArxivRollBench is a rolling benchmark built from recent arXiv papers. It evaluates whether large language models can reason over fresh scientific text through three task formats: sequencing, cloze, and next-fragment prediction.
Open documentationArxivRollBench is a rolling benchmark built from recent arXiv papers. It evaluates whether large language models can reason over fresh scientific text through three task formats: sequencing, cloze, and next-fragment prediction.
Open documentationAutomationBench evaluates agents on realistic business workflows across sales, marketing, operations, support, finance, and HR. EvalScope runs the public tasks, simulated SaaS services, and assertion-based scoring provided by Zapier's official Python package.
Open documentationBabyVision is a visual perception benchmark that evaluates the fundamental visual abilities of multimodal large language models through tasks inspired by infant and early childhood visual development. It focuses on fine-grained discrimination, spatial perce…
Open documentationBBH (BIG-Bench Hard) is a subset of 23 challenging tasks from the BIG-Bench benchmark that are specifically selected because language models initially struggled with them. These tasks require complex reasoning abilities that benefit from Chain-of-Thought (C…
Open documentationThe BC2GM (BioCreative II Gene Mention) dataset is a widely used corpus for gene mention recognition, consisting of 20,000 sentences from MEDLINE abstracts where gene and protein names have been manually annotated by domain experts.
Open documentationThe BC4CHEMD (BioCreative IV CHEMDNER) dataset is a corpus of 10,000 PubMed abstracts with 84,355 chemical entity mentions manually annotated by experts for chemical named entity recognition.
Open documentationThe BC5CDR corpus is a manually annotated resource of 1,500 PubMed articles developed for the BioCreative V challenge, containing over 4,400 chemical mentions, 5,800 disease mentions, and 3,100 chemical-disease interactions.
Open documentationBFCL (Berkeley Function Calling Leaderboard) v3 is the first comprehensive and executable function call evaluation benchmark for assessing LLMs' ability to invoke functions. It evaluates various forms of function calls, diverse scenarios, and executability.
Open documentationBFCL-v4 (Berkeley Function-Calling Leaderboard V4) is a comprehensive benchmark for evaluating agentic function-calling capabilities of LLMs. It tests web search, memory operations, and format sensitivity as building blocks for agentic applications.
Open documentationBhashaBench-Multi (Ayurveda) is a domain-specific multiple-choice benchmark evaluating LLM knowledge of Ayurvedic medicine across 22 Indic languages. Each question originates in English and is machine translated (with LLM-judged translation quality scores)…
Open documentationBhashaBench-Multi (Finance) is a domain-specific multiple-choice benchmark evaluating LLM knowledge of finance across 22 Indic languages. Each question originates in English and is machine translated (with LLM-judged translation quality scores) into the tar…
Open documentationBhashaBench-Multi (Krishi) is a domain-specific multiple-choice benchmark evaluating LLM knowledge of agriculture (Krishi) across 22 Indic languages. Each question originates in English and is machine translated (with LLM-judged translation quality scores)…
Open documentationBhashaBench-Multi (Legal) is a domain-specific multiple-choice benchmark evaluating LLM knowledge of Indian law across 22 Indic languages. Each question originates in English and is machine translated (with LLM-judged translation quality scores) into the ta…
Open documentationBhashaBench-Ayur is the predecessor of BhashaBench-Multi's ayur domain: a domain-specific multiple-choice benchmark evaluating LLM knowledge of Ayurvedic medicine, covering English and Hindi.
Open documentationBhashaBench-Finance is the predecessor of BhashaBench-Multi's finance domain: a domain-specific multiple-choice benchmark evaluating LLM knowledge of finance, covering English and Hindi.
Open documentationBhashaBench-Krishi is the predecessor of BhashaBench-Multi's krishi domain: a domain-specific multiple-choice benchmark evaluating LLM knowledge of agriculture (Krishi), covering English and Hindi.
Open documentationBhashaBench-Legal is the predecessor of BhashaBench-Multi's legal domain: a domain-specific multiple-choice benchmark evaluating LLM knowledge of Indian law, covering English and Hindi.
Open documentationBigCodeBench is an easy-to-use benchmark for solving practical and challenging tasks via code. It evaluates the true programming capabilities of large language models (LLMs) in a more realistic setting with diverse function calls from 139 popular libraries…
Open documentationBigCodeBench-Hard is a curated subset of BigCodeBench containing 148 tasks that are more aligned with real-world programming tasks. These tasks require more complex reasoning and multi-step problem solving.
Open documentationBiomixQA is a curated biomedical question-answering dataset designed to evaluate AI models on biomedical knowledge and reasoning. It has been utilized to validate the Knowledge Graph based Retrieval-Augmented Generation (KG-RAG) framework across different L…
Open documentationBLINK is a benchmark designed to evaluate the core visual perception abilities of Multimodal Large Language Models (MLLMs). It transforms 14 classic computer vision tasks into 3,807 multiple-choice questions with single or multiple images and visual prompts.
Open documentationBroadTwitterCorpus is a dataset of tweets collected over stratified times, places, and social uses. The goal is to represent a broad range of activities, giving a dataset more representative of the language used in this hardest of social media formats to pr…
Open documentationBrowseComp is an OpenAI benchmark for evaluating browsing and search agents. It contains 1,266 hard-to-find, fact-seeking questions with short, verifiable answers. EvalScope loads the mirrored dataset from ModelScope (evalscope/browsecomp).
Open documentationCCBench (Chinese Culture Bench) is an extension of MMBench specifically designed to evaluate multimodal models' understanding of Chinese traditional culture. It covers various aspects of Chinese cultural heritage through visual question answering.
Open documentationCC-OCR V2 is a challenging OCR benchmark tailored to real-world enterprise document processing. It deliberately over-samples the hard and corner cases that prior OCR benchmarks under-represent, such as photographed and scanned tables, handwritten formulas,…
Open documentationC-Eval is a comprehensive Chinese evaluation benchmark designed to assess the knowledge and reasoning abilities of language models in Chinese. It covers 52 subjects ranging from STEM to humanities and social sciences, with questions from middle school to pr…
Open documentationChartQA is a benchmark designed to evaluate question-answering capabilities over charts and data visualizations. It tests both visual reasoning and logical understanding of various chart types including bar charts, line graphs, and pie charts.
Open documentationCharXiv is a comprehensive chart understanding benchmark from NeurIPS 2024 that evaluates multimodal large language models on realistic scientific charts from arXiv papers. It tests both low-level chart element perception (descriptive) and high-level reason…
Open documentationChinese SimpleQA is a Chinese question-answering dataset designed to evaluate the performance of language models on simple factual questions. It tests the model's ability to understand and generate correct answers in Chinese across various knowledge domains.
Open documentationCL-bench represents a step towards building LMs with this fundamental capability (Context Learning), making them more intelligent and advancing their deployment in real-world scenarios. This benchmark is specifically designed to evaluate a model's ability t…
Open documentationClaw-Eval evaluates assistant agents on realistic personal-assistant workflows that require tool use, file and fixture access, multimodal inputs, and simulated user interactions. EvalScope runs the pinned official Claw-Eval Python runner, Docker sandbox, an…
Open documentationCMATH is a Chinese elementary school mathematics benchmark containing 1,698 problems across grades 1-6. It evaluates the mathematical reasoning capabilities of language models on Chinese-language math word problems at increasing difficulty levels.
Open documentationC-MMLU (Chinese Massive Multitask Language Understanding) is a comprehensive Chinese evaluation benchmark covering 67 subjects across STEM, humanities, social sciences, and China-specific topics. It evaluates models' knowledge and reasoning in Chinese conte…
Open documentationCMMU (Chinese Massive Multi-discipline Multimodal Understanding) includes manually collected multimodal questions from college exams, quizzes, and textbooks, covering six core disciplines in Chinese. It is the Chinese counterpart to MMMU.
Open documentationCMMU is a novel Chinese multi-modal benchmark designed to evaluate domain-specific knowledge across seven foundational subjects: math, biology, physics, chemistry, geography, politics, and history. It tests multimodal understanding in Chinese educational co…
Open documentationCoinFlip is a symbolic reasoning benchmark that tests LLMs' ability to track binary state changes through sequences of actions. Each problem involves determining a coin's final state (heads/tails) after various flipping operations.
Open documentationCommon Voice 15 is a massively multilingual speech corpus collected by Mozilla, covering 114 languages with thousands of hours of validated speech data from volunteers worldwide.
Open documentationCommonsenseQA is a benchmark for evaluating AI models' ability to answer questions that require commonsense reasoning about the world. Questions are designed to require background knowledge not explicitly stated in the question.
Open documentationCompetition-MATH is a comprehensive benchmark of 12,500 challenging competition mathematics problems collected from AMC, AIME, and other prestigious math competitions. It is designed to evaluate the advanced mathematical reasoning capabilities of language m…
Open documentationCoNLL-2003 is a classic Named Entity Recognition (NER) benchmark introduced at the Conference on Computational Natural Language Learning 2003. It contains news articles annotated with four entity types.
Open documentationThe CoNLL++ dataset is a corrected and cleaner version of the test set from the widely-used CoNLL2003 NER benchmark. It provides improved annotation quality for evaluating named entity recognition systems on news text.
Open documentationCopious corpus is a gold standard corpus for biodiversity entity recognition, consisting of 668 documents downloaded from the Biodiversity Heritage Library with over 26K sentences and more than 28K entities covering taxonomic and ecological information.
Open documentationCountQA probes object counting, a basic perceptual skill that multimodal models are largely unevaluated on. Its images were hand-captured in everyday environments and deliberately feature high object density, clutter and occlusion, so counting cannot be sol…
Open documentationCrossNER is a fully-labeled collection of named entity recognition (NER) data spanning over five diverse domains: AI, Literature, Music, Politics, and Science. It enables cross-domain NER evaluation and domain adaptation research.
Open documentationData-Collection is a flexible framework for mixing multiple evaluation datasets into a unified evaluation suite. It enables comprehensive model assessment using carefully selected samples from various benchmarks.
Open documentationDeepSWE is a coding-agent benchmark for evaluating repository-level software engineering tasks. EvalScope integrates it through Pier and runs each benchmark sample as one Pier Python API job.
Open documentationDeepSearchQA is a Google DeepMind benchmark for evaluating deep research agents on difficult multi-step information-seeking tasks across the open web. It contains 900 prompts spanning 17 domains and is designed to measure exhaustive answer-set generation ra…
Open documentationDocMath-Eval is a comprehensive benchmark focused on numerical reasoning within specialized domains. It requires models to comprehend long and specialized documents and perform numerical reasoning to answer questions.
Open documentationDocVQA (Document Visual Question Answering) is a benchmark designed to evaluate AI systems' ability to answer questions based on document images such as scanned pages, forms, invoices, and reports. It requires understanding complex document layouts, structu…
Open documentationDrivelology Binary Classification evaluates models' ability to identify "drivelology" - a unique linguistic phenomenon characterized as "nonsense with depth." These are utterances that are syntactically coherent yet pragmatically paradoxical, emotionally lo…
Open documentationDrivelology Multi-label Classification evaluates models' ability to categorize "drivelology" text into rhetorical technique categories: inversion, wordplay, switchbait, paradox, and misdirection. Each text may belong to multiple categories.
Open documentationDrivelology Narrative Selection evaluates models' ability to understand the underlying narrative of "drivelology" text - linguistic utterances that are syntactically coherent yet pragmatically paradoxical, emotionally loaded, or rhetorically subversive.
Open documentationDrivelology Narrative Writing evaluates models' ability to generate detailed descriptions illustrating the implicit narrative of "drivelology" text - linguistic utterances that are syntactically coherent yet pragmatically paradoxical, emotionally loaded, or…
Open documentationDROP (Discrete Reasoning Over Paragraphs) is a challenging reading comprehension benchmark that requires models to perform discrete reasoning operations over text passages. Unlike simple extractive QA, DROP questions require numerical reasoning, counting, a…
Open documentationEmbSpatial-Bench is a benchmark for evaluating embodied spatial understanding of large vision-language models (LVLMs). The benchmark is automatically derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective: close, far,…
Open documentationEQ-Bench is a benchmark for evaluating language models on emotional intelligence tasks. It assesses the ability to predict likely emotional responses of characters in dialogues by rating the intensity of possible emotional reactions.
Open documentationERQA (Embodied Reasoning QA) is a benchmark for evaluating spatial reasoning and embodied understanding capabilities of multimodal large language models. It tests models' ability to reason about trajectories, actions, spatial relationships, and task plannin…
Open documentationEvalMuse is a text-to-image benchmark that evaluates the quality and semantic alignment of generated images using fine-grained analysis with the FGA-BLIP2Score metric.
Open documentationThe FinNER dataset is a corpus of financial agreements from public U.S. Security and Exchange Commission (SEC) filings, annotated with Person, Organization, Location, and Miscellaneous entities to support information extraction for credit risk assessment.
Open documentationFLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) is a massively multilingual benchmark covering 102 languages for evaluating automatic speech recognition (ASR), spoken language understanding, and speech translation.
Open documentationFRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems. It evaluates factuality, retrieval accuracy, and reasoning abilities in long-context scenarios.
Open documentationGAIA (General AI Assistants) is a benchmark of 450+ questions targeting next-generation LLMs with tool use, web browsing and multi-step reasoning. Each question has an unambiguous short answer and is bucketed into one of three difficulty levels.
Open documentationGDPval evaluates whether models can complete realistic economically valuable work tasks and produce requested deliverable files. This adapter targets OpenAI's public 220-task gold subset mirrored on ModelScope as openai-mirror/gdpval.
Open documentationGEdit-Bench (Grounded Edit Benchmark) is an image editing benchmark grounded in real-world usage scenarios. It provides comprehensive evaluation of image editing models across diverse editing tasks with LLM-based judging.
Open documentationGenAI-Bench is a comprehensive text-to-image benchmark featuring 1600 prompts designed to evaluate image generation models across diverse categories and complexity levels.
Open documentationGeneralArena is a custom benchmark designed to evaluate the performance of large language models in a competitive setting, where models are pitted against each other in custom tasks to determine their relative strengths and weaknesses.
Open documentationGeneral-FunctionCalling is a customizable benchmark for evaluating function calling (tool use) capabilities of language models. It tests both the decision to call tools and the accuracy of generated function calls.
Open documentationGeneral-MCQ is a customizable multiple-choice question answering benchmark for evaluating language models. It supports flexible data formats and variable number of answer choices.
Open documentationGeneral-QA is a customizable question answering benchmark for evaluating language models on open-ended text generation tasks. It supports flexible data formats and configurable evaluation metrics.
Open documentationGeneral Text-to-Image is a customizable benchmark adapter for evaluating text-to-image generation models with user-provided prompts and images.
Open documentationGeneral-VMCQ is a customizable visual multiple-choice question answering benchmark for multimodal models. It uses MMMU-style format with image/video/audio placeholders in text, supporting flexible media inputs.
Open documentationGeneral-VQA is a customizable visual question answering benchmark for evaluating multimodal models. It supports OpenAI-compatible message format with flexible image/video/audio input (local paths, URLs, or base64).
Open documentationGeniaNER is a large-scale biomedical NER dataset consisting of 2,000 MEDLINE abstracts with over 400,000 words and almost 100,000 annotations for biological terms. It is one of the most comprehensive resources for biomedical entity recognition.
Open documentationGPQA (Graduate-Level Google-Proof Q&A) Diamond is a challenging benchmark of 198 multiple-choice questions written by domain experts in biology, physics, and chemistry. The questions are designed to be extremely difficult, requiring PhD-level expertise to a…
Open documentationGSM8K (Grade School Math 8K) is a high-quality dataset of 8.5K linguistically diverse grade school math word problems created by human problem writers. The dataset is specifically designed to evaluate and improve the multi-step mathematical reasoning capabi…
Open documentationGSM8K-Indic translates the GSM8K grade-school math word problems into 10 Indic languages, each available in native script and a romanized (Latin transliteration) variant, plus the original English.
Open documentationGSM8K-V is a purely visual multi-image mathematical reasoning benchmark that systematically transforms each GSM8K math word problem into its visual counterpart. It enables clean within-item comparison across modalities for multimodal math evaluation.
Open documentationHallusionBench is an advanced diagnostic benchmark designed to evaluate image-context reasoning and detect hallucination tendencies in Large Vision-Language Models (LVLMs). It specifically tests models' susceptibility to language hallucination and visual il…
Open documentationHaluEval is a large collection of generated and human-annotated hallucinated samples for evaluating the performance of LLMs in recognizing hallucination. It provides a comprehensive benchmark for assessing model reliability and factual accuracy.
Open documentationHarveyNER is a dataset with fine-grained locations annotated in tweets, collected during Hurricane Harvey. It presents unique challenges with complex and long location mentions in informal crisis-related descriptions.
Open documentationHealthBench is a comprehensive benchmark designed to measure AI capabilities for health-related tasks. Built in partnership with 262 physicians from 60 countries, it includes 5,000 realistic health conversations with custom physician-created rubrics.
Open documentationHellaSwag is a benchmark for evaluating commonsense natural language inference, specifically testing a model's ability to complete sentences describing everyday situations. The dataset uses adversarial filtering to create challenging distractors that are gr…
Open documentationHellaSwag-Hindi is a Hindi translation of the HellaSwag commonsense sentence-completion benchmark's full validation set. The context stem stays in English; the 4 candidate continuations are translated into Hindi, so the model must connect an English scenari…
Open documentationHiPhO is the first benchmark dedicated to high school physics Olympiads with human-aligned evaluation. It compiles 13 recent Olympiad exams (2024-2025) spanning international and regional competitions, with mixed modalities that range from text-only problem…
Open documentationHumanity's Last Exam (HLE) is a comprehensive language model benchmark consisting of 2,500 questions across a broad range of subjects. Created jointly by the Center for AI Safety and Scale AI, it represents one of the most challenging academic benchmarks av…
Open documentationHMMT February 2025 (MathArena) is a challenging evaluation benchmark derived from the Harvard-MIT Mathematics Tournament (HMMT) February 2025 competition, one of the most prestigious and difficult high school math contests globally.
Open documentationHMMT February 2026 is a challenging evaluation benchmark derived from the Harvard-MIT Mathematics Tournament (HMMT) February 2026 competition, one of the most prestigious and difficult high school math contests globally.
Open documentationHMMT November 2025 (MathArena) is a challenging evaluation benchmark derived from the Harvard-MIT Mathematics Tournament (HMMT) November 2025 competition, one of the most prestigious and difficult high school math contests globally. It is a different contes…
Open documentationHPD-v2 (Human Preference Dataset v2) is a text-to-image benchmark that evaluates generated images based on human preferences. It uses the HPSv2.1 score metric trained on large-scale human preference data.
Open documentationHumanEval is a benchmark for evaluating the code generation capabilities of language models. It consists of 164 hand-written Python programming problems with function signatures, docstrings, and comprehensive test cases.
Open documentationHumanEval Plus is a rigorous extension of OpenAI's HumanEval benchmark, designed to address high false-positive rates in code generation evaluation. It augments the original test cases with tens of thousands of automatically generated inputs to expose edge-…
Open documentationIFBench is a benchmark designed to evaluate how reliably AI models follow novel, challenging, and diverse verifiable instructions, with a strong focus on out-of-domain generalization. Developed by AllenAI, it addresses overfitting and data contamination iss…
Open documentationIFEval (Instruction-Following Eval) is a benchmark for evaluating how well language models follow explicit, verifiable instructions. It contains prompts with specific formatting, content, or structural requirements that can be objectively verified.
Open documentationIMO-AnswerBench is a benchmark of 400 challenging problems sourced from the International Mathematical Olympiad (IMO) Shortlists. It covers four major mathematical domains and is designed to evaluate advanced mathematical reasoning capabilities of language…
Open documentationBoolQ-Indic is a translation of the BoolQ yes/no reading-comprehension benchmark into 10 Indic languages plus English, for evaluating multilingual passage understanding.
Open documentationIndicParam is a graduate-level benchmark evaluating LLM understanding of low- and extremely low-resource Indic languages. All 13,207 multiple-choice questions are sourced from official UGC-NET language question papers and answer keys, presented in each lang…
Open documentationInfoVQA (Infographic Visual Question Answering) is a benchmark designed to evaluate AI models' ability to answer questions based on information-dense images such as charts, graphs, diagrams, maps, and infographics. It focuses on understanding complex visual…
Open documentationIQuiz is a Chinese benchmark for evaluating AI models on intelligence quotient (IQ) and emotional quotient (EQ) questions. It tests logical reasoning, pattern recognition, and social-emotional understanding through multiple-choice questions.
Open documentationThe JNLPBA dataset is a widely-used resource for bio-entity recognition, consisting of 2,404 MEDLINE abstracts from the GENIA corpus annotated for five key molecular biology entity types. It is a standard benchmark for biomedical NER.
Open documentationThe JNLPBA-Rare dataset is a specialized subset of the JNLPBA test set created to evaluate zero-shot performance on its least frequent entity types: RNA and cell line. It tests model ability to recognize rare biomedical entities.
Open documentationJobBench evaluates agentic systems on realistic professional work tasks that require reading reference files, producing deliverables, and reconciling multi-source information. This adapter uses the ModelScope dataset evalscope/job-bench.
Open documentationK2-Vendor-Verifier checks whether a third-party deployment of Kimi-K2 faithfully reproduces the official Moonshot AI API's tool-calling behavior. It replays the official evaluation prompt set against a vendor endpoint and compares finishreason and tool-call…
Open documentationKimi-Vendor-Verifier is a pre-flight compliance check for Kimi K2 / K2-Thinking deployments. It sends synthetic probe requests to verify that the vendor API correctly rejects non-default values of immutable decoding parameters (temperature, topp, presencepe…
Open documentationKINA (Knowledge Index of Noah's Ark) is a high-density multidisciplinary knowledge benchmark for evaluating whether large language models can solve expert-level questions across 261 fine-grained disciplines. It is the first benchmark to incorporate discipli…
Open documentationLibriSpeech is a large-scale corpus of approximately 1,000 hours of read English speech derived from audiobooks. It is one of the most widely used benchmarks for evaluating automatic speech recognition (ASR) systems.
Open documentationLiveCodeBench is a contamination-free benchmark for evaluating code generation models on real-world competitive programming problems. It continuously collects new problems from coding platforms to ensure models haven't seen the test data during training.
Open documentationLoCoMo evaluates very long-term conversational memory in two-person multi-session dialogues. This adapter supports the official question-answering task from locomo10.json.
Open documentationLogiQA is a benchmark for evaluating logical reasoning abilities, sourced from expert-written questions originally designed for testing human logical reasoning skills in standardized examinations.
Open documentationLogicVista evaluates the fundamental logical reasoning abilities of multimodal large language models in visual contexts. Every item is a multiple-choice question whose answer options are drawn inside the image (diagrams, puzzles, sequences, charts), so a mo…
Open documentationLongBench v2 is a challenging benchmark for evaluating long-context understanding of large language models. It covers a wide variety of real-world tasks that require reading and comprehending long documents (ranging from a few thousand to over 2 million tok…
Open documentationLongMemEval evaluates long-term interactive memory in chat assistants. Each question is answered from a timestamped multi-session user-assistant history.
Open documentationMaritimeBench is a benchmark for evaluating AI models on maritime-related multiple-choice questions in Chinese. It consists of specialized questions related to maritime knowledge, navigation, marine engineering, and seafaring operations.
Open documentationMaritime-OCR-Bench is a comprehensive evaluation benchmark for assessing multimodal large model capabilities on OCR-related tasks. The current released set contains 1,888 manually curated samples across five task types.
Open documentationMATH-500 is a curated subset of 500 problems from the MATH benchmark, designed to evaluate the mathematical reasoning capabilities of language models. It covers five difficulty levels across various mathematical topics including algebra, geometry, number th…
Open documentationMathQA is a large-scale dataset for mathematical word problem solving, gathered by annotating the AQuA-RAT dataset with fully-specified operational programs using a new representation language. It contains diverse math problems requiring multi-step reasoning.
Open documentationMathVerse is an all-around visual math benchmark designed for equitable and in-depth evaluation of Multimodal Large Language Models (MLLMs). It contains 2,612 high-quality, multi-subject math problems with diagrams, transformed into 15K test samples across…
Open documentationMATH-Vision (MATH-V) is a meticulously curated dataset of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions. It evaluates mathematical reasoning abilities in multimodal settings.
Open documentationMathVista is a comprehensive benchmark for mathematical reasoning in visual contexts. It combines newly created datasets with existing benchmarks to evaluate models on diverse visual mathematical reasoning tasks across multiple domains.
Open documentationMBPP (Mostly Basic Python Problems) is a benchmark consisting of approximately 1,000 crowd-sourced Python programming problems designed for entry-level programmers. It evaluates a model's ability to understand problem descriptions and generate correct Pytho…
Open documentationMBPP Plus is a fortified version of the MBPP benchmark, created to improve evaluation reliability for basic Python programming synthesis. It addresses quality issues in the original dataset and significantly increases test coverage for each problem.
Open documentationMCP-Atlas is a Scale AI benchmark for evaluating tool-use competency with real Model Context Protocol (MCP) servers. It contains public tasks with prompts, allowed tool lists, ground-truth tool trajectories, and expert claims used for LLM-as-judge coverage…
Open documentationMeasureBench is a comprehensive benchmark for evaluating the ability of vision-language models (VLMs) to read values from measuring instruments. It covers both real-world photographs and synthetically generated images of 26 instrument types across 4 design…
Open documentationMedMCQA is a large-scale multiple-choice question answering dataset designed to address real-world medical entrance exam questions. It contains over 194K questions covering diverse medical topics from Indian medical entrance examinations (AIIMS, NEET-PG).
Open documentationMedXpertQA is an expert-level medical multiple-choice benchmark designed to evaluate advanced medical knowledge and reasoning. It contains separate text-only and multimodal tracks built from challenging medical examination questions and reviewed by licensed…
Open documentationMGSM (Multilingual Grade School Math) is a benchmark designed to evaluate multilingual mathematical reasoning capabilities of language models. It extends GSM8K to 11 typologically diverse languages, testing whether models can perform chain-of-thought reason…
Open documentationMIA-Bench is a multimodal instruction-following benchmark designed to evaluate vision-language models on their ability to follow complex, compositional instructions grounded in images. Each sample contains an image paired with a multi-component instruction,…
Open documentationMicroVQA is an expert-curated benchmark for multimodal reasoning in microscopy-based scientific research. It evaluates AI models' ability to understand and reason about microscopy images across various scientific domains.
Open documentationMILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset for assessing LLM performance across 11 Indic languages. It spans 8 domains and 41 subjects, combining translated general-knowledge questions with culturally spec…
Open documentationMinerva-Math is a benchmark designed to evaluate advanced mathematical and quantitative reasoning capabilities of language models. It consists of 272 challenging problems sourced primarily from MIT OpenCourseWare courses, covering university and graduate-le…
Open documentationMiniMax-Vendor-Verifier is a multi-validator deployment-correctness check for MiniMax M2 / M2.5 / M2.7 vendors. Each prompt row carries an optional checktype tag that routes it through specific validators, plus an always-on erroronlyreasoning detector for t…
Open documentationMiniWoB evaluates whether a multimodal agent can complete short browser tasks such as clicking buttons, filling forms, scrolling and dragging items.
Open documentationThe MIT-Movie-Trivia dataset, originally created for slot filling in movie domain dialogues, has been modified for NER by merging and filtering slot types. It tests recognition of movie-related entities in conversational queries.
Open documentationThe MIT-Restaurant dataset is a collection of restaurant review text specifically curated for training and testing NLP models for Named Entity Recognition. It contains sentences from real reviews with annotations in BIO format.
Open documentationMMBench is a systematically designed benchmark for evaluating vision-language models across 20 fine-grained ability dimensions. It uses a novel CircularEval strategy and provides both English and Chinese versions for cross-lingual evaluation.
Open documentationMMStar is an elite vision-indispensable multimodal benchmark designed to ensure genuine visual dependency in evaluation. Each sample is carefully curated to require actual visual understanding, minimizing data leakage and testing advanced multimodal capabil…
Open documentationMMAU (Massive Multitask Audio Understanding) is a comprehensive benchmark for evaluating audio understanding capabilities of multimodal large language models across diverse audio tasks.
Open documentationMMLU (Massive Multitask Language Understanding) is a comprehensive evaluation benchmark designed to measure knowledge acquired during pretraining. It covers 57 subjects across STEM, humanities, social sciences, and other domains, ranging from elementary to…
Open documentationMMLU-Pro is an enhanced version of MMLU with increased difficulty and reasoning requirements. It features 10 answer choices instead of 4 and includes more challenging questions requiring deeper reasoning across 14 diverse domains.
Open documentationMMLU-Redux is an improved version of the MMLU benchmark with corrected answers. It addresses known errors in the original MMLU dataset by fixing incorrect ground truth labels, missing correct options, and ambiguous questions.
Open documentationMMMLU (Multilingual Massive Multitask Language Understanding) is a multilingual extension of the MMLU benchmark. It evaluates the multilingual knowledge and reasoning capabilities of language models across 14 languages, covering 57 subjects from the origina…
Open documentationMMMU (Massive Multi-discipline Multimodal Understanding) is a comprehensive benchmark designed to evaluate multimodal models on expert-level tasks requiring college-level subject knowledge and deliberate reasoning. It covers 30 subjects across 6 core discip…
Open documentationMMMU-PRO is an enhanced multimodal benchmark designed to rigorously assess the genuine understanding capabilities of advanced AI models across multiple modalities. It builds upon the original MMMU benchmark with key improvements that make evaluation more ch…
Open documentationMRI-MCQA is a specialized benchmark composed of multiple-choice questions related to Magnetic Resonance Imaging (MRI). It evaluates AI models' understanding of MRI physics, protocols, image acquisition, and clinical applications.
Open documentationMSR-VTT is a large-scale open-domain video captioning benchmark for evaluating video-to-text generation. The native adapter groups records by videoid, so multiple annotation rows for one video become one sample with multiple reference captions.
Open documentationMSVD is a classic video captioning benchmark with short web videos annotated by many human captions. The native adapter treats each video as one evaluation sample and uses all available captions as references.
Open documentationMulti-IF is a benchmark designed to evaluate LLM capabilities in multi-turn instruction following within a multilingual environment. It tests the ability to follow complex instructions across multiple conversation turns in different languages.
Open documentationMultiNERD is a large-scale, multilingual, and multi-genre dataset for fine-grained Named Entity Recognition, automatically generated from Wikipedia and Wikinews. It covers 10 languages and 15 distinct entity categories.
Open documentationMultiPL-E HumanEval is a multilingual code generation benchmark derived from OpenAI's HumanEval. It extends the original HumanEval to 18 programming languages, enabling cross-lingual evaluation of code generation capabilities.
Open documentationMultiPL-E MBPP is a multilingual code generation benchmark derived from MBPP (Mostly Basic Python Programming). It extends the original MBPP to 18 programming languages, enabling cross-lingual evaluation of code generation capabilities.
Open documentationMusicTrivia is a curated multiple-choice benchmark for evaluating AI models on music knowledge. It covers both classical and modern music topics including composers, musical periods, instruments, and popular artists.
Open documentationMuSR (Multistep Soft Reasoning) is a benchmark for evaluating complex reasoning abilities through narrative-based problems. It includes murder mysteries, object placements, and team allocation scenarios requiring multi-step inference.
Open documentationMVBench is a public multimodal video understanding benchmark covering temporal perception, attribute/state reasoning, symbolic ordering, and high-level cognition. This native adapter uses the ModelScope PKU-Alignment/MVBench mirror by default, which provide…
Open documentationThe NCBI disease corpus is a manually annotated resource of PubMed abstracts designed for disease name recognition and normalization. It provides a gold standard for evaluating disease named entity recognition systems.
Open documentationNeedle in a Haystack is a benchmark focused on evaluating information retrieval capabilities in long-context scenarios. It tests a model's ability to find specific information (needles) within large documents (haystacks).
Open documentationOCRBench is a comprehensive evaluation benchmark designed to assess the OCR (Optical Character Recognition) capabilities of Large Multimodal Models. It covers five key OCR-related tasks with 1,000 manually verified question-answer pairs.
Open documentationOCRBench v2 is a large-scale bilingual text-centric benchmark with the most comprehensive set of OCR tasks (4x more than OCRBench v1), covering 31 diverse scenarios including street scenes, receipts, formulas, diagrams, and more.
Open documentationOfficeQA is a grounded reasoning benchmark by Databricks, built for evaluating model/agent performance on end-to-end grounded reasoning tasks over U.S. Treasury Bulletin documents (1939-2025).
Open documentationolmOCR-Bench evaluates end-to-end document transcription: a model reads one rendered PDF page and returns the full Markdown transcription of that page, which is then checked against human-written unit tests instead of a single reference answer.
Open documentationOlympiadBench is an Olympiad-level bilingual multimodal scientific benchmark featuring 8,476 problems from mathematics and physics competitions, including the Chinese college entrance exam (CEE). It provides rigorous evaluation of advanced scientific reason…
Open documentationOmniBench is a pioneering universal multimodal benchmark designed to rigorously evaluate MLLMs' capability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously.
Open documentationThis adapter preserves EvalScope's original 981-page OmniDocBench TSV integration for compatibility with existing evaluations.
Open documentationOmniDocBench v1.6 evaluates end-to-end document parsing for text, formulas, tables, layout, and reading order. This adapter is intentionally restricted to the official v1.6 data and scoring contract.
Open documentation$OneMillion-Bench ($1M-Bench) evaluates how well language models and agents complete economically valuable, expert-level professional work. The public release contains 400 bilingual tasks written and reviewed by domain experts across finance, healthcare, in…
Open documentationOntoNotes Release 5.0 is a large, multilingual corpus containing text in English, Chinese, and Arabic across various genres. It is richly annotated with multiple layers of linguistic information including syntax, predicate-argument structure, word sense, na…
Open documentationMRCR (Memory-Recall with Contextual Retrieval) is OpenAI's benchmark for evaluating retrieval and recall capabilities in long-context scenarios. It tests whether models can correctly extract and use specific information (needles) embedded in long prompts.
Open documentationPerceptionBench is a benchmark from Moonshot AI that evaluates the atomic visual perception capabilities of multimodal large language models. It is built bottom-up: the earliest failure points of frontier MLLMs on 42 existing benchmarks were diagnosed to de…
Open documentationPerspectiveGap evaluates whether a model can compose orchestration prompts for multi-agent systems while routing only the context each sub-agent needs.
Open documentationPerspectiveGap evaluates whether a model can compose orchestration prompts for multi-agent systems while routing only the context each sub-agent needs.
Open documentationPhyX is the first large-scale benchmark for physical reasoning in realistic, visually grounded scenarios. This is its multiple-choice variant: each university-level physics problem is presented with a figure and four answer options, and the model has to nam…
Open documentationPhyX is the first large-scale benchmark for physical reasoning in realistic, visually grounded scenarios. This is its open-ended variant: no options are shown, so the model has to derive the answer of a university-level physics problem from the figure and s…
Open documentationPIQA (Physical Interaction QA) is a benchmark for evaluating AI models' understanding of physical commonsense - how objects interact in the physical world and what happens when we manipulate them.
Open documentationPLawBench is a rubric-based benchmark that evaluates large language models on real-world Chinese legal practice. It mirrors the workflow of a practising lawyer across three hierarchical levels: eliciting facts during a public legal consultation, analysing a…
Open documentationPMC-VQA is a large-scale medical visual question answering benchmark built from figures of biomedical papers in the PubMed Central Open Access subset. This integration evaluates the manually verified testclean split, the 2,000-question subset the authors re…
Open documentationPolyMath is a multilingual mathematical reasoning benchmark covering 18 languages and 4 difficulty levels with 9,000 high-quality problem samples. It ensures difficulty comprehensiveness, language diversity, and high-quality translation for discriminative m…
Open documentationPOPE (Polling-based Object Probing Evaluation) is a benchmark specifically designed to evaluate object hallucination in Large Vision-Language Models (LVLMs). It tests models' ability to accurately identify objects present in images through yes/no questions.
Open documentationPRBench (Professional Reasoning Benchmark) evaluates open-ended reasoning on realistic, high-stakes Finance and Legal problems. Its expert-authored conversations and fine-grained rubrics measure whether a response is accurate, useful, auditable, and appropr…
Open documentationProcessBench is a benchmark for evaluating AI models on mathematical reasoning process verification. It tests the ability to identify errors in step-by-step mathematical solutions across various difficulty levels from GSM8K to OmniMath.
Open documentationPubMedQA is a biomedical question answering dataset designed to evaluate models' ability to reason over biomedical research texts. It contains questions derived from PubMed abstracts with yes/no/maybe answers.
Open documentationQASC (Question Answering via Sentence Composition) is a question-answering dataset with a focus on multi-hop sentence composition. It consists of 9,980 8-way multiple-choice questions about grade school science, requiring models to combine multiple facts to…
Open documentationRACE (ReAding Comprehension from Examinations) is a large-scale reading comprehension benchmark collected from Chinese middle school and high school English examinations. It tests comprehensive reading comprehension abilities.
Open documentationRealWorldQA is a benchmark contributed by XAI designed to evaluate multimodal AI models' understanding of real-world spatial and physical environments. It uses authentic images from everyday scenarios to test practical visual comprehension.
Open documentationRef-Adv-s is the public 1,142-case subset of Ref-Adv, a referring expression comprehension benchmark designed to test whether multimodal large language models can distinguish a target from hard same-category visual distractors instead of relying on groundin…
Open documentationRefCOCO is a dataset for training and evaluating models on Referring Expression Comprehension (REC). It contains images, object bounding boxes, and free-form natural-language expressions that uniquely describe target objects within MSCOCO images.
Open documentationResearchRubrics evaluates Deep Research agents on realistic, open-ended research tasks. Each task pairs a user prompt with expert-written, fine-grained rubrics covering explicit and implicit requirements, information synthesis, references, communication qua…
Open documentationSanskriti is a multiple-choice trivia benchmark testing knowledge of Indian states' culture, history, and geography, sourced from state-specific attributes (art, cuisine, festivals, etc.) with Wikipedia-backed answers. From the SANSKRITI paper (arXiv:2506.1…
Open documentationSciCode is a challenging benchmark designed to evaluate language model capabilities in generating code for solving realistic scientific research problems. It covers 16 subdomains from 5 major domains: Physics, Math, Material Science, Biology, and Chemistry.
Open documentationScienceQA is a multimodal benchmark consisting of multiple-choice science questions derived from elementary and high school curricula. It covers diverse subjects including natural science, social science, and language science, with questions accompanied by…
Open documentationSciQ is a crowdsourced science exam question dataset covering Physics, Chemistry, Biology, and other scientific domains. Most questions include supporting evidence paragraphs.
Open documentationScreenSpot-Pro is a GUI grounding benchmark built from authentic high-resolution screenshots of professional desktop software. Given a natural-language instruction, a model must locate the target UI element on the screen, which stresses fine-grained localiz…
Open documentationSEED-Bench-2-Plus is a large-scale benchmark designed to evaluate Multimodal Large Language Models (MLLMs) on text-rich visual understanding tasks. It contains 2.3K multiple-choice questions with precise human annotations across real-world scenarios.
Open documentationSeed-TTS-Eval is an objective benchmark for zero-shot text-to-speech and voice conversion evaluation. It uses out-of-domain English and Mandarin samples from Common Voice and DiDiSpeech-2, and the official evaluation focuses on intelligibility and speaker c…
Open documentationSimpleQA is a benchmark by OpenAI designed to evaluate language models' ability to answer short, fact-seeking questions accurately. It focuses on measuring factual accuracy with clear grading criteria for correct, incorrect, and not-attempted answers.
Open documentationSimpleVQA is the first comprehensive multimodal benchmark to evaluate the factuality ability of MLLMs to answer natural language short questions. It features high-quality, challenging queries with static and timeless reference answers.
Open documentationSIQA (Social Interaction QA) is a benchmark for evaluating social commonsense intelligence - understanding people's actions and their social implications. Unlike benchmarks focusing on physical knowledge, SIQA tests reasoning about human behavior.
Open documentationSkillsBench evaluates whether coding agents can discover and apply task-bundled Agent Skills. Each task contains an instruction, an optional skill directory, a Docker environment, an oracle solution, and a verifier. EvalScope builds the task Docker image, r…
Open documentationSLAKE is a bilingual (English / Chinese) radiology visual question answering benchmark built by physicians on CT, MRI and X-Ray images. Questions cover both purely visual properties of the scan and medical knowledge that has to be recalled on top of what th…
Open documentationSuperGPQA is a large-scale multiple-choice question answering dataset designed to evaluate model generalization across diverse fields. It contains 26,000+ questions from 50+ fields, with each question featuring 10 answer options.
Open documentationSURDS benchmarks fine-grained spatial understanding and reasoning by vision-language models in realistic driving scenes. It is derived from the six-camera nuScenes dataset and evaluates object-centric and relational spatial skills without supplying depth ma…
Open documentationSWE-bench Lite is a focused subset of SWE-bench containing 300 Issue-Pull Request pairs from 11 popular Python repositories. It provides a more accessible entry point for evaluating automated software engineering capabilities.
Open documentationSWE-bench Lite Agentic is the agentic-mode evaluation of SWE-bench Lite, a focused subset of SWE-bench containing 300 Issue-Pull Request pairs from 11 popular Python repositories. The model autonomously drives a multi-turn agent loop inside a per-instance D…
Open documentationSWE-bench Multilingual Agentic is the agentic-mode evaluation of SWE-bench Multilingual, a 300-task SWE-bench-style benchmark spanning 42 repositories and 9 programming languages. The model autonomously explores, edits, and submits a patch through a multi-t…
Open documentationSWE-benchPro is a challenging benchmark from Scale AI evaluating LLMs/Agents on long-horizon software engineering tasks across multiple programming languages. Given a codebase and an issue, the model must autonomously explore the repository, edit source fil…
Open documentationSWE-bench Verified is a human-validated subset of 500 samples from SWE-bench, designed to test systems' ability to automatically resolve real-world GitHub issues. Each sample represents a genuine bug fix or feature implementation from popular Python reposit…
Open documentationSWE-bench Verified Agentic is the agentic-mode evaluation of SWE-bench Verified, a human-validated subset of 500 samples from SWE-bench. Unlike the oracle single-turn variant, the model must autonomously explore the repository, run shell commands, edit sour…
Open documentationSWE-bench Verified Mini is a compact subset of SWE-bench Verified, containing 50 carefully selected samples that maintain the same distribution of performance, test pass rates, and difficulty as the full dataset while requiring only 5GB of storage instead o…
Open documentationSWE-bench Verified Mini Agentic is the agentic-mode evaluation of SWE-bench Verified Mini, a compact 50-sample subset that maintains the same distribution of performance, test pass rates, and difficulty as the full Verified set while requiring only 5GB of s…
Open documentationτ²-bench (Tau Squared Bench) is an extension and enhancement of the original τ-bench. It evaluates conversational AI agents in domain-specific scenarios with expanded capabilities including telecom domain support.
Open documentationτ³-bench (Tau Cubed Bench) is the v1.0.0 release of the tau-bench family. It extends τ²-bench with a knowledge-retrieval domain, voice/audio-native evaluation, and 75+ task fixes across the existing domains.
Open documentationτ-bench (Tau Bench) is a benchmark for evaluating conversational AI agents that interact with users through domain-specific API tools and policy guidelines. It simulates dynamic, multi-turn conversations where a language model acts as both the user and the…
Open documentationTerminal-Bench v2 is a command-line benchmark suite that evaluates AI agents on 89 real-world, multi-step terminal tasks. Tasks range from compiling and debugging to system administration, running within isolated containers with rigorous validation.
Open documentationTerminal-Bench v2.1 is an improved iteration of Terminal-Bench 2.0, with 26 task fixes addressing bugs, timeout adjustments, and reward hacking prevention. Recommended over v2.0 for new evaluations.
Open documentationTIFA-160 is a text-to-image benchmark with 160 carefully curated prompts designed to evaluate the faithfulness and quality of generated images using automated VQA-based evaluation.
Open documentationTIR-Bench (Thinking-with-Images Reasoning Benchmark) is a comprehensive multimodal benchmark that evaluates agentic visual reasoning capabilities of vision-language models. It covers diverse task categories requiring spatial, compositional, and multi-step v…
Open documentationToolBench-Static is a benchmark for evaluating AI models' ability to use tools and execute function calls in a static evaluation setting. It tests models on realistic tool-use scenarios with ground-truth comparisons.
Open documentationToolathlon is an agent benchmark for realistic, long-horizon tool use across many MCP-backed software environments. This EvalScope benchmark is a wrapper around the official Toolathlon remote evaluation service, not a local reimplementation of the MCP envir…
Open documentationTORGO is a specialized database of dysarthric speech designed for evaluating ASR systems on speakers with motor speech disorders. It contains aligned acoustic and articulatory data from speakers with cerebral palsy (CP) or amyotrophic lateral sclerosis (ALS).
Open documentationTriviaQA is a large-scale reading comprehension dataset containing over 650K question-answer-evidence triples. Questions are collected from trivia enthusiast websites and paired with Wikipedia articles as evidence documents.
Open documentationTriviaQA-Indic-MCQ reformats TriviaQA trivia questions as 4-way multiple-choice questions, translated into 10 Indic languages plus English, for evaluating multilingual world-knowledge recall.
Open documentationTruthfulQA is a benchmark designed to measure whether language models generate truthful answers to questions. It focuses on questions where humans might give false answers due to misconceptions, superstitions, or false beliefs.
Open documentationTVBench is a temporal video understanding benchmark for evaluating whether multimodal models can reason over dynamic visual events rather than isolated frames. It covers a broad set of video reasoning skills, including action recognition, action counting, t…
Open documentationTweebank-NER is an English Twitter corpus created by annotating the syntactically-parsed Tweebank V2 with four types of named entities: Person, Organization, Location, and Miscellaneous. It addresses NER challenges in informal social media text.
Open documentationTweetNER7 is a large-scale NER dataset featuring over 11,000 tweets from 2019-2021, annotated with seven entity types to facilitate the study of short-term temporal shifts in social media language.
Open documentationVideo-MME-v2 is a public comprehensive video understanding benchmark. It contains 800 videos, 3,200 multiple-choice QA instances, and word-level subtitles with timestamps. The native adapter uses the shared DatasetHub abstraction for both annotation loading…
Open documentationVisFactor evaluates foundational visual cognition in multimodal large language models using 20 vision-centric subtests adapted from the Factor-Referenced Cognitive Test (FRCT). It isolates abilities that support higher-level visual reasoning instead of meas…
Open documentationVisuLogic is a benchmark for evaluating visual reasoning capabilities of Multimodal Large Language Models (MLLMs), independent of textual reasoning. It features carefully constructed visual reasoning tasks that are inherently difficult to articulate using l…
Open documentationVLMs Are Biased (VLMBias) evaluates whether vision-language models answer objective visual questions from the image or fall back to memorized prior knowledge. It uses counterfactual images whose visible properties conflict with familiar concepts, such as an…
Open documentationVQAv2 is the balanced Visual Question Answering benchmark built on COCO images. It evaluates whether multimodal models can answer open-ended natural-language questions grounded in image content.
Open documentationVBench is a benchmark designed for evaluating visual search capabilities within multimodal reasoning systems. It focuses on actively locating and identifying specific visual information in high-resolution images, crucial for fine-grained visual understanding.
Open documentationVTCBench (Vision-Text Compression Benchmark) evaluates long-context understanding when text is represented as rendered images, and compares it with a pure-text baseline.
Open documentationWenetSpeech is a large-scale Mandarin Chinese speech corpus with over 10,000 hours of multi-domain transcribed audio data, designed for speech recognition research.
Open documentationWideSearch evaluates search agents on broad web information-seeking tasks. Each task asks the agent to collect many atomic facts and return one structured Markdown table. EvalScope uses the ModelScope bytedance-community/WideSearch dataset.
Open documentationWinogrande is a large-scale benchmark for commonsense reasoning, specifically designed to test pronoun resolution in the Winograd Schema Challenge format. It contains 44K problems that require understanding of physical and social commonsense.
Open documentationWMT2024++ is a comprehensive machine translation benchmark based on the WMT 2024 news translation task. It supports 54 language pairs with English as the source language, enabling evaluation of translation quality across diverse target languages.
Open documentationThe WNUT2017 dataset is a collection of user-generated text from various social media platforms, like Twitter and YouTube, specifically designed for named entity recognition tasks focusing on emerging and unusual entities.
Open documentationWorldVQA is a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models (MLLMs). It measures models' ability to ground and name visual entities across a stratified taxonomy, spanning from common head-class objects…
Open documentationZebraLogicBench is a comprehensive evaluation framework for assessing LLM reasoning performance on logic grid puzzles derived from constraint satisfaction problems (CSPs). It tests systematic logical reasoning abilities.
Open documentationZeroBench is a challenging visual reasoning benchmark for Large Multimodal Models (LMMs). It consists of 100 high-quality, manually curated questions covering numerous domains, reasoning types, and image types designed to be beyond current model capabilities.
Open documentationEach catalog entry routes to generated documentation and carries task type, modality and metrics. Missing metadata causes the site check to fail.
Add an adapter, declare BenchmarkMeta, run its narrow smoke test, then generate docs. Ask an AI agent to follow the project skill for the surrounding workflow.
Open full setup guideOpen contributions · quality first
Implement a standard interface.
Verify a narrow runnable path.
Declare version and metrics.
Generate bilingual details.
Make it discoverable.
