What actually changed?
Separate task behavior from prompt, rubric, data and scoring-scale changes.
EvalScope keeps evaluation inputs, reviews, and outputs connected. This page explores how to tell a real evaluator improvement from a change in data, prompts, or scoring.

THREE QUESTIONS EVERY EVALUATOR CHANGE SHOULD ANSWER
Separate task behavior from prompt, rubric, data and scoring-scale changes.
Audit judge consistency and failure modes instead of treating a score as an oracle.
Measure the consequence when an update changes a ranking or recommendation.
A REPRODUCIBLE VALIDATION PATH
Freeze a baseline, introduce one controlled change, test it on held-out tasks, inspect the evidence, and record the decision with its artifacts. Preserved raw outputs and explicit claims keep any apparent improvement reviewable.
Tasks, environments and raw artifacts.
One controlled change at a time.
Held-out tasks and declared metrics.
Review evidence before accepting a score.
Keep the selection and its rationale.
Labels separate available capabilities, external evidence, and open research directions.
Versioned benchmark metadata, structured judge contracts, output artifacts, local reports and API-first evaluation workflows.
Self-evolving harness evaluation, evaluator drift guardrails and selection-regret protocols. These are research directions, not product claims.
Preserve requests, tool calls and outcomes alongside the task result.
Browse benchmark metadata and evaluation versions across modalities.
Reports, predictions, reviews and configs remain linked.