Inputs and outcomes
Keep the task, final response and completion state together so a result can be checked in context.
AGENT & HARNESS
Capture model requests, tool calls, token usage and final outcomes for agent tasks — with a path for external replay scenarios too.

MEASURE · TRACE · IMPROVE
Every step counts
Task and context
qwen-plus
python_exec
Evidence returned
Model continues
Final answer: 18

Multi-dimensional evidence
Keep the task, final response and completion state together so a result can be checked in context.
Record tool inputs, outputs and failure states instead of reducing a multi-step run to a single score.
Inspect retries, intervention signals and step count to understand how the agent reached its answer.
Use runtime signals alongside correctness when choosing an agent design for a real workload.
Use EvalScope’s native loop evaluation with configurable tools, environments and step limits.
Bridge external harnesses and preserve their scenario provenance, trajectory artifacts and validity state.
AGENT BENCHMARKS & DRIFT
Benchmark tasks make agent behavior comparable today. Freezing artifacts, auditing judges and testing held-out tasks are an exploring research protocol — not a shipped drift-prevention claim.
Keep task, environment, requests and raw artifacts together.
Separate a scoring-contract change from a behavior change.
Use held-out evidence before claiming an improved harness.