AGENT & HARNESS

Measure the loop, not just the last answer.

Capture model requests, tool calls, token usage and final outcomes for agent tasks — with a path for external replay scenarios too.

Native agent evaluationExternal scenariosFunction calling
Abstract agent trace loop

MEASURE · TRACE · IMPROVE

A measured agent loop.

Every step counts

1

Prompt

Task and context

2

Model

qwen-plus

3

Tool call

python_exec

4

Tool result

Evidence returned

5

Nudge

Model continues

6

Submit

Final answer: 18

Existing EvalScope Agent trace with Python tool calls and submitted answer
Existing local Agent run · Python tool trace, submitted answer and verifiable result

Key measurements for agent evaluation.

Multi-dimensional evidence

01
Outcome evidence

Inputs and outcomes

Keep the task, final response and completion state together so a result can be checked in context.

02
Tool evidence

Calls and results

Record tool inputs, outputs and failure states instead of reducing a multi-step run to a single score.

03
Trajectory health

Steps and recovery

Inspect retries, intervention signals and step count to understand how the agent reached its answer.

04
Runtime profile

Tokens, timing and cost

Use runtime signals alongside correctness when choosing an agent design for a real workload.

Available Today

Native AgentLoop

Use EvalScope’s native loop evaluation with configurable tools, environments and step limits.

function_callingreacttool use
Available Today

External Agent Bridge

Bridge external harnesses and preserve their scenario provenance, trajectory artifacts and validity state.

AgentXAIPerfreplay

AGENT BENCHMARKS & DRIFT

Keep an evolving harness tied to its evidence.

Benchmark tasks make agent behavior comparable today. Freezing artifacts, auditing judges and testing held-out tasks are an exploring research protocol — not a shipped drift-prevention claim.

SWE-benchGAIAAgentBenchExploring
01

Freeze the run

Keep task, environment, requests and raw artifacts together.

02

Audit the judge

Separate a scoring-contract change from a behavior change.

03

Check selection

Use held-out evidence before claiming an improved harness.

Keep the research boundary visible.

Explore research protocol