VISUALIZATION

Inspect the evidence behind a metric.

EvalScope reports connect aggregate results with sample reviews, agent details and serving charts. Run locally; no hosted customer data is required.

Abstract visualization evidence path

FROM OUTPUTS TO INSIGHTS

One local dashboard, one evidence trail.

Same reports · better stays with you

1

Outputs directory

outputs/

2

Scan runs

Discover evaluations.

3

Compare

Contrast results.

4

Inspect samples

Open reviews.

5

Share report

HTML / JSON

FOUR EVIDENCE VIEWS

Evaluation

Aggregates, comparisons and benchmark reports.

Agent

Task details, model requests and tool event evidence.

Performance

Load curves, latency, TTFT and throughput charts.

Run artifacts

Configuration, raw outputs and shareable HTML reports.

Four views, one evidence trail.

Real Dashboard screenshots

Light Dashboard overview and trend screenshot
Light Performance runs screenshot
Light evaluation summary screenshot
Light Agent Trace screenshot

Aggregate → sample → artifact.

Use the report to locate the sample, then its prediction and review. The same output folder preserves configs and raw artifacts for a local audit.

reports/predictions/reviews/configs/
open-local-dashboard.sh
pip install 'evalscope[service]'
evalscope service --outputs ./outputs

LOCAL EVIDENCE, NOT A DETACHED CHART

Keep the chart tied to the run that produced it.

A report is useful when you can move from a number to its samples, config and raw artifacts without losing the thread.

See the evaluation path
01reports/Aggregate decisions
02reviews/Sample evidence
03configs/Runnable context
04outputs/Raw record