Versioned semantics
Benchmark metadata publishes evaluation versions, so prompt, mapping, data or scoring changes can be visible.
MODEL EVALUATION
Evaluate LLMs, VLMs, AIGC and RAG workloads with benchmark metadata, sample reviews, scoring contracts and reports that retain context.

BROAD MODEL COVERAGE
Models · tasks · datasets · metrics
The selected dataset and version define what the run is allowed to measure.
Transparent · reproducible · extensible
Benchmark metadata publishes evaluation versions, so prompt, mapping, data or scoring changes can be visible.
Structured judge outputs fail closed when invalid; malformed replies do not quietly become scores.
Predictions, reviews, reports and configs remain together for local inspection and dashboard drill-down.