GET STARTED

Start with your coding agent, safely.

Install the EvalScope project skill, then describe the task you want to complete. Your agent will clarify the model, endpoint, and other context it needs before running it.

Ask your agent to install the Skill

Install the EvalScope project skill from https://github.com/modelscope/evalscope/tree/main/skills/evalscope. Read it after installation, then tell me what you need from me before using EvalScope.

CHOOSE YOUR TASK

Give your coding agent a bounded job.

Same principle · more possibilities

Evaluate a model

Run a bounded evaluation.

Choose the model and evaluation context with the agent, then review the saved results.

Use the EvalScope project skill to evaluate a model served through an OpenAI-compatible API on five GSM8K samples. Ask me for the model name, endpoint, and credentials configuration you need before running it, then report the results and where the outputs were saved.

COMMAND LINE

Run a small evaluation from the terminal.

Install the service tools, run five samples against an OpenAI-compatible endpoint, or select the matching eval type for another supported protocol; then open the saved outputs locally.

01

Install service toolsAdds the CLI and local dashboard.

02

Run five samplesUses the endpoint and API key supplied by your provider.

03

Inspect artifactsOpens reports, predictions and configs from ./outputs.

quickstart.sh
pip install 'evalscope[service]'
export MODEL_API_KEY='…'
evalscope eval \
  --model your-model --eval-type openai_api \
  --api-url https://api.example.com/v1 \
  --api-key "$MODEL_API_KEY" \
  --datasets gsm8k --limit 5
evalscope service --outputs ./outputs

Troubleshoot only what changed.

My API endpoint uses the OpenAI protocol.

Use the API-first route: set the URL, model name, and credential environment variable from your service. For another protocol, select its matching eval type.

I do not have a GPU.

Nothing on this page requires one. Use a remote compatible API and avoid local checkpoint model types.

Where are my results?

EvalScope prints an output directory. Preserve it; reports, predictions, reviews and configs are organized under that run.