npx (npx ezvals run evals/).
ezvals COMMAND -h prints a command’s flags.
ezvals run
.py files and *.eval.ts, *.eval.mts, *.eval.js and *.eval.mjs files, and evals in every language found run together in one run. Python files whose names start with _ are skipped, as are hidden directories, directories starting with __, node_modules and venv.
Output. The human-readable output goes to stderr, grouped by file. Each eval function gets one line, as soon as all its results are in, with a mark for its outcome:
Each line also shows how many of the eval’s cases or trials passed, score averages where they apply, and the latency. A failures section follows, listing each failed or errored result (up to 20) with its failing scores’ notes or its error, input and output. A summary at the end counts results by outcome, gives each score’s average (plus pass@k and pass^k for trials), and names the saved run file.
With
--json, stdout is only the run JSON, so a script or coding agent can read it while you still see the human output.
Exit codes. Failed evals don’t make the command fail. Read the JSON to decide whether a CI job should fail.
ezvals serve
127.0.0.1 only. If the port is taken, it tries the next nine ports.
When you open a saved run whose eval files still exist, you can keep running evals in it. Otherwise it opens read-only.
ezvals regrade
ezvals query
--json prints them as a JSON array. See Querying results for the tables.
ezvals export
ezvals skills
add needs at least one of --agents, --claude, --codex, --cursor, --windsurf, --kiro or --roo. Every subcommand accepts -g, --global to use your home directory instead of the project. See Using with coding agents.
Configuration
ezvals.json in the directory you run from sets defaults for run and serve. Command-line flags take precedence over it. The file is created when you save settings in the web UI, and you can also write it yourself.
ezvals.json
concurrency, timeout and trials also apply to runs started from the web UI, and concurrency applies to regrades.
