Skip to main content
This page takes you from an empty project to a saved run you can review in the web UI.
1

Install

Install the package for your language. Both ship the same ezvals CLI and web UI.
The Python package has no Python dependencies. The npm package needs Node 22.18 or newer, which runs .ts eval files directly with no build step. In a TypeScript project, run every command on this page as npx ezvals.
2

Let your coding agent write the evals

Install the EZVals skill, which teaches your agent how to plan evals, pick graders and read results:
Then ask for what you want, for example:
The agent writes the eval files, runs them, and reports back. See Using with coding agents for other install options and prompts.
3

Or write one by hand

Python evals can live in any .py file (files starting with _ are skipped). TypeScript evals live in files ending in .eval.ts (or .eval.mts, .eval.js, .eval.mjs).
4

Run it

ezvals run prints a short per-eval summary with a pass or fail mark for each eval, followed by a failures section with each failing eval’s reason, and saves the run under .ezvals/sessions/. Add --json to print the full run as JSON, which is what coding agents usually read.
5

Review it in the browser

This opens the web UI at http://127.0.0.1:8000. It lists every eval it found but runs nothing until you press the Run button. Open a result to see its input, output and scores. See Web UI.
6

Iterate and compare

Change your agent or prompt, run again, and compare the two runs side by side in the UI. Giving runs names in a shared session makes this easy:
Every run starts fresh worker processes, so edits to eval code are always picked up. See Sessions & comparing runs.

What’s in your project now

An optional ezvals.json holds defaults such as concurrency and timeout. It’s created only when you save settings from the UI, or you can write it yourself. See Configuration.

Next

How EZVals works

The mental model behind evals, runs and sessions.

Writing evals

Inputs, outputs, context, timeouts and errors.