The web UI shows one or a few runs at a time. ezvals query loads every saved run into an in-memory SQLite database, so you (or your coding agent) can answer questions across runs: did the new prompt regress anything, which evals are flaky, which ones got slower.
Rows print as a table; add --json to get a JSON array. ezvals query --schema prints the tables and some example queries. Run it from your project root: it reads the runs in .ezvals/sessions/, or under results_dir if you set one in ezvals.json.
Tables
results.passed is 1 when the result passed: no error, at least one pass/fail score, and all of them passed.
results.trial is the trial number, or 0 for evals without trials.
labels, input, output, reference, metadata and trace_data hold JSON. Read them with json_extract, for example json_extract(metadata, '$.model').
- Join
results and scores on run_id and row.
Example queries
Compare pass rates across a session
Failures in the latest run
Average a numeric score by run
Flaky evals
Evals where some trials passed and some didn’t:
You rarely need to write these yourself. Ask your coding agent something like “which evals regressed between the last two runs?” and it will write the query.