Skip to main content
Before you switch models or ship a new prompt, run your existing evals against both versions and compare the results eval by eval. Run configs do this without changing any eval code.

1. Read the model from the config

Your eval, or its target, reads the settings from ctx.config instead of hard-coding them:

2. Define a config per model

ezvals.json
A config can hold anything your code reads: a model, a prompt version, a retrieval setting.

3. Run each config in one session

Each run is named after its config. In the web UI you can instead choose the config before starting each run.

4. Compare

The comparison view shows the pass rate, each score and latency per run, and each eval’s outputs side by side. Look at the evals whose outcome changed, not only at the totals. To do the same from the terminal, query the session.
Models are nondeterministic, so a difference of one or two evals may be noise. Add --trials 3 to both runs and compare pass^k to see which model is reliably better.

Alternative: models as cases

To get every model in a single run, put the model in each case instead. A case’s input replaces the eval’s input rather than merging with it, so put the whole input, prompt included, in every case, and have the target read the model from ctx.input:
This is handy for a quick check. Run configs scale better: every eval runs against every model without editing each eval, and the comparison view lines up the runs for you.