Skip to main content
Put your agent call in a target and your checks in the eval body. You can then share one agent call across many evals, and change graders or judge prompts without paying for new agent runs.

Targets

A target runs before the eval body and receives the same ctx. It can set fields with ctx.store(...), or just return the output. A returned value becomes ctx.output.
You can also set target in file defaults, so every eval in a file uses the same agent call. A single case can override it too.

Regrading

Regrading scores a finished run’s stored outputs again, using your current eval code. The eval body and evaluators run again, but the target is skipped. Instead, ctx starts with the stored input, output, latency, metadata and trace_data, as the target left them. This works the same in Python and TypeScript.
In the web UI, regrade the whole run or the selected rows with Regrade in the header, or a single result from its detail page. The run is updated in place: new scores replace the old ones. Manual score edits are replaced too, but annotations are kept. These results are skipped, and the skipped count is reported:
  • results from evals without a target
  • results that errored or never finished
  • results from evals that return several results at once
Regrading runs the eval code as it is now, so the eval file must still exist.
A typical loop:
  1. Run the agent once.
  2. Read the failures.
  3. Tighten the judge prompt or assertion.
  4. ezvals regrade.
  5. Repeat until the scores match your own judgment.
Only then run the agent again.