Skip to main content
Some qualities, such as tone, helpfulness or following a policy, can’t be checked with string matching. An LLM judge can grade them, but only once you’ve checked that it agrees with you.

Write the judge as an evaluator

Writing the judge as an evaluator lets you reuse it on any eval, and its reasoning is saved as the score’s notes.
To judge every eval in a file, put the judge in file defaults as evaluators. Tips for judge prompts:
  • Ask for pass/fail against clear criteria rather than a 1 to 10 rating. Pass/fail is easier to agree with and to act on.
  • Ask for the reason, and store it in notes, so you can tell a wrong answer from a wrong judge.
  • Judge one quality per judge. Use several judges with different keys instead of one prompt that checks everything.

Calibrate it

A judge is only useful if it agrees with you. Because the agent call lives in a target, you can refine the judge against the same stored outputs:
1

Run once

ezvals serve evals/ and run the evals. This is the only step that calls your agent.
2

Review

Read the results in the web UI. Where the judge is wrong, correct the score and add an annotation explaining why.
3

Tighten the prompt

Change the judge prompt, or ask your coding agent to: “Read my score corrections in the last run and update the judge prompt.”
4

Regrade

Choose Regrade in the UI, or run ezvals regrade <run_id>. The judge scores the stored outputs again.
Repeat until the judge’s scores match yours, then trust it on new runs.
Regrading replaces your score corrections with the judge’s new scores. Annotations are kept, and each correction stays in the result’s correction_history, so you can still check whether the new judge agrees with you.