Write the judge as an evaluator
Writing the judge as an evaluator lets you reuse it on any eval, and its reasoning is saved as the score’s notes.evaluators.
Tips for judge prompts:
- Ask for pass/fail against clear criteria rather than a 1 to 10 rating. Pass/fail is easier to agree with and to act on.
- Ask for the reason, and store it in
notes, so you can tell a wrong answer from a wrong judge. - Judge one quality per judge. Use several judges with different keys instead of one prompt that checks everything.
Calibrate it
A judge is only useful if it agrees with you. Because the agent call lives in a target, you can refine the judge against the same stored outputs:1
Run once
ezvals serve evals/ and run the evals. This is the only step that calls your agent.2
Review
Read the results in the web UI. Where the judge is wrong, correct the score and add an annotation explaining why.
3
Tighten the prompt
Change the judge prompt, or ask your coding agent to: “Read my score corrections in the last run and update the judge prompt.”
4
Regrade
Choose Regrade in the UI, or run
ezvals regrade <run_id>. The judge scores the stored outputs again.Regrading replaces your score corrections with the judge’s new scores. Annotations are kept, and each correction stays in the result’s
correction_history, so you can still check whether the new judge agrees with you.
