Set trials
On an eval, a case or in file defaults:ezvals run --trials 5, "trials": 5 in ezvals.json, or the Trials setting in the web UI. A run-wide value overrides per-eval trials.
Read the results
Each trial is its own result, with an id likeevals.py::test_booking~2 and the fields trial (starting at 1) and trial_of (the eval’s id). You can open, annotate and rerun trials one at a time. A trial passes by the usual rule: no error, at least one pass/fail score, and all pass/fail scores passed.
Runs with trials report two extra numbers, in the UI stats and in the run JSON as pass_at_k and pass_all_k:
A big gap between the two means the agent can do the task but doesn’t do it consistently. That usually points to an ambiguous prompt, sampling temperature or a brittle tool, not a missing capability. To list the flaky evals, query the results.

