Skip to main content
One run can make a flaky agent look solid, or a solid one look broken. Trials run each eval N times and tell you both “can it do this?” and “does it do this every time?”.

Set trials

On an eval, a case or in file defaults:
To run every eval N times without changing code, use ezvals run --trials 5, "trials": 5 in ezvals.json, or the Trials setting in the web UI. A run-wide value overrides per-eval trials.

Read the results

Each trial is its own result, with an id like evals.py::test_booking~2 and the fields trial (starting at 1) and trial_of (the eval’s id). You can open, annotate and rerun trials one at a time. A trial passes by the usual rule: no error, at least one pass/fail score, and all pass/fail scores passed. Runs with trials report two extra numbers, in the UI stats and in the run JSON as pass_at_k and pass_all_k: A big gap between the two means the agent can do the task but doesn’t do it consistently. That usually points to an ambiguous prompt, sampling temperature or a brittle tool, not a missing capability. To list the flaky evals, query the results.
Trials multiply cost and time. Use 3 to 5 trials on evals you suspect are flaky, or pass --trials for an occasional reliability check, and raise --concurrency to match.