Skip to main content
Cases work when every example is checked the same way. Agents that use tools often need a different check per task: one reads a live API for the right answer, another checks a side effect such as a calendar booking. In EZVals, each of those is its own eval function.

The evals

All four evals share one target, and each checks its task in its own way.

Why it’s built this way

  • File defaults set the target and dataset once, so each eval is only its input and its check.
  • The reference is set inside the eval when the right answer is only known at run time, such as live weather or stock prices. It’s saved with the result, so you can see what the agent was compared against.
  • Side effects are checked directly. The booking eval asks the calendar what happened rather than trusting the agent’s reply, and stores the booking in metadata for debugging.
  • All four share a dataset, so they appear together in the results and in the summary.

Run it

Use separate functions like this when each task needs its own data source or its own check. Use cases when many inputs share one check.