Skip to main content
This is the full Python API, for when you or your coding agent need to look something up. The guides under Writing evals explain when to use each part.

@eval

Marks a function as an eval. It works bare (@eval) or with options (@eval(...)). The function can be sync or async. Give it a parameter annotated EvalContext to receive the context. An eval with target, cases or input_loader requires that parameter. Files are discovered by ezvals run and ezvals serve. See CLI for which files are included.

Eval options

Return value. Return nothing (the context becomes the result), an EvalResult, or a list of EvalResult for several results. Anything else is an error. File defaults. A module-level ezvals_defaults dict sets any option except cases and input_loader for every eval in the file. See File defaults.

EvalContext

ctx.store(...)

Sets every argument that isn’t None, and returns the context.
  • scores: one score or a list. A score with an existing key replaces it.
  • metadata and trace_data merge into what’s already there.
  • messages and trace_url set trace_data.messages and trace_data.trace_url.

Score

A dict with a key and at least one of passed or value: True/False and bare numbers are shorthand for {"passed": ...} and {"value": ...} with the default key. See When a result passes.

EvalResult

What an eval produces and what evaluators receive. A dataclass with these fields: An evaluator can also return a new EvalResult, which replaces the result.

TraceData

A dict that also supports attribute access. Two keys are shown specially in the web UI:
  • messages: a conversation transcript, rendered as chat. Most common message formats are recognized.
  • trace_url: a link to the trace in another tool.
Other keys are shown as JSON.

EvalCase

A TypedDict for type-checking cases and input_loader output. The keys are id, input, reference, dataset, labels, metadata, default_score_key, timeout, target, evaluators and trials. An unknown key is an error.

run()

Runs evals the way ezvals run --json does, saves the run, and returns the run as a dict, including saved_path.
Raises ValueError with the CLI’s error message if the run can’t start.

run_evals()

Runs eval functions or paths in the current process and returns a list of EvalResult. Nothing is saved. Useful in notebooks and tests.
evals is a list that can mix @eval functions and file or directory paths.