Skip to main content
Use cases when many inputs share the same check. You write the logic once, and each row becomes its own result that you can filter, rerun and compare.

Define cases

Pass cases, a list of dicts (plain objects in TypeScript). Each case sets input and reference on ctx, and an optional id names it:
This produces test_sentiment[pos], test_sentiment[neg] and test_sentiment[2]. A case without an id is numbered by its position. When a case needs more than one value, make input a dict: {"input": {"text": "...", "expected_intent": "refund"}}.

What a case can override

A case can set any eval option except cases and input_loader: input, reference, dataset, labels, metadata, default_score_key, timeout, trials, target and evaluators.
  • A field the case sets replaces the eval’s value. This includes input: a case’s input doesn’t combine with the decorator’s input.
  • labels and metadata merge with the eval’s values. On a metadata key clash, the case wins.
  • Setting a field to None (null) clears the eval’s value.

Grids

Build a grid with a list comprehension (flatMap in TypeScript):
To compare whole runs rather than rows, such as one run per model, use run configs instead.

Load cases at discovery time

input_loader (inputLoader) is a function, sync or async, that returns cases in the same shape. It runs when evals are discovered, not when the file is imported. That makes it a good fit for fetching from a database, a dataset API or a file:
A loader can’t be combined with input, reference or cases. If it raises, the eval records a single errored result that starts with Input loader failed:.

Run specific cases

name@id means the same as name[id] and doesn’t need shell quoting.