Skip to main content
A RAG agent can fail in two ways: it gives the wrong answer, or it gives the right-sounding answer without support in what it retrieved. This recipe scores both across a list of question and answer pairs.

The eval

grounding_judge and correctness_judge stand in for your own LLM calls. See LLM judge for how to write and calibrate them.

Why it’s built this way

  • Cases turn a list of questions into one result each. Real suites often have hundreds, which you can load from a file with an input loader.
  • The target calls the agent and keeps the retrieved documents in metadata. They show up next to each result in the web UI, so when an answer is wrong you can see straight away whether retrieval or generation failed.
  • Two named scores keep the failure modes apart. The summary shows a bar for each, so “grounding dropped, correctness held” is visible at a glance.
  • Because the agent call is in a target, you can change either judge and regrade without querying the agent again.

Run it

Ask your coding agent to look at the failures: “Which questions failed grounding, and what did retrieval return for them?”