5.1 · The design
- Ground truth first. For each eval question, compute the expected answer with direct SQL on the warehouse — before any agent runs. Where the data can’t answer (coded fields with no dictionary), the expected answer is an honest “unanswerable as posed.”
- Baseline with the ontology OFF. Run every question through the platform with schema inference only, and capture answers + SQL. Sequence matters: capture the control before ontology content enters the workspace.
- Install (Module 4), re-run identical questions, diff.
5.2 · Script it
Drive the runs headlessly — each question one chat via the platform API with the ontology tool toggled, or viarefinery agent run for agent-shaped evals; write results to a JSONL you can resume. Score against ground truth; grade judgment questions by hand.
5.3 · What lift actually looks like
- Routing fixes — questions that previously searched the wrong schema and declared data missing now find it (a routing README is often the single highest-lift file).
- Definitional pins — “total X” stops having two defensible answers; responses cite the governed definition they used.
- Honest refusals — where no governed definition exists, the agent says so and asks for the dictionary instead of fabricating. For evaluators whose criteria are accuracy and trust, that’s a feature.
- And an honest ledger — expect the baseline to beat your first-draft ground truth on a question or two. Revise the ground truth and say so; the eval is scored against data, not pride.
✅ Checkpoint
- Ground truth existed before the first agent run, honest “unanswerable” entries included
- Baseline captured with the ontology off, before install
- The lift table separates routing fixes, definitional pins, and honest refusals