Skip to main content
Goal: a scripted evaluation showing exactly what your ontology changed — the artifact that turns “trust me” into a table.

5.1 · The design

  • Ground truth first. For each eval question, compute the expected answer with direct SQL on the warehouse — before any agent runs. Where the data can’t answer (coded fields with no dictionary), the expected answer is an honest “unanswerable as posed.”
  • Baseline with the ontology OFF. Run every question through the platform with schema inference only, and capture answers + SQL. Sequence matters: capture the control before ontology content enters the workspace.
  • Install (Module 4), re-run identical questions, diff.

5.2 · Script it

Drive the runs headlessly — each question one chat via the platform API with the ontology tool toggled, or via refinery agent run for agent-shaped evals; write results to a JSONL you can resume. Score against ground truth; grade judgment questions by hand.

5.3 · What lift actually looks like

  • Routing fixes — questions that previously searched the wrong schema and declared data missing now find it (a routing README is often the single highest-lift file).
  • Definitional pins — “total X” stops having two defensible answers; responses cite the governed definition they used.
  • Honest refusals — where no governed definition exists, the agent says so and asks for the dictionary instead of fabricating. For evaluators whose criteria are accuracy and trust, that’s a feature.
  • And an honest ledger — expect the baseline to beat your first-draft ground truth on a question or two. Revise the ground truth and say so; the eval is scored against data, not pride.

✅ Checkpoint

  • Ground truth existed before the first agent run, honest “unanswerable” entries included
  • Baseline captured with the ontology off, before install
  • The lift table separates routing fixes, definitional pins, and honest refusals