Skip to main content

A workshop of short, hands-on modules for data engineers and analytics engineers who need to trust — and prove — answer accuracy. The thesis throughout: accuracy is checked, not asserted. Golden datasets pinned to numbers the org already trusts, eval cases that exercise edge conditions, drift detection on a schedule, reconciliation workflows for when systems disagree, and a validation agent that runs the moment data lands.

Total time: about 2 hours.

Builds on Build Your Ontology, End to End and Ontology Operations — this workshop assumes governed definitions exist and doesn’t re-teach the build; it teaches the proof.

Who this is for

  • Data/analytics engineers accountable for “is this number right?”
  • Ontology owners who need tests, not vibes, behind their definitions
  • Teams burned by a silent upstream change that shipped wrong numbers for a week
  • Anyone preparing for an audit, a board cycle, or a regulator

What you’ll be able to do

After moduleYou can…
0Map every way an answer goes wrong to the defense that catches it
1Build golden queries pinned to numbers the org already trusts
2Generate and curate eval cases that exercise definitions and edge cases
3Run drift detection on a schedule, distinguishing data drift from definition drift
4Reconcile two disagreeing systems to the exact fork, and record the resolution
5Deploy a data-quality agent triggered by ETL completion
6Write the quality runbook: owners, cadences, and what a red alert triggers

Workshop modules

ModuleTime
0 · The Trust Stack10 min
1 · Golden Datasets20 min
2 · Validation Sets & Eval Cases20 min
3 · Drift Detection on a Schedule20 min
4 · Reconciliation Workflows20 min
5 · The Data-Quality Feed Agent15 min
6 · The Quality Runbook15 min

How to use this workshop

  1. Prerequisites: a connected warehouse and an existing ontology with governed metrics (the build workshops above if not).
  2. Work in order — golden datasets (Module 1) are the foundation every later module schedules, extends, or operationalizes.
  3. Swap [bracketed] placeholders for your metrics, tables, and trusted sources.

T-minus-1-day pre-flight: run the workshop’s key prompts end-to-end in the session workspace — confirm logins, connectors, and the features this workshop touches are enabled for every attendee; have a fallback demo workspace ready in case a customer connector fails live. Pacing: when a long prompt is running (2–4 min), fire it first, then discuss the concept while Ana works — never watch a spinner in silence. If behind schedule, cut sections marked optional, never the checkpoints. Group sessions: attendees drive, you narrate; collect every miss or wrong answer in a shared doc — that list is the ontology backlog. With 10+ attendees on one warehouse, stagger the heavy prompts.