validation/golden-queries.md — verified against a synthetic wealth warehouse on Databricks (horizon 2026-05-31, window 2025-06-01…2026-05-31). Yours will differ; the point is that the same call gives the same number every time.
5.1 · AUM (the firm-wide number)
Prompt
You’ll see: AUM from the
account_value series at one valuation date (golden: **58.6B) used for billing. point_in_time ≠ average_12m; the surface pins both (notes/aum-definition.md).5.2 · Net flows (organic growth)
Prompt
You’ll see: the organic-growth number isolated from market beta (golden: net **−648,148,543 + out −$2,320,157,552), sourced from
account_value.net_external_flow (notes/aum-definition.md).5.3 · Return (TWR vs MWR)
Prompt
You’ll see: the GIPS-standard TWR (golden: +10.15%), with client flows geometrically removed, and MWR (Modified Dietz) exposed explicitly — never silently swapped. They answer different questions: manager skill vs. the client’s dollar experience (
notes/return-definition.md).5.4 · Allocation and concentration
Prompt
You’ll see: allocation weights that sum to 1.000 (golden: US_EQUITY 0.4436 · FIXED_INCOME 0.3446 · INTL_EQUITY 0.1571 · CASH 0.0547), and a concentration breach list (golden: 201,120 holdings >10% across 60,701 accounts, max weight 0.8601) — both point-in-time, both vs. the account denominator.
5.5 · Effective fee rate (revenue quality)
Prompt
You’ll see: realized yield as fee revenue / average AUM × 10,000 (golden: 58,582,267,503 = 52.34 bps), with the billing-basis average AUM denominator — not point-in-time (
notes/fee-definition.md).Why everyone gets the same number — Metrics like AUM, return, and fee yield can be computed several ways (point-in-time vs average AUM, TWR vs MWR, gross vs net). The ontology pins one governed definition — with the decision recorded in
ontology/notes/ — so Finance, the CIO office, and the advisors stop disagreeing.5.6 · When the answer isn’t governed yet — watch the model grow
Now ask something from your shortlist that the starter doesn’t already cover. This is the important beat: a starter pack is a head start, not the finished model. (Even the synthetic dataset has honest gaps — no benchmark-return series, and undated closures — recorded openly ingolden-queries.md and schema-mapping.md.)
Prompt
You’ll see: Ana explore only the frontier (not re-derive the whole warehouse), answer, and propose a write-back — a new metric committed to your repo with provenance. Review and merge it, and the next person who asks gets the governed answer for free. That’s the malleable loop: the ontology you ship is the one you grow, and it gets more complete every time you use it.
✅ Checkpoint
- All five metric families answered through governed surfaces, SQL shown
- You can point to the notes file explaining at least one metric’s definition decision
- Ana proposed a write-back for a not-yet-governed question — and you saw it land as a PR