> ## Documentation Index
> Fetch the complete documentation index at: https://docs.textql.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Module 7 · Validate Numbers & Make It Yours

> Trust in insurance analytics is earned on the first matching number — and lost the first time a loss ratio is quoted on the wrong basis. The starter’s golden values are already … (~20 min)

## 7.1 · Reconcile against a number someone already trusts

Trust in insurance analytics is earned on the **first matching number** — and lost the first time a loss ratio is quoted on the wrong basis. The starter's golden values are already **pinned and verified** against the synthetic warehouse (see `validation/golden-queries.md` — loss ratio 0.5469, combined 0.7969, reserve ratio 0.7143, retention 0.5004). Against **your** warehouse, reconcile each governed surface to a number an actuary or finance lead already trusts:

```text Prompt theme={null}
Run each governed surface against my warehouse and compare to a reference number I trust (loss ratio, combined ratio, reserve position, retention). For each, show the SQL and the basis, and flag any drift. Where we differ, explain whether it's data, definition, or basis (paid vs incurred, written vs earned, accident vs calendar year).
```

<Check>
  **You'll see:** accuracy checked, not asserted — and a triage of any mismatch into data vs. definition vs. basis. The decisive moment is the first time the accident-year loss ratio lands exactly where the actuary expected.
</Check>

## 7.2 · Assert the invariants

Even before you have an external reference, some numbers must agree with each other. The golden queries assert these:

```text Prompt theme={null}
Check the cross-surface invariants from validation/golden-queries.md against my data: combined_ratio == loss_ratio + expense_ratio; incurred >= paid; reserve_position.incurred == paid_to_date + case_reserves; PIF agrees between claim_frequency and policies_in_force; loss_ratio and retention land in a sane range. Report any that don't hold.
```

<Check>
  **You'll see:** internal consistency proven — if combined ≠ loss + expense, or PIF disagrees between two surfaces, something is wrong *before* a stakeholder ever sees the number.
</Check>

## 7.3 · Customize a definition — the written-vs-earned premium lesson

Your carrier inevitably defines something differently — a loss-ratio basis (net vs. gross, with or without LAE), a retention basis, an in-force definition. But the starter's **flagship field lesson** is one every carrier hits, and it makes the perfect worked example because it's a *real bug that was found and fixed* in this very repo:

<Warning>
  **The \~12× premium-inflation bug** — The reference model pre-earns premium into a **monthly** `premium_earned` fact — one row per policy per month. **Written premium** is booked *once* at inception for the whole term, so it repeats across every one of those monthly rows. An early version of `earned_premium` did `SUM(written_premium)` off that monthly series — which **inflated written premium \~12×** (roughly one duplicate per month of term). The fix: report **earned** premium (sum the monthly earned series — that *is* correct, it recognizes over time) and pull **written** premium from the **policy grain**, never the earned series. Documented in `notes/premium-definition.md` and `validation/golden-queries.md` → *Issue found & fixed*.
</Warning>

```text Prompt theme={null}
Walk me through the written-vs-earned premium distinction in notes/premium-definition.md. Then check my warehouse: is written_premium stored on the monthly premium_earned series? If so, prove the trap — show that SUM(written_premium) off the monthly fact inflates ~12x vs. written premium counted once at the policy grain. Confirm earned_premium.tql sums the EARNED series for ratios and sources WRITTEN from the policy grain. If our carrier defines the loss-ratio basis differently (net of reinsurance, w/ or w/o LAE), update the governed surface in our repo, record the decision and the rejected default in a notes file, add a golden-query test pinning the value, and open a PR.
```

<Check>
  **You'll see:** the most expensive ratio bug in P\&C demonstrated on your own data and then guarded against — the earned-vs-written discipline confirmed in the surface, and any carrier-specific basis change landing as a reviewable PR **in your repo** with a pinned golden value. The template stays pristine upstream; your adaptations are yours.
</Check>

## 7.4 · Localize the vocabulary

`ontology/notes/glossary.md` holds the canonical insurance terms — policy, policyholder, coverage, written vs. earned premium, loss/combined ratio, incurred/case/IBNR/ultimate, frequency/severity, retention, PIF, line of business, peril — each with a **variance column** flagging where your carrier or line of business diverges (life "policy ≈ contract on a life" vs. P\&C term policy; gross vs. net of reinsurance; renewal-as-flag vs. successor-policy link; PIF as bound vs. issued vs. paid).

```text Prompt theme={null}
Walk the glossary's variance column for our line(s) of business. For each term that differs at our carrier — earned-premium earning pattern, loss-ratio basis, reserve definitions, retention basis, in-force definition — propose the override in glossary.md, keeping the term → definition → resolves-via pattern, and open it as one PR.
```

<Check>
  **You'll see:** the vocabulary localized in one reviewable pass — so "earned premium," "incurred," and "in force" mean *your* carrier's thing, everywhere, from now on.
</Check>

<Note>
  **Two habits as you make it yours** — **1 · Write for the search box.** As you extend the kit, keep a short README per folder and repeat the phrases your teams actually use (metric names, synonyms, team names) in the prose — future threads find context by *search*, not browsing.<br /><br />
  **2 · Let usage drive the roadmap.** Stand up a weekly gap-review playbook: mine repeated questions, manual SQL, and mid-thread corrections; have Ana draft small reviewable patches; a named owner approves. The kit is the seed — usage is what grows it. (See [Ontology Operations](/workshops/ontology-operations/overview) Module 4.)
</Note>

### ✅ Checkpoint

* [ ] Governed surfaces reconciled to a trusted reference; any drift triaged (data / definition / basis)
* [ ] The written-vs-earned trap demonstrated and guarded against on your data
* [ ] One definition is now yours — PR'd, noted, and pinned with a golden-query test
