Model risk is the risk of adverse consequences from decisions based on incorrect or misused model output. That definition, drawn from banking supervision, is worth borrowing intact, because it contains the two failure modes that actually occur: the model is wrong, or the model is right and used for something it was never built to do. The second is more common and less discussed.
Three lines, sized for a real department
The three-lines model is often dismissed as unaffordable below a certain size. In practice it can be implemented by three named people rather than three departments.
- First line — the owner. The person who uses the output and is accountable for the decision. They maintain the model inventory entry, document assumptions, and own the override log.
- Second line — the reviewer. Someone in the finance function who did not build the model and does not use it operationally. Their job is to confirm that the documented purpose matches the actual use, and that the assumptions still hold.
- Third line — internal or external audit. Periodic independent assessment, ideally aligned to the audit plan already in place.
For entities without internal audit, the second-line role can sit with a finance committee member or a peer entity under a reciprocal arrangement. Independence matters more than seniority.
The inventory is the foundation
Nearly every failure investigation begins with the discovery that nobody had a list. A model inventory for a mid-sized entity is a spreadsheet with one row per model and roughly a dozen columns: name, owner, purpose, tier, inputs, source systems, output consumers, last validation date, next validation date, known limitations, override rate, and vendor.
Two entries are almost always missing when a first inventory is compiled: the revenue forecast workbook that has been maintained by one person for eleven years, and whatever the actuary supplies for pension and OPEB. Both are models. Both carry model risk. Neither is usually treated as such.
Validation without a quantitative team
Validation does not require a doctorate. It requires four activities, performed honestly.
Conceptual soundness
Does the method fit the question? A model trained on pre-2020 spending patterns to detect anomalies in a post-2020 grant environment has a conceptual problem no amount of tuning fixes.
Outcome analysis
Compare predictions to what happened. For triage models, this means pulling a random sample of unflagged items and reviewing them. Reviewing only flagged items measures precision and tells you nothing about what the model missed, which is the number that matters.
Benchmarking
Compare against a deliberately naive alternative: a simple threshold rule, last year's value, a random sample of equal size. A surprising share of production models fail this comparison. That is useful information, not an embarrassment — a simple rule that performs as well is cheaper to run, easier to explain, and easier to audit.
Stability
Check whether the input distribution has shifted. A chart-of-accounts restructuring, an ERP migration, or a change in procurement thresholds will all move the distribution, and the model will keep producing confident output regardless.
Override rate as a leading indicator
Track the proportion of recommendations reversed by reviewers, segmented by reviewer and by category. A rising override rate concentrated in one category usually signals drift. A high override rate spread evenly usually signals that the model was never fit for purpose. A near-zero override rate is not good news — it typically means reviewers have stopped reviewing.
Documentation that survives staff turnover
The practical standard is that a competent successor should be able to reproduce the model's output, understand its limitations, and decide whether to keep using it, without speaking to the person who built it. That means recording the data sources and extraction logic, the transformations, the parameters, the evaluation results with dates, the known failure modes, and the decisions that were considered and rejected. The last item is the one that saves successors the most time and is almost never written down.
Reporting upward
A quarterly model risk summary to the finance committee that runs to one page: number of models by tier, validations completed and overdue, override rates against threshold, incidents, and changes to the inventory. Governing bodies rarely want more than this, and producing it forces the inventory to stay current — which is, in the end, most of the benefit.
This publication is general information and is not legal, accounting, audit or financial advice. See our Disclaimer. Found an error? Write to [email protected] — we correct in place and note what changed.