# ML Model Review Checklist

**Model:** _______________  **Owner:** _______________  **Date:** _______

## 1. The problem

- [ ] The prediction target is defined precisely, including edge cases
- [ ] **A trivial baseline exists and has been measured** — always predict the majority class,
      predict yesterday's value, use the existing rule
- [ ] The model beats that baseline by a margin worth the operational cost
- [ ] The decision made from the prediction is documented

**Baseline score:** ____ **Model score:** ____ **Difference worth the complexity?** yes / no

🔴 A model that beats a trivial baseline by two points is usually not worth the pipeline,
monitoring and on-call that come with it. Measure the baseline first, not last.

## 2. Data leakage — the defect that produces excellent results and useless models

- [ ] **No feature contains information unavailable at prediction time**
- [ ] No feature derived from the target
- [ ] Time-ordered data split by time, never randomly
- [ ] Preprocessing fitted on training data only, then applied to validation
- [ ] Duplicates removed before splitting, not after
- [ ] Groups (customer, device, session) kept entirely within one split

**Suspiciously high score?** Assume leakage before assuming success. It is the single most common
cause of a model that performs brilliantly in evaluation and fails immediately in production.

## 3. Validation

- [ ] Held-out test set, used **once**, at the end
- [ ] Validation strategy matches how the model will be used
- [ ] Cross-validation where data is limited
- [ ] Performance reported with a range, not a single number
- [ ] Results reproducible from a clean run — seed and versions pinned

## 4. Metrics

- [ ] The metric matches the business consequence
- [ ] **Class imbalance accounted for** — accuracy on a 99:1 split is meaningless
- [ ] False positives and false negatives costed separately
- [ ] Threshold chosen deliberately, not left at 0.5
- [ ] Performance checked **per segment**, not only in aggregate

| Segment | Score | Volume |
|---|---|---|
| | | |

An aggregate score can hide a model that fails badly for a specific group. Check before someone
else does.

## 5. Fairness and explanation

- [ ] Checked for disparate performance across groups where relevant
- [ ] Proxies for protected characteristics considered (postcode, education, gaps)
- [ ] Can explain an individual prediction if challenged
- [ ] Decisions with legal or significant effect on people have human review

## 6. Production

- [ ] **Training and serving use the same feature computation** — a separate reimplementation is
      the classic source of silent skew
- [ ] Input validation at the boundary; behaviour defined for out-of-range values
- [ ] Latency and cost per prediction measured
- [ ] Model version recorded with every prediction
- [ ] Predictions logged with inputs for later analysis
- [ ] Fallback when the model is unavailable

## 7. Monitoring

- [ ] Input distribution monitored for drift
- [ ] Prediction distribution monitored
- [ ] **Actual outcomes captured, so accuracy can be measured in production**
- [ ] Alert thresholds set, routed to a named person
- [ ] Scheduled review, because the world changes even when the code does not

**How will we know this model has degraded?** _______________

If the answer is "someone will notice", it will degrade unnoticed for months.

## 8. Retraining

- [ ] Trigger defined — schedule, drift, or performance drop
- [ ] Retraining is reproducible
- [ ] New model compared to current before replacing it
- [ ] Rollback to the previous model tested

## Sign-off

| | Name | Date |
|---|---|---|
| Built by | | |
| Reviewed by | | |
| Approved for production | | |
