# Model Card

> Fill this in before deployment, not after. Its purpose is to make the limitations explicit while
> the people who know them are still on the project — because in a year the model will still be
> running and they will not be here.
>
> Keep it in the repository next to the training code.

**Model:** _______________  **Version:** ____  **Date:** _______
**Owner (person):** _______________  **Status:** development / production / retired

## 1. What it does

**Prediction target, precisely — including edge cases:** _______________

**The decision made from this prediction:** _______________

**Who or what consumes the output:** _______________

**Is a human in the loop?** yes / no — where: _______________

## 2. Intended use

**In scope:** _______________

🔴 **Out of scope — uses this model must NOT be put to:** _______________

The out-of-scope line is the most valuable on the page. A model built to rank an internal queue
will eventually be suggested for customer messaging, and by then nobody remembers why that is a bad
idea.

## 3. Training data

| | |
|---|---|
| Source | |
| Period covered | |
| Rows / examples | |
| How labels were produced | |
| Label agreement between annotators | |
| Known gaps or under-represented groups | |
| Personal data included | yes / no — lawful basis |

**Seasonality or events NOT represented:** _______________
(A model trained on 14 months has seen one December. Say so.)

## 4. Method

| | |
|---|---|
| Algorithm / architecture | |
| Key hyperparameters | |
| **Split strategy** (time-based? grouped?) | |
| Test set used how many times | |
| Leakage review completed | yes / no — features removed: |
| Reproducible from clean checkout | yes / no |

## 5. Performance

| Metric | Value | On which set |
|---|---|---|
| **Trivial baseline** | | |
| **Existing process / rule** | | |
| This model | | |

**By segment** — an aggregate figure hides the segment where it fails:

| Segment | Volume share | Performance |
|---|---|---|
| | | |

**Decision threshold, and why that value:** _______________

## 6. Where it fails

**Weakest segments:** _______________

**Why:** _______________

**Does it fail visibly or silently?** visibly / silently

**What a wrong prediction costs, on the cases where it is weakest:** _______________

## 7. Fairness

| | |
|---|---|
| Groups checked for disparate performance | |
| Result | |
| Proxies considered (postcode, education, gaps) | |
| Not checked, and why | |

Recording "not checked, and why" is more useful than silence. Silence reads as "checked and fine".

## 8. Operations

| | |
|---|---|
| Where it runs | |
| Feature computation shared with training? | yes / no |
| Cost per prediction | |
| p95 latency | |
| **Outcomes captured, so production accuracy is measurable** | yes / no |
| Monitoring, and alert thresholds | |
| Who is alerted | |
| Retraining trigger | |
| Rollback to previous version tested | yes / no |

🔴 If outcomes are not captured, production accuracy cannot be measured and you are relying on the
evaluation score from launch day indefinitely.

## 9. Review

| | |
|---|---|
| Approved for production by | |
| Date | |
| Review due | |
| Conditions of approval | |

**What would cause us to retire this model:** _______________

## Change log

| Version | Date | What changed | Performance impact |
|---|---|---|---|
| | | | |
