Model Card Template
A one-page record of what a model is, what it was trained on, where it works and where it does not, filled in before deployment so the limitations are stated by the people who know them.
Markdown. No sign-up, no email.
Fill this in before deployment, not after. Its purpose is to make the limitations explicit while the people who know them are still on the project, because in a year the model will still be running and they will not be here.
Keep it in the repository next to the training code.
Model: _______________ Version: ____ Date: _______ Owner (person): _______________ Status: development / production / retired
1. What it does#
Prediction target, precisely, including edge cases: _______________
The decision made from this prediction: _______________
Who or what consumes the output: _______________
Is a human in the loop? yes / no, where: _______________
2. Intended use#
In scope: _______________
Out of scope: uses this model must NOT be put to: _______________
The out-of-scope line is the most valuable on the page. A model built to rank an internal queue will eventually be suggested for customer messaging, and by then nobody remembers why that is a bad idea.
3. Training data#
| Source | |
| Period covered | |
| Rows / examples | |
| How labels were produced | |
| Label agreement between annotators | |
| Known gaps or under-represented groups | |
| Personal data included | yes / no, lawful basis |
Seasonality or events NOT represented: _______________ (A model trained on 14 months has seen one December. Say so.)
4. Method#
| Algorithm / architecture | |
| Key hyperparameters | |
| Split strategy (time-based? grouped?) | |
| Test set used how many times | |
| Leakage review completed | yes / no, features removed: |
| Reproducible from clean checkout | yes / no |
5. Performance#
| Metric | Value | On which set |
|---|---|---|
| Trivial baseline | ||
| Existing process / rule | ||
| This model |
By segment. An aggregate figure hides the segment where it fails:
| Segment | Volume share | Performance |
|---|---|---|
Decision threshold, and why that value: _______________
6. Where it fails#
Weakest segments: _______________
Why: _______________
Does it fail visibly or silently? visibly / silently
What a wrong prediction costs, on the cases where it is weakest: _______________
7. Fairness#
| Groups checked for disparate performance | |
| Result | |
| Proxies considered (postcode, education, gaps) | |
| Not checked, and why |
Recording "not checked, and why" is more useful than silence. Silence reads as "checked and fine".
8. Operations#
| Where it runs | |
| Feature computation shared with training? | yes / no |
| Cost per prediction | |
| p95 latency | |
| Outcomes captured, so production accuracy is measurable | yes / no |
| Monitoring, and alert thresholds | |
| Who is alerted | |
| Retraining trigger | |
| Rollback to previous version tested | yes / no |
If outcomes are not captured, production accuracy cannot be measured and you are relying on the evaluation score from launch day indefinitely.
9. Review#
| Approved for production by | |
| Date | |
| Review due | |
| Conditions of approval |
What would cause us to retire this model: _______________
Change log#
| Version | Date | What changed | Performance impact |
|---|---|---|---|