# AI Regression Testing Checklist

**System:** _______________  **Change:** _______________  **Date:** _______

## 1. Suite structure

- [ ] **Golden set** — must always pass. Small, stable, high-consequence
- [ ] **Regression set** — every production failure, permanently
- [ ] **Exploratory set** — newer cases, allowed to fail while investigated
- [ ] The three are kept **separate**, not merged
- [ ] Promotion path defined: exploratory → golden once stable

## 2. Pinning — you cannot detect drift against a moving baseline

- [ ] **Model version pinned** — not "latest"
- [ ] Temperature and sampling fixed
- [ ] Prompt versioned in the repository
- [ ] Corpus snapshotted for the golden set
- [ ] Scorer version fixed

## 3. What triggers a run

- [ ] Prompt change
- [ ] **Model version change** — the most common cause of silent degradation
- [ ] Retrieval or chunking change
- [ ] Tool definition change
- [ ] **Document corpus change** — the system changed even though the code did not
- [ ] Scheduled run regardless of changes

## 4. What counts as a regression here

- [ ] **Correctness dropped** — a passing case now fails
- [ ] **Consistency dropped** — 10/10 became 7/10. Same nominal result, different system
- [ ] **Cost rose at flat quality** — prompt grew, score did not move
- [ ] **Refusal behaviour changed** — it now answers what it used to decline

## 5. Reading the result

- [ ] **Per-case deltas reported**, not only the aggregate
- [ ] Aggregate treated as a dashboard number, not a verdict
- [ ] Cases that broke identified individually
- [ ] Distinguished: correctness failure vs consistency failure
- [ ] Cause checked in order — prompt, model, corpus, tools
- [ ] Asked whether the old behaviour was actually correct

> 🔴 A suite reporting "84%, was 84%" can be hiding three broken cases offset by three improved ones.

## 6. Hygiene

- [ ] **Every production failure added the same day**, not "when we get time"
- [ ] Outputs archived, not just verdicts
- [ ] Test case updates are deliberate and noted, not silent
- [ ] Cases that have never failed and never will are pruned

## 7. Speed — a slow suite gets skipped exactly when it matters

- [ ] Golden set runs in minutes
- [ ] Full suite runs nightly
- [ ] Runs parallelised
- [ ] Results cached on unchanged prompt + model + input
- [ ] Fails fast on structural checks before paying a judge

## 8. CI

- [ ] Golden set blocks on failure
- [ ] Someone is accountable for a red build
- [ ] Alerting on consistency drop, not just pass/fail
- [ ] Alerting on cost per completed task rising

## Sign-off

| | Name | Date |
|---|---|---|
| Run by | | |
| Result | Pass / Fail / Accepted with exceptions | |
