# AI Chaos Engineering Checklist

**System:** _______________  **Experiment:** _______________  **Date:** _______

> Not "unplug things and see". A hypothesis test with a stop condition.

## 1. Design the experiment

- [ ] **Steady state defined and measured first** — what normal looks like, numerically
- [ ] **Hypothesis written before running it**
- [ ] Blast radius stated: environment, traffic, data
- [ ] **Stop condition stated** — what ends the experiment immediately
- [ ] Team present and available
- [ ] Announced — an unannounced experiment that looks like an incident wastes everyone's time
- [ ] Off-peak

```
Steady state:   ____________________________________
Hypothesis:     ____________________________________
Blast radius:   ____________________________________
Stop condition: ____________________________________
```

## 2. Provider failures to inject

- [ ] **Complete outage** — sustained errors from the model API
- [ ] **Slow, not down** — 30s instead of 2s. Harder and more common
- [ ] Rate limited — sustained 429s
- [ ] Malformed responses — truncated JSON, empty content, unexpected schema
- [ ] **Partial degradation** — responses arrive but are poor. No error to catch

## 3. Surrounding failures

- [ ] Retrieval unavailable — does it refuse, or answer ungrounded from training data?
- [ ] Retrieval returns empty — not an error. Does it say "I don't know"?
- [ ] Tool failures (agents) — timeout, error, wrong shape
- [ ] Stale corpus — nothing errors, answers confidently wrong
- [ ] Spend cap reached mid-task

## 4. "Graceful" verified — define before testing

- [ ] **Honest** — user told something is wrong, not shown a spinner forever
- [ ] **Bounded** — one slow dependency does not stall unrelated features
- [ ] **Lossless** — queued work survives, or is discarded with the user informed
- [ ] **Idempotent** — a retry cannot duplicate an effect
- [ ] **Recoverable** — resumes without intervention when the dependency returns
- [ ] **Visible** — something alerted. Silent degradation lasts longest

## 5. The findings people usually get

- [ ] Timeouts longer than intended, threads piling up
- [ ] Retry without backoff, extending the outage
- [ ] No user-facing message at all
- [ ] Non-idempotent operations retried, producing duplicates
- [ ] Queued work lost on restart

## 6. Afterwards

- [ ] Result recorded against the hypothesis — matched or not
- [ ] Surprises recorded. These are the value
- [ ] Findings fed into runbooks and the incident plan
- [ ] Fixes made, then **experiment re-run to verify**
- [ ] Next experiment scheduled

## Record

| | |
|---|---|
| Hypothesis held? | Yes / No / Partly |
| What surprised us | |
| Actions, with owners and dates | |
| Re-run scheduled for | |

## Sign-off

| | Name | Date |
|---|---|---|
| Experiment lead | | |
| System owner | | |
