Checklist · Chaos Engineering
AI Chaos Engineering Checklist
Run a controlled failure experiment against an AI system, hypothesis, blast radius, stop conditions, the failures worth injecting, and what "graceful" has to mean.
Markdown. No sign-up, no email.
System: _______________ Experiment: _______________ Date: _______
Not "unplug things and see". A hypothesis test with a stop condition.
1. Design the experiment#
- [ ] Steady state defined and measured first. What normal looks like, numerically
- [ ] Hypothesis written before running it
- [ ] Blast radius stated: environment, traffic, data
- [ ] Stop condition stated. What ends the experiment immediately
- [ ] Team present and available
- [ ] Announced: an unannounced experiment that looks like an incident wastes everyone's time
- [ ] Off-peak
Steady state: ____________________________________
Hypothesis: ____________________________________
Blast radius: ____________________________________
Stop condition: ____________________________________
2. Provider failures to inject#
- [ ] Complete outage. Sustained errors from the model API
- [ ] Slow, not down. 30s instead of 2s. Harder and more common
- [ ] Rate limited, sustained 429s
- [ ] Malformed responses: truncated JSON, empty content, unexpected schema
- [ ] Partial degradation. Responses arrive but are poor. No error to catch
3. Surrounding failures#
- [ ] Retrieval unavailable: does it refuse, or answer ungrounded from training data?
- [ ] Retrieval returns empty, not an error. Does it say "I don't know"?
- [ ] Tool failures (agents), timeout, error, wrong shape
- [ ] Stale corpus, nothing errors, answers confidently wrong
- [ ] Spend cap reached mid-task
4. "Graceful" verified: define before testing#
- [ ] Honest. User told something is wrong, not shown a spinner forever
- [ ] Bounded. One slow dependency does not stall unrelated features
- [ ] Lossless. Queued work survives, or is discarded with the user informed
- [ ] Idempotent. A retry cannot duplicate an effect
- [ ] Recoverable. Resumes without intervention when the dependency returns
- [ ] Visible. Something alerted. Silent degradation lasts longest
5. The findings people usually get#
- [ ] Timeouts longer than intended, threads piling up
- [ ] Retry without backoff, extending the outage
- [ ] No user-facing message at all
- [ ] Non-idempotent operations retried, producing duplicates
- [ ] Queued work lost on restart
6. Afterwards#
- [ ] Result recorded against the hypothesis, matched or not
- [ ] Surprises recorded. These are the value
- [ ] Findings fed into runbooks and the incident plan
- [ ] Fixes made, then experiment re-run to verify
- [ ] Next experiment scheduled
Record#
| Hypothesis held? | Yes / No / Partly |
| What surprised us | |
| Actions, with owners and dates | |
| Re-run scheduled for |
Sign-off#
| Name | Date | |
|---|---|---|
| Experiment lead | ||
| System owner |