Checklist · Chaos Engineering

AI Chaos Engineering Checklist

Run a controlled failure experiment against an AI system, hypothesis, blast radius, stop conditions, the failures worth injecting, and what "graceful" has to mean.

Markdown. No sign-up, no email.

System: _______________ Experiment: _______________ Date: _______

Not "unplug things and see". A hypothesis test with a stop condition.

1. Design the experiment#

  • [ ] Steady state defined and measured first. What normal looks like, numerically
  • [ ] Hypothesis written before running it
  • [ ] Blast radius stated: environment, traffic, data
  • [ ] Stop condition stated. What ends the experiment immediately
  • [ ] Team present and available
  • [ ] Announced: an unannounced experiment that looks like an incident wastes everyone's time
  • [ ] Off-peak
Steady state:   ____________________________________
Hypothesis:     ____________________________________
Blast radius:   ____________________________________
Stop condition: ____________________________________

2. Provider failures to inject#

  • [ ] Complete outage. Sustained errors from the model API
  • [ ] Slow, not down. 30s instead of 2s. Harder and more common
  • [ ] Rate limited, sustained 429s
  • [ ] Malformed responses: truncated JSON, empty content, unexpected schema
  • [ ] Partial degradation. Responses arrive but are poor. No error to catch

3. Surrounding failures#

  • [ ] Retrieval unavailable: does it refuse, or answer ungrounded from training data?
  • [ ] Retrieval returns empty, not an error. Does it say "I don't know"?
  • [ ] Tool failures (agents), timeout, error, wrong shape
  • [ ] Stale corpus, nothing errors, answers confidently wrong
  • [ ] Spend cap reached mid-task

4. "Graceful" verified: define before testing#

  • [ ] Honest. User told something is wrong, not shown a spinner forever
  • [ ] Bounded. One slow dependency does not stall unrelated features
  • [ ] Lossless. Queued work survives, or is discarded with the user informed
  • [ ] Idempotent. A retry cannot duplicate an effect
  • [ ] Recoverable. Resumes without intervention when the dependency returns
  • [ ] Visible. Something alerted. Silent degradation lasts longest

5. The findings people usually get#

  • [ ] Timeouts longer than intended, threads piling up
  • [ ] Retry without backoff, extending the outage
  • [ ] No user-facing message at all
  • [ ] Non-idempotent operations retried, producing duplicates
  • [ ] Queued work lost on restart

6. Afterwards#

  • [ ] Result recorded against the hypothesis, matched or not
  • [ ] Surprises recorded. These are the value
  • [ ] Findings fed into runbooks and the incident plan
  • [ ] Fixes made, then experiment re-run to verify
  • [ ] Next experiment scheduled

Record#

Hypothesis held?Yes / No / Partly
What surprised us
Actions, with owners and dates
Re-run scheduled for

Sign-off#

NameDate
Experiment lead
System owner

Back to Chaos Engineering

Get new material when it is published

Everything here is free and stays free. There is no form in front of any document. If you want to know when new guides and templates go up, leave an email.

Roughly monthly. Unsubscribe in one click. We do not share your address, and we will not call you.