Checklist · Performance Testing

AI Performance Testing Checklist

Load-test a system whose backend is a rate-limited external API, what to measure, the AI-specific failure modes, and why cost belongs in a performance test.

Markdown. No sign-up, no email.

System: _______________ Date: _______ Tested by: _______________

You do not control the bottleneck. The model runs on someone else's infrastructure, behind a rate limit, with latency you cannot tune. That changes what you measure and what you can fix.

1. Targets set before testing#

  • [ ] Target set on p95, not the average
  • [ ] Time-to-first-token target set (if streaming)
  • [ ] Total completion target set
  • [ ] Expected concurrency stated
  • [ ] Cost ceiling per request stated
  • [ ] Behaviour at capacity decided in advance: queue, degrade, or shed

2. Measure#

  • [ ] Time to first token
  • [ ] Total completion time
  • [ ] p50 / p95 / p99. Not the mean
  • [ ] Throughput at target concurrency
  • [ ] Rate-limit rejections. Your real ceiling
  • [ ] Cost per completed request, including retries
  • [ ] Tokens per request
  • [ ] Queue depth and wait time

3. Tests#

  • [ ] Baseline. Single request, repeated. Best case and its variance
  • [ ] Sustained load. Expected concurrency, sustained period
  • [ ] Spike. Sudden burst. Queue, shed or fail?
  • [ ] Long input. Near the context limit
  • [ ] Degraded backend. Provider slow or erroring. The failure you will actually get
  • [ ] Cost under load. Retries multiply spend invisibly

4. AI-specific failure modes checked#

  • [ ] Rate limits, not CPU, confirmed as the constraint
  • [ ] Retry amplification. A timeout triggers a retry and you pay for both
  • [ ] Context growth impact on latency and cost
  • [ ] Cold start (self-hosted models)
  • [ ] Streaming masking degradation. Time-to-first-token flat while total time worsens

5. Levers verified#

  • [ ] Caching in place for repeated or near-identical requests
  • [ ] Prompt length tracked as a metric
  • [ ] Number of retrieved chunks tuned
  • [ ] Cheaper model evaluated for the same task
  • [ ] Streaming used where the interface allows
  • [ ] Explicit queue with position indicator, not an unbounded wait
  • [ ] Deliberate degradation path defined

6. Targets by interaction type#

Set the target from what the interaction actually needs, then measure against it.

InteractionReasonable targetYours
Interactive chatFirst token < 1s, total < 5s
Form assist / autocompleteTotal < 2s
Document processingMinutes acceptable, progress required
Batch / overnightThroughput matters, latency does not

7. Agent-specific#

Complete if the system takes actions rather than answering.

  • [ ] End-to-end completion time measured, not per-call latency
  • [ ] Steps per task recorded. Step count multiplies everything downstream
  • [ ] Cost per completed task, not per call
  • [ ] Step count regression treated as a performance regression even when the task still succeeds
  • [ ] Behaviour measured when a tool is slow rather than failing

8. What you can actually change#

You do not control the model, so verify each lever is available:

  • [ ] Caching, the largest single lever at volume
  • [ ] Prompt length reduction
  • [ ] Fewer retrieved chunks
  • [ ] A cheaper model that still passes the evaluation set
  • [ ] Streaming, where the interface allows
  • [ ] Explicit queue with a position indicator
  • [ ] Deliberate degradation path rather than a timeout

9. Results#

MeasureTargetMeasuredPass
Time to first token (p95)
Total completion (p95)
Throughput at target concurrency
Rate-limit rejection rate
Cost per completed request
Steps per task (agents)

10. Findings#

#FindingImpactActionOwner

Sign-off#

NameDate
Tested by
Accepted by

Back to Performance Testing

Get new material when it is published

Everything here is free and stays free. There is no form in front of any document. If you want to know when new guides and templates go up, leave an email.

Roughly monthly. Unsubscribe in one click. We do not share your address, and we will not call you.