RAG Testing Checklist
Verify a retrieval-augmented system before it reaches users, retrieval quality, grounding, permissions and the failure cases most teams skip. Free, editable Markdown.
About the numbers here. Numbers here are thresholds the checklist proposes, not measurements.
Markdown. No sign-up, no email.
System: _______________ Date: _______ Completed by: _______________
Test retrieval and generation separately. If retrieval is the constraint, work on the generator is wasted effort.
1. Before testing#
- [ ] Document corpus identified, with a named owner per document set
- [ ] Someone is accountable for each document being current
- [ ] Chunking strategy recorded, split on structure, not character count
- [ ] Heading path / document title prepended to each chunk
- [ ] Metadata stored for filtering: owner, department, date, access level
- [ ] Embedding model and version pinned
- [ ] Test question set built from real user questions
2. Retrieval: test this first#
- [ ] Question set paired with the passage that should answer each
- [ ] Hit rate / recall@k measured
- [ ] Precision@k measured, how much returned content was relevant
- [ ] Mean reciprocal rank measured, position matters
- [ ] Coverage measured, questions with no relevant passage returned
- [ ] Hybrid search in place (keyword + semantic)
- [ ] Exact identifiers tested, invoice numbers, case references, codes
- [ ] Multi-hop questions identified and their behaviour understood
- [ ] Chunk size validated: chunks contain enough to answer, not just to match
If hit rate is 60%, your system's accuracy ceiling is 60%. Fix this before prompting.
3. Generation: with correct passages supplied#
- [ ] Groundedness measured: every claim traceable to supplied context
- [ ] Answer relevance measured: it addresses the question actually asked
- [ ] Completeness measured: uses all relevant supplied information
- [ ] Refusal accuracy measured. Given passages that do not answer, it says so
- [ ] Over-refusal measured: it does not refuse when the answer was present
- [ ] Sources returned with every answer
- [ ] Answer length constrained
4. The failure cases teams skip#
- [ ] Unanswerable questions. Must refuse, not fabricate
- [ ] Ambiguous questions. Two documents disagree; must surface both
- [ ] Outdated documents. Old and current versions both present
- [ ] Empty context. Retrieval returns nothing. Must not answer from training data
- [ ] Poisoned context. A deliberate falsehood in a document. Must repeat it (proves grounding)
- [ ] Injected instructions. A document containing "ignore the above"
- [ ] Very long context: near the limit, where behaviour degrades quietly
5. Permissions: this is a security test#
- [ ] Retrieval filtered by the requesting user's permissions, not a service account
- [ ] Same question run as users with different entitlements
- [ ] Each sees only what they are entitled to
- [ ] Failure treated as a security incident, not a quality bug
- [ ] Permission logic tested after any change to retrieval
6. Operations#
- [ ] Retrieval confidence logged in production
- [ ] Low-confidence answers identifiable after the fact
- [ ] Source click-through measured
- [ ] Every reported bad answer added to the offline suite permanently
- [ ] Suite re-run on prompt change, model change, embedding change and corpus change
- [ ] Scheduled run regardless of changes, the corpus moves on its own
7. Cost#
- [ ] Number of chunks retrieved tuned, more context is not reliably better
- [ ] Cost per answer measured
- [ ] Caching in place for repeated questions
Sign-off#
| Name | Date | |
|---|---|---|
| Completed by | ||
| Reviewed by |
Known limitations released with:
| # | Limitation | Impact | Accepted by |
|---|---|---|---|