# RAG Testing Checklist

**System:** _______________  **Date:** _______  **Completed by:** _______________

> Test retrieval and generation **separately**. If retrieval is the constraint, work on the
> generator is wasted effort.

## 1. Before testing

- [ ] Document corpus identified, with a **named owner per document set**
- [ ] Someone is accountable for each document being current
- [ ] Chunking strategy recorded — split on structure, not character count
- [ ] Heading path / document title prepended to each chunk
- [ ] Metadata stored for filtering: owner, department, date, access level
- [ ] Embedding model and version pinned
- [ ] Test question set built from **real user questions**

## 2. Retrieval — test this first

- [ ] Question set paired with the passage that should answer each
- [ ] **Hit rate / recall@k** measured
- [ ] Precision@k measured — how much returned content was relevant
- [ ] Mean reciprocal rank measured — position matters
- [ ] Coverage measured — questions with *no* relevant passage returned
- [ ] **Hybrid search in place** (keyword + semantic)
- [ ] Exact identifiers tested — invoice numbers, case references, codes
- [ ] Multi-hop questions identified and their behaviour understood
- [ ] Chunk size validated: chunks contain enough to answer, not just to match

> 🔴 **If hit rate is 60%, your system's accuracy ceiling is 60%.** Fix this before prompting.

## 3. Generation — with correct passages supplied

- [ ] **Groundedness** measured: every claim traceable to supplied context
- [ ] Answer relevance measured: it addresses the question actually asked
- [ ] Completeness measured: uses all relevant supplied information
- [ ] **Refusal accuracy measured** — given passages that do not answer, it says so
- [ ] Over-refusal measured — it does not refuse when the answer *was* present
- [ ] Sources returned with every answer
- [ ] Answer length constrained

## 4. The failure cases teams skip

- [ ] **Unanswerable questions** — must refuse, not fabricate
- [ ] **Ambiguous questions** — two documents disagree; must surface both
- [ ] **Outdated documents** — old and current versions both present
- [ ] **Empty context** — retrieval returns nothing. Must not answer from training data
- [ ] **Poisoned context** — a deliberate falsehood in a document. Must repeat it (proves grounding)
- [ ] **Injected instructions** — a document containing "ignore the above"
- [ ] Very long context — near the limit, where behaviour degrades quietly

## 5. Permissions — this is a security test

- [ ] Retrieval filtered by the **requesting user's** permissions, not a service account
- [ ] Same question run as users with different entitlements
- [ ] Each sees only what they are entitled to
- [ ] Failure treated as a **security incident**, not a quality bug
- [ ] Permission logic tested after any change to retrieval

## 6. Operations

- [ ] Retrieval confidence logged in production
- [ ] Low-confidence answers identifiable after the fact
- [ ] Source click-through measured
- [ ] **Every reported bad answer added to the offline suite permanently**
- [ ] Suite re-run on prompt change, model change, embedding change and corpus change
- [ ] Scheduled run regardless of changes — the corpus moves on its own

## 7. Cost

- [ ] Number of chunks retrieved tuned — more context is not reliably better
- [ ] Cost per answer measured
- [ ] Caching in place for repeated questions

## Sign-off

| | Name | Date |
|---|---|---|
| Completed by | | |
| Reviewed by | | |

**Known limitations released with:**

| # | Limitation | Impact | Accepted by |
|---|---|---|---|
| | | | |
