# RAG Implementation Checklist

**System:** _______________  **Date:** _______

## 1. Before writing any code

- [ ] **Who owns each document set, and is responsible for it being current?**
      If the answer is "nobody", fix that first — the most common RAG failure is a stale document
      retrieved perfectly
- [ ] Corpus scoped — start with one team's documents, not everything
- [ ] Twenty real questions written down, with known answers
- [ ] Decided what happens when the answer is not in the corpus
- [ ] Confirmed the knowledge is actually written down, not in people's heads
- [ ] Fine-tuning ruled out — facts belong in documents, not weights

## 2. Chunking

- [ ] **Split on structure** — sections, headings, paragraphs. Not fixed character counts
- [ ] Small overlap so a thought spanning a boundary survives
- [ ] **Document title and heading path prepended to every chunk**
- [ ] Chunks large enough to contain an answer, not just to match a query
- [ ] Metadata stored: owner, department, date, access level, document type

## 3. Retrieval

- [ ] **Hybrid search — keyword AND semantic.** Semantic alone fails on exact identifiers
- [ ] Metadata filtering available to narrow before ranking
- [ ] Embedding model and version pinned
- [ ] Number of chunks retrieved tuned — a cost lever and a quality lever at once
- [ ] Re-indexing process defined for when documents change

## 4. Generation

- [ ] Instruction to **answer only from supplied passages**
- [ ] **Explicit refusal instruction** — "if the passages do not contain the answer, say so"
- [ ] Supplied content clearly delimited as data, not instructions
- [ ] Output length constrained
- [ ] Model and version pinned

## 5. Sources

- [ ] **Every answer shows the chunks it used**
- [ ] Sources are clickable and lead to the real document
- [ ] Document date shown, so a reader can judge currency

> This is not cosmetic. It is the mechanism by which users catch mistakes, and the difference
> between a system people trust and one they quietly stop using.

## 6. Permissions — a security requirement

- [ ] Retrieval filtered by the **requesting user's** permissions
- [ ] Filtering happens at query time, not after generation
- [ ] System does not use a privileged service account
- [ ] Tested with users of different entitlements
- [ ] Failure treated as a security incident

## 7. Measurement — before launch

- [ ] **Retrieval measured separately from generation**
- [ ] Hit rate / recall@k recorded
- [ ] Groundedness measured
- [ ] Refusal accuracy measured on unanswerable questions
- [ ] Over-refusal measured
- [ ] Baseline recorded so later changes are comparable

> 🔴 If retrieval hit rate is 60%, your system's ceiling is 60%. No prompt work raises it.

## 8. Operations

- [ ] Re-index triggered when documents change
- [ ] Retrieval confidence logged
- [ ] Source click-through measured
- [ ] **Every reported bad answer added to the test set permanently**
- [ ] Document owners notified when their content produces a bad answer
- [ ] Scheduled re-test — the corpus changes even when the code does not

## 9. Cost

- [ ] Cost per answer measured
- [ ] Caching for repeated questions
- [ ] Chunk count tuned against quality, not maximised

## Sign-off

| | Name | Date |
|---|---|---|
| Built by | | |
| Content owner | | |
| Approved for users | | |
