# Production Prompt Checklist

**Prompt / feature:** _______________  **Owner:** _______________  **Date:** _______

## 1. The task

- [ ] The prompt states one task, not three
- [ ] Success is defined — you can say what a correct output looks like
- [ ] The audience and register are stated where they matter
- [ ] Length constrained
- [ ] Examples included where the task shape is unusual

Examples do more work than instructions for anything with a specific output shape. Two good ones
usually beat a paragraph of description.

## 2. Untrusted content

- [ ] Retrieved or user-supplied content is **clearly delimited as data**
- [ ] The prompt states that content inside the delimiters is not an instruction
- [ ] No user content is concatenated directly into the instruction section
- [ ] Nothing the model reads can directly trigger a tool with real effect

🔴 Any text your system reads can contain instructions aimed at it. In a chatbot that produces a
bad answer; in a system with tools it produces an action. Delimiting is the minimum, not the
solution — see [AI agents](/knowledge/ai-agents/) for the architectural part.

## 3. Refusal and uncertainty

- [ ] There is an **explicit instruction on what to do when the answer is not available**
- [ ] The exact refusal wording is written down
- [ ] Refusal behaviour is tested on questions known to be unanswerable
- [ ] Over-refusal measured too

**The refusal text:** _______________

A system that always answers is not confident, it is unmeasured.

## 4. Output the calling code can rely on

- [ ] Output format specified precisely — schema, fields, types
- [ ] Structured output enforced by the API where available, not requested politely
- [ ] The parser handles a malformed response without crashing
- [ ] Behaviour defined for an empty or truncated response
- [ ] No secrets, internal identifiers or system details echoed back to users

## 5. Configuration

- [ ] **Model and version pinned** — not a floating "latest" alias
- [ ] Temperature and sampling settings deliberate, and recorded
- [ ] Token limit generous enough to complete the task
- [ ] Timeout and retry policy set
- [ ] Cost per call measured

Pinning matters more than it appears. A provider updating a model underneath you changes your
system's behaviour with no change on your side, and the symptom is a quality complaint nobody can
reproduce.

## 6. Version control and testing

- [ ] The prompt lives **in version control**, not in a database field or a console
- [ ] Changes go through review like code
- [ ] **A test set of at least 20 real cases exists**, with expected outcomes
- [ ] The suite is re-run before any prompt change ships
- [ ] Pass rate recorded per run, with the number of runs
- [ ] Every reported bad output is added to the test set permanently

🔴 A prompt edited directly in production, with no test set, is an untracked change to system
behaviour. It is the least controlled thing in most AI deployments.

## 7. Safety and compliance

- [ ] Disclosure that the user is interacting with AI, where required
- [ ] Generated content labelled where required
- [ ] Guidance on what the system must not discuss or do
- [ ] Personal data handling consistent with policy
- [ ] Logging excludes sensitive content, or is retained appropriately

## 8. Before shipping

- [ ] Run against the test set; result recorded
- [ ] Run several times on the same input — variation observed and acceptable
- [ ] Tried deliberately awkward inputs: empty, very long, wrong language, adversarial
- [ ] Cost and latency at expected volume checked
- [ ] Someone other than the author has read the prompt

## Sign-off

| | Name | Date |
|---|---|---|
| Author | | |
| Reviewer | | |
| Test set pass rate | | |
| Approved | | |
