Checklist · Hallucination Testing

Hallucination Testing Checklist

Measure fabrication in an AI system, grounding checks, the no-answer test most teams skip, identifier verification and the over-refusal counterweight. Free, editable Markdown.

About the numbers here. Numbers here are thresholds the checklist proposes, not measurements.

Markdown. No sign-up, no email.

System: _______________ Date: _______ Tested by: _______________

Do not test "is it true?": that requires someone who already knows the answer. Test "is it grounded?". Does every claim trace to something the system was actually given.

1. Setup#

  • [ ] Retrieved context captured alongside every output, grounding cannot be checked without it
  • [ ] Model and version pinned
  • [ ] Temperature recorded
  • [ ] Test set includes questions the corpus can and cannot answer

2. The core test: source attribution#

  • [ ] Output split into individual factual claims
  • [ ] Each claim checked against the supplied context
  • [ ] Unsupported claims counted → this is your fabrication rate
  • [ ] Judge model is different from the generator
  • [ ] Judge asked a narrow question ("is this claim supported: yes/no"), not a broad one
  • [ ] Judge calibrated against human review on a sample

3. The no-answer test: highest value, usually skipped#

  • [ ] Questions included whose answer is genuinely absent from the corpus
  • [ ] Correct behaviour defined: say so explicitly
  • [ ] Refusal rate measured
  • [ ] Refusal instruction present in the prompt ("if the passages do not contain the answer, say so")

A system scoring 95% on answerable questions and fabricating on 100% of unanswerable ones is dangerous precisely where the user is most dependent on it.

4. Identifier verification: cheapest, catches the worst errors#

  • [ ] Every number in the output checked against source
  • [ ] Every date checked
  • [ ] Every code, reference and identifier checked
  • [ ] Every proper noun checked
  • [ ] Percentages in prose checked against a source that actually contains them

5. Consistency#

  • [ ] Same question asked repeatedly (5+ runs)
  • [ ] Answers compared for contradiction
  • [ ] Inconsistency treated as strong evidence of fabrication
  • [ ] Consistency not treated as proof of correctness

6. Context manipulation#

  • [ ] Poisoned context. A deliberate falsehood in a document. It should be repeated (proves grounding)
  • [ ] Empty context. Nothing supplied. Anything returned is coming from training data
  • [ ] Irrelevant context. Passages that do not answer. Must refuse
  • [ ] Contradictory context. Two sources disagree. Must surface both

7. The counterweight#

  • [ ] Over-refusal measured. Refuses when the answer was present
  • [ ] Tracked alongside groundedness

Eliminating fabrication by refusing everything is not a fix. Teams that optimise only for groundedness drift there without noticing.

8. Where fabrication concentrates: verify each#

  • [ ] When retrieval returned nothing useful
  • [ ] Specific identifiers and citations
  • [ ] Numbers appearing in prose
  • [ ] Bridging between two passages
  • [ ] Attribution ("according to the policy…")
  • [ ] Long outputs. Fabrication rate rises with length

9. Metrics#

MetricValueTarget
Groundedness (claims traceable)primary
Refusal accuracyhigh
Identifier accuracy~100%
Consistency across runsdirectional
Over-refusal ratelow

10. Mitigations verified in place#

  • [ ] Explicit refusal instruction in the prompt
  • [ ] Sources shown in the interface. Makes remaining errors catchable
  • [ ] Retrieval quality addressed before prompt tuning
  • [ ] Output length constrained
  • [ ] Precise values extracted programmatically, not generated

Sign-off#

NameDate
Tested by
Reviewed by

Back to Hallucination Testing

Get new material when it is published

Everything here is free and stays free. There is no form in front of any document. If you want to know when new guides and templates go up, leave an email.

Roughly monthly. Unsubscribe in one click. We do not share your address, and we will not call you.