# AI Benchmarking Checklist

**Benchmark:** _______________  **Version:** _______  **Date:** _______

> If you cannot say exactly what was held constant, you have a number, not a benchmark.

## 1. Pin everything — each of these changes the result

- [ ] **Model and version string** — "latest" is not a version
- [ ] Temperature and all sampling parameters
- [ ] Prompt version
- [ ] Retrieved context, if applicable
- [ ] Input set, fixed and ordered
- [ ] Number of runs per case
- [ ] Scorer model version and rubric text
- [ ] **Date** — providers update models under the same name

## 2. Input set

- [ ] Fixed, and **never changed silently**
- [ ] Adding cases creates a new version, stated in the report
- [ ] Composition recorded: typical / edge / should-refuse counts
- [ ] Drawn from real work, not invented

## 3. Execution

- [ ] Each case run multiple times (5 default)
- [ ] **Distribution reported, not just the mean**
- [ ] Order randomised if any state persists between runs
- [ ] Deterministic scoring used wherever possible
- [ ] Scorer pinned and versioned

## 4. Always include

- [ ] **Cost per completed task**
- [ ] p50 and p95 latency
- [ ] Refusal accuracy

> A benchmark measuring only quality will select an expensive, slow model for a task where a cheap
> one sufficed.

## 5. Fair comparison

- [ ] Same inputs, same scorer, same run count across candidates
- [ ] **Equal tuning effort** — otherwise you benchmarked your own effort
- [ ] Ties reported as ties, not resolved on a tenth of a point
- [ ] Started from the cheapest candidate

## 6. Report format

```
Benchmark:    ______________ v___
Date:         ______________
Cases:        ___ (___ typical, ___ edge, ___ should-refuse)
Runs:         ___ per case
Model:        <exact version string>
Temperature:  ___
Prompt:       v___ (commit ______)
Scorer:       ______________

Result:       structural ___%  ·  semantic ___% (σ ___)
              refusal accuracy ___%
              cost $____ per completed task
              p50 ___s / p95 ___s
```

- [ ] Report contains every field above
- [ ] Anyone reading it could reproduce the run

## 7. Reading published benchmarks

- [ ] **Contamination considered** — old benchmarks appear in training data
- [ ] Understood as peak, not typical, conditions
- [ ] Not compared across sources — harnesses differ
- [ ] Task mismatch assessed against your actual work
- [ ] Used to **rule models out**, never to rule one in

## 8. Re-run triggers

- [ ] Model version change
- [ ] Prompt change
- [ ] Scorer change
- [ ] Scheduled regardless

## Sign-off

| | Name | Date |
|---|---|---|
| Run by | | |
| Reviewed by | | |
