Checklist · Benchmarking

AI Benchmarking Checklist

Build a benchmark whose results actually compare, what must be pinned, how to avoid the traps in published benchmarks, and the reporting format that makes a result trustworthy.

Markdown. No sign-up, no email.

Benchmark: _______________ Version: _______ Date: _______

If you cannot say exactly what was held constant, you have a number, not a benchmark.

1. Pin everything: each of these changes the result#

  • [ ] Model and version string. "latest" is not a version
  • [ ] Temperature and all sampling parameters
  • [ ] Prompt version
  • [ ] Retrieved context, if applicable
  • [ ] Input set, fixed and ordered
  • [ ] Number of runs per case
  • [ ] Scorer model version and rubric text
  • [ ] Date. Providers update models under the same name

2. Input set#

  • [ ] Fixed, and never changed silently
  • [ ] Adding cases creates a new version, stated in the report
  • [ ] Composition recorded: typical / edge / should-refuse counts
  • [ ] Drawn from real work, not invented

3. Execution#

  • [ ] Each case run multiple times (5 default)
  • [ ] Distribution reported, not just the mean
  • [ ] Order randomised if any state persists between runs
  • [ ] Deterministic scoring used wherever possible
  • [ ] Scorer pinned and versioned

4. Always include#

  • [ ] Cost per completed task
  • [ ] p50 and p95 latency
  • [ ] Refusal accuracy

A benchmark measuring only quality will select an expensive, slow model for a task where a cheap one sufficed.

5. Fair comparison#

  • [ ] Same inputs, same scorer, same run count across candidates
  • [ ] Equal tuning effort. Otherwise you benchmarked your own effort
  • [ ] Ties reported as ties, not resolved on a tenth of a point
  • [ ] Started from the cheapest candidate

6. Report format#

Benchmark:    ______________ v___
Date:         ______________
Cases:        ___ (___ typical, ___ edge, ___ should-refuse)
Runs:         ___ per case
Model:        <exact version string>
Temperature:  ___
Prompt:       v___ (commit ______)
Scorer:       ______________

Result:       structural ___%  ·  semantic ___% (σ ___)
              refusal accuracy ___%
              cost $____ per completed task
              p50 ___s / p95 ___s
  • [ ] Report contains every field above
  • [ ] Anyone reading it could reproduce the run

7. Reading published benchmarks#

  • [ ] Contamination considered. Old benchmarks appear in training data
  • [ ] Understood as peak, not typical, conditions
  • [ ] Not compared across sources, harnesses differ
  • [ ] Task mismatch assessed against your actual work
  • [ ] Used to rule models out, never to rule one in

8. Re-run triggers#

  • [ ] Model version change
  • [ ] Prompt change
  • [ ] Scorer change
  • [ ] Scheduled regardless

Sign-off#

NameDate
Run by
Reviewed by

Back to Benchmarking

Get new material when it is published

Everything here is free and stays free. There is no form in front of any document. If you want to know when new guides and templates go up, leave an email.

Roughly monthly. Unsubscribe in one click. We do not share your address, and we will not call you.