Checklist · Model Evaluation

Model Evaluation Checklist

Select a model with evidence rather than a leaderboard, building the evaluation set, pinning conditions, comparing fairly on cost as well as quality. Free, editable Markdown.

Markdown. No sign-up, no email.

Task: _______________ Date: _______ Evaluated by: _______________

The question is not which model is best. It is what is the cheapest model that passes my evaluation set: because pricing spans a thousand-fold range and quality on well-specified business tasks does not.

1. Define the task first#

  • [ ] The task is stated specifically enough to be checked
  • [ ] "Correct" is defined as properties, not a single expected string
  • [ ] Consequence of a wrong answer understood, this sets the required bar
  • [ ] Latency requirement stated
  • [ ] Volume estimate stated

2. Build the evaluation set#

  • [ ] 20 to 50 cases drawn from real work
  • [ ] Typical cases
  • [ ] Edge cases: empty, malformed, unusually long, wrong language
  • [ ] Every known past failure, permanently
  • [ ] Should-refuse cases
  • [ ] Adversarial cases, injection, contradictory instructions
  • [ ] Expected outcome recorded for each

3. Pin the conditions#

  • [ ] Model version string recorded, not "latest"
  • [ ] Temperature and sampling parameters fixed
  • [ ] Prompt version fixed and identical across candidates
  • [ ] Retrieved context fixed, if applicable
  • [ ] Number of runs per case fixed (5 is a reasonable default)
  • [ ] Scorer version fixed
  • [ ] Date recorded. Providers update models under the same name

4. Score#

  • [ ] Deterministic checks: parses, fields present, arithmetic, no personal data
  • [ ] Rubric for semantic quality, max 4 dimensions
  • [ ] Judge is a different model from any candidate
  • [ ] Pass rate reported per case, not a single aggregate
  • [ ] Variance reported, not just the mean

5. Compare fairly#

  • [ ] Same inputs, same scorer, same run count for every candidate
  • [ ] Equal tuning effort. Do not polish one prompt and not another
  • [ ] Started from the cheapest candidate, moving up only on failure
  • [ ] Ties reported as ties, not resolved on a tenth of a point
Model AModel BModel C
Deterministic pass rate
Rubric score
Pass rate at production temperature
p50 / p95 latency
Cost per completed task
Refusal accuracy
Injection resistance

6. The criteria not on any leaderboard#

  • [ ] Data handling terms. Is your input used for training?
  • [ ] Regional data residency
  • [ ] Rate limits at your actual volume
  • [ ] Deprecation policy and notice period
  • [ ] Context window, if genuinely needed
  • [ ] Provider stability

7. Decide and record#

  • [ ] Chosen model recorded, with the reason
  • [ ] Rejected candidates recorded, with why, stops the debate recurring
  • [ ] Model version pinned in configuration
  • [ ] Evaluation set committed alongside the prompt

8. Re-evaluation triggers#

  • [ ] Model version change. A full run before shipping
  • [ ] Prompt change
  • [ ] Data distribution shift
  • [ ] New failure mode observed in production
  • [ ] Scheduled regardless, quarterly floor

Sign-off#

NameDate
Evaluated by
Decision approved by

Back to Model Evaluation

Get new material when it is published

Everything here is free and stays free. There is no form in front of any document. If you want to know when new guides and templates go up, leave an email.

Roughly monthly. Unsubscribe in one click. We do not share your address, and we will not call you.