AI evaluation
How to evaluate an AI product
The short answer
Evaluate the whole product against the real task. Define what good looks like, build representative and adversarial cases, measure the failures that matter, set thresholds, and keep the evaluation suite as a regression test.
Start with the decision the product must support
Generic model benchmarks say little about whether a product is safe or useful in its operating context. Begin with the target user task, the acceptable outcome, the consequential failure modes and what happens after the model responds.
Build a representative evaluation set
Include common cases, difficult cases, ambiguous inputs, missing information, edge cases and examples that previously caused failure. The set should mirror the inputs the product sees in real use, hard cases included.
Measure more than answer quality
- Task success and failure severity.
- Unsupported or fabricated claims.
- Tool selection and action accuracy.
- Calibration and handling of uncertainty.
- Latency, cost and retry behaviour.
- Human review burden and override patterns.
Choose thresholds before launch
A metric becomes useful once the team agrees what level is acceptable. Set thresholds by consequence: an unsafe action deserves far more weight than a cosmetic wording slip.
Keep evaluation in the delivery loop
Model, prompt, retrieval and tool changes can all create regressions. Run the evaluation suite on every change as the product evolves, and add new failure cases as you discover them.

