A test matters only if it can stop the release
An evaluation deck has little value when every result leads to launch. Define critical failures, escalation owners and rollback triggers before the scores arrive, then rerun the decision after any meaningful system change.
The hypothetical 99-of-100 release review separates ordinary quality, critical disclosure and coverage, then carries the failure into monitoring and recovery.
Use the release table to join benchmark results, adversarial cases, live samples and incident evidence in one go-or-stop decision.
Keep in mind: Passing the listed tests does not prove safety beyond their coverage, especially after the model, prompt, tools, data or users change.
In this article6 sections
Editorial note: The release table is a planning example, not a safety certification. Passing a finite test set cannot establish performance outside its coverage or after the model, data, prompts, tools or users change.
Decide which failure can stop the release
Imagine a test run with 99 acceptable answers and one answer that reveals a restricted customer record. Calling it “99% accurate” hides the decision a release owner actually faces. The restricted disclosure is not interchangeable with a slightly awkward sentence.
The figures in this article are hypothetical. They illustrate how to structure a release review, not the results of a real product test. Define critical failures, acceptance rules and response owners before running the evaluation. Otherwise, a high average can become an excuse to explain away the one case that contradicts the system’s basic promise.
A release table with three separate conclusions
Keep ordinary quality, critical failures and coverage limits in different fields. They answer different questions. A small clean test set can establish that those examples passed, while leaving many user groups or operating conditions untested.
Scroll the table sideways to see every column.
| Evidence | Result in the example | Release implication |
|---|---|---|
| 100 routine test questions | 99 acceptable answers | Useful summary of this set, not the whole population. |
| Critical disclosure check | One restricted record exposed | Block this release under the predeclared rule. |
| Coverage review | No tests for scanned documents or French inputs | Do not claim support for those conditions. |
| Rollback rehearsal | Not completed | Recovery readiness remains unproved. |
Keep the denominator visible
“One failure” means little without the number and kind of opportunities. Record failures divided by evaluated cases, but also describe how the cases were chosen. One hundred easy questions copied from a demonstration cannot represent all the work users will bring.
Keep a held-out set that was not used to tune the prompt. Add cases from actual reported problems where lawful and appropriate, removing unnecessary personal details. Repeated runs can reveal variability, but repeated copies of one easy question do not replace broader coverage. Report which change each evaluation is meant to assess: model, prompt, tools, retrieval corpus or policy.
Connect the failed case to a live signal
A pre-release test is useful only if the failure remains visible after deployment. For the fictional disclosure case, define what evidence triggers escalation, who can disable the affected path and how access is contained. Avoid collecting sensitive prompts indiscriminately in the name of monitoring.
Track unsupported claims, access failures, unresolved outcomes and correction effort separately. Sample across the kinds of requests actually received. A sudden change in request mix can make an old score less relevant even when the model version is unchanged. Monitoring needs both an operating measure and a person who will act on it.
Continue with the original sources
These claim-relevant primary and first-party references support the reporting above. Open them for technical detail, current requirements and subsequent updates.
- nist.govNIST: Adversarial Machine Learning taxonomy ↗A vocabulary for attacks and mitigations. It is not a checklist that proves a deployed system secure.
- developers.openai.comOpenAI: Working with evals ↗First-party documentation on defining test criteria, datasets and evaluation runs; not independent evidence that a particular model performs well.
- nist.govNIST: Generative AI Risk Management Profile ↗Risk-management background, including confabulation and information integrity. It does not certify the examples or prescribe our scoring thresholds.
- genai.owasp.orgOWASP: Excessive Agency ↗Describes risks from excessive functionality, permissions and autonomy, and ways to limit them.
Finished reading? Save that here without waiting for a timer.
Rewritten throughout on September 21, 2026, with a complete worked method, explicit evidence limits and checked primary sources.
See something we should fix or clarify? Read the corrections policy or tell the newsroom. Material changes are noted here.
