EDITOR’S NOTE

A test matters only if it can stop the release

An evaluation deck has little value when every result leads to launch. Define critical failures, escalation owners and rollback triggers before the scores arrive, then rerun the decision after any meaningful system change.

The hypothetical 99-of-100 release review separates ordinary quality, critical disclosure and coverage, then carries the failure into monitoring and recovery.

Use the release table to join benchmark results, adversarial cases, live samples and incident evidence in one go-or-stop decision.

Keep in mind: Passing the listed tests does not prove safety beyond their coverage, especially after the model, prompt, tools, data or users change.

4 named sources, checked belowRead the source notes ↓
In this article6 sections

Editorial note: The release table is a planning example, not a safety certification. Passing a finite test set cannot establish performance outside its coverage or after the model, data, prompts, tools or users change.

Decide which failure can stop the release

Imagine a test run with 99 acceptable answers and one answer that reveals a restricted customer record. Calling it “99% accurate” hides the decision a release owner actually faces. The restricted disclosure is not interchangeable with a slightly awkward sentence.

The figures in this article are hypothetical. They illustrate how to structure a release review, not the results of a real product test. Define critical failures, acceptance rules and response owners before running the evaluation. Otherwise, a high average can become an excuse to explain away the one case that contradicts the system’s basic promise.

A release table with three separate conclusions

Keep ordinary quality, critical failures and coverage limits in different fields. They answer different questions. A small clean test set can establish that those examples passed, while leaving many user groups or operating conditions untested.

Scroll the table sideways to see every column.

Illustrative release review for a support assistant
EvidenceResult in the exampleRelease implication
100 routine test questions99 acceptable answersUseful summary of this set, not the whole population.
Critical disclosure checkOne restricted record exposedBlock this release under the predeclared rule.
Coverage reviewNo tests for scanned documents or French inputsDo not claim support for those conditions.
Rollback rehearsalNot completedRecovery readiness remains unproved.

Keep the denominator visible

“One failure” means little without the number and kind of opportunities. Record failures divided by evaluated cases, but also describe how the cases were chosen. One hundred easy questions copied from a demonstration cannot represent all the work users will bring.

Keep a held-out set that was not used to tune the prompt. Add cases from actual reported problems where lawful and appropriate, removing unnecessary personal details. Repeated runs can reveal variability, but repeated copies of one easy question do not replace broader coverage. Report which change each evaluation is meant to assess: model, prompt, tools, retrieval corpus or policy.

Make adversarial cases correspond to real authority

A red-team case should test a failure path the system could actually take. Put a misleading instruction inside a retrieved document and see whether it changes tool authority. Ask for a record belonging to another user. Change a permission after an answer is cached. Use an invalid date or conflicting source to test whether the system admits uncertainty.

Run these checks in an authorized environment. Document the expected behaviour and the observed result, including partial failures. A refusal in the chat window does not prove that no restricted material appeared in logs or previews. Inspect the relevant boundary rather than judging only the final prose.

Connect the failed case to a live signal

A pre-release test is useful only if the failure remains visible after deployment. For the fictional disclosure case, define what evidence triggers escalation, who can disable the affected path and how access is contained. Avoid collecting sensitive prompts indiscriminately in the name of monitoring.

Track unsupported claims, access failures, unresolved outcomes and correction effort separately. Sample across the kinds of requests actually received. A sudden change in request mix can make an old score less relevant even when the model version is unchanged. Monitoring needs both an operating measure and a person who will act on it.

Rollback is a procedure, not a button label

Record which model, prompt, retrieval index, tools and configuration form the known working version. A model rollback cannot undo a permission change in a connected service or recover a message already sent. Identify those separate recovery tasks and test the feasible ones before release.

For the example, the decision is stop, investigate the restricted-record path, repair it and rerun relevant regression and access tests. The next review must still state what remains untested. NIST and OWASP provide risk and security guidance; they do not certify our illustrative thresholds or any system that copies this table.

EVIDENCE & FURTHER READING

Continue with the original sources

These claim-relevant primary and first-party references support the reporting above. Open them for technical detail, current requirements and subsequent updates.

Finished reading? Save that here without waiting for a timer.

Corrections & updates

Rewritten throughout on September 21, 2026, with a complete worked method, explicit evidence limits and checked primary sources.

See something we should fix or clarify? Read the corrections policy or tell the newsroom. Material changes are noted here.