A polished answer can still lose
Judge answers against the same records and predeclared stop conditions. A correct sum cannot rescue an invented authorization.
The complete Cedar Hall exercise, answer key and local worksheet let readers inspect the evidence, record scores and export a review.
Use the sheet when two tools both look good in a demo and you need to compare them on the work and mistakes that matter to you.
Keep in mind: The weights in this article are editorial examples. They are not a universal benchmark or proof that a system will behave the same way in production.
In this article6 sections
Editorial note: The 20-point scorecard and its weights are editorial examples for the fictional task shown. A real evaluation needs test cases, reviewers and failure limits chosen for its own users and consequences.
Fix the acceptance rule before reading the answers
A comparison becomes unfair when you invent the scoring rules after seeing which answer you prefer. Start with the task, the permitted sources and the mistakes that would make an answer unusable. Style can matter, but it should not rescue a response that changes a spending decision.
The Cedar Hall exercise below contains the complete source pack and two answers written for teaching. They are not outputs captured from commercial models. The worksheet lets you record your own judgment; the exercise does not establish a model ranking. Its purpose is to make disagreements specific enough that another reviewer can check them.
The task and its three records
Write a two-sentence public summary of the Cedar Hall equipment project. Include the approved spending limit and the unresolved decision. Cite the record numbers. A stop condition applies if an answer invents spending authorization or converts a proposed date into a binding one.
Scroll the table sideways to see every column.
| Record | Text |
|---|---|
| R1: committee note, June 3 | Approved up to $2,400 for two microphones and installation. A portable speaker was discussed but not approved. |
| R2: quote, June 4 | Microphones cost $1,600; installation costs $500; an optional speaker costs $700. All amounts include tax. The quote is not an order. |
| R3: coordinator email, June 5 | Hold the purchase until the room booking is confirmed. June 20 is a proposed event date; the venue has not confirmed it. |
Read both answers before scoring
Answer A: “Cedar Hall approved up to $2,400 for two microphones and installation; the quoted $2,100 leaves $300 within that limit (R1–R2). Purchasing remains on hold pending the room booking, and June 20 is only proposed (R3).”
Answer B: “Cedar Hall approved a $2,800 microphone, installation and speaker package for its confirmed June 20 event (R1–R2). The coordinator can proceed with purchasing because the quote establishes the final cost (R3).”
Both answers have citations. Check what those citations support before considering the tone. Mark each material claim as supported, contradicted or not established. Keep the purchase hold separate from the arithmetic; a correct addition cannot authorize an expense.
Five dimensions, with evidence beside each score
The worksheet uses five scores from zero to four, for a maximum of 20. Zero means unusable on that dimension, two means substantial correction is needed and four means the written criterion is met. Intermediate scores need an explanation. These weights are our illustrative choice, not an industry standard.
Scroll the table sideways to see every column.
| Dimension | What earns four points |
|---|---|
| Factual accuracy | All material claims match the records. |
| Coverage | The spending limit and unresolved decision are included. |
| Traceability | The cited records actually support the associated claims. |
| Instruction following | Two sentences, relevant scope and no invented authorization. |
| Practical usability | The summary can be used without substantive repair. |
Answer key: the expensive error is permission
Answer A’s arithmetic is $1,600 + $500 = $2,100, leaving $300 below the authorized ceiling. It keeps the purchase on hold and the date tentative. Awarding four in all five dimensions is reasonable for this exercise, although another reviewer might prefer a different formulation of the outstanding booking decision.
Answer B adds three quoted prices correctly but includes a speaker that was not approved. It also calls the date confirmed and contradicts the purchase hold. The predeclared stop condition rejects B. Giving it points for brevity or arithmetic must not turn it into an accepted result.
Now alter only R3 so that the room is confirmed but the purchase remains on hold for another reason. The room-status judgment should change; the purchase instruction should still fail. This variation checks whether reviewers can separate two conditions that happened to move together in the first case.
Turn this exercise into a comparison of your own
Use the worksheet below to preserve the same task, prompt, source text, settings and stop condition for both answers. Keep each actual output unchanged beside its score, system name, disclosed version and account tier. The exported JSON stays on your device; the worksheet does not send your source pack to a model. Avoid entering confidential material without authorization.
For a real tool comparison, keep inputs and allowed tools consistent and record the product, date and disclosed settings. Add ordinary, ambiguous, conflicting and unanswered cases. Reserve some examples from prompt development. Repeat cases when variability could change your decision, and report the limits of the sample instead of declaring a universal winner.
Finally, count review time and rejected attempts. If a fictional batch takes 60 minutes and produces eight accepted summaries, the effort is 7.5 minutes per accepted summary. A higher raw score can still be less useful if it requires more checking or produces a critical failure.
AI release evaluation and blocking failuresDecide which failures should prevent use even when the total score looks good.
Compare two answers yourself
Give both answers the same task and source material. Record your reasons before revealing product names. This tool adds your scores; it does not check facts or grade an AI model.
Nothing you enter is sent to us or saved in this page. Download your record before leaving. Avoid entering confidential material.
0 = unusable · 2 = substantial correction needed · 4 = meets the criterion. Use 1 and 3 for intermediate cases.
Keep incomplete records if useful, but do not report them as completed evaluations. There is no automatic passing score or universal winner.
Continue with the original sources
These claim-relevant primary and first-party references support the reporting above. Open them for technical detail, current requirements and subsequent updates.
- developers.openai.comOpenAI: Working with evals ↗First-party documentation on defining test criteria, datasets and evaluation runs; not independent evidence that a particular model performs well.
- nist.govNIST: Generative AI Risk Management Profile ↗Risk-management background, including confabulation and information integrity. It does not certify the examples or prescribe our scoring thresholds.
- canada.caTreasury Board: Guide on the use of generative AI ↗Federal workplace guidance on checking outputs and managing information. Its institutional requirements are not a universal rule for every Canadian business.
Finished reading? Save that here without waiting for a timer.
September 23: expanded the comparison worksheet to retain prompts, source text, settings and full outputs.
See something we should fix or clarify? Read the corrections policy or tell the newsroom. Material changes are noted here.
