EDITOR’S NOTE

A polished answer can still lose

Judge answers against the same records and predeclared stop conditions. A correct sum cannot rescue an invented authorization.

The complete Cedar Hall exercise, answer key and local worksheet let readers inspect the evidence, record scores and export a review.

Use the sheet when two tools both look good in a demo and you need to compare them on the work and mistakes that matter to you.

Keep in mind: The weights in this article are editorial examples. They are not a universal benchmark or proof that a system will behave the same way in production.

3 named sources, checked belowRead the source notes ↓
In this article6 sections

Editorial note: The 20-point scorecard and its weights are editorial examples for the fictional task shown. A real evaluation needs test cases, reviewers and failure limits chosen for its own users and consequences.

Fix the acceptance rule before reading the answers

A comparison becomes unfair when you invent the scoring rules after seeing which answer you prefer. Start with the task, the permitted sources and the mistakes that would make an answer unusable. Style can matter, but it should not rescue a response that changes a spending decision.

The Cedar Hall exercise below contains the complete source pack and two answers written for teaching. They are not outputs captured from commercial models. The worksheet lets you record your own judgment; the exercise does not establish a model ranking. Its purpose is to make disagreements specific enough that another reviewer can check them.

The task and its three records

Write a two-sentence public summary of the Cedar Hall equipment project. Include the approved spending limit and the unresolved decision. Cite the record numbers. A stop condition applies if an answer invents spending authorization or converts a proposed date into a binding one.

Scroll the table sideways to see every column.

Fictional Cedar Hall records: the complete source pack
RecordText
R1: committee note, June 3Approved up to $2,400 for two microphones and installation. A portable speaker was discussed but not approved.
R2: quote, June 4Microphones cost $1,600; installation costs $500; an optional speaker costs $700. All amounts include tax. The quote is not an order.
R3: coordinator email, June 5Hold the purchase until the room booking is confirmed. June 20 is a proposed event date; the venue has not confirmed it.

Read both answers before scoring

Answer A: “Cedar Hall approved up to $2,400 for two microphones and installation; the quoted $2,100 leaves $300 within that limit (R1–R2). Purchasing remains on hold pending the room booking, and June 20 is only proposed (R3).”

Answer B: “Cedar Hall approved a $2,800 microphone, installation and speaker package for its confirmed June 20 event (R1–R2). The coordinator can proceed with purchasing because the quote establishes the final cost (R3).”

Both answers have citations. Check what those citations support before considering the tone. Mark each material claim as supported, contradicted or not established. Keep the purchase hold separate from the arithmetic; a correct addition cannot authorize an expense.

Five dimensions, with evidence beside each score

The worksheet uses five scores from zero to four, for a maximum of 20. Zero means unusable on that dimension, two means substantial correction is needed and four means the written criterion is met. Intermediate scores need an explanation. These weights are our illustrative choice, not an industry standard.

Scroll the table sideways to see every column.

A rubric for this source-based task
DimensionWhat earns four points
Factual accuracyAll material claims match the records.
CoverageThe spending limit and unresolved decision are included.
TraceabilityThe cited records actually support the associated claims.
Instruction followingTwo sentences, relevant scope and no invented authorization.
Practical usabilityThe summary can be used without substantive repair.

Answer key: the expensive error is permission

Answer A’s arithmetic is $1,600 + $500 = $2,100, leaving $300 below the authorized ceiling. It keeps the purchase on hold and the date tentative. Awarding four in all five dimensions is reasonable for this exercise, although another reviewer might prefer a different formulation of the outstanding booking decision.

Answer B adds three quoted prices correctly but includes a speaker that was not approved. It also calls the date confirmed and contradicts the purchase hold. The predeclared stop condition rejects B. Giving it points for brevity or arithmetic must not turn it into an accepted result.

Now alter only R3 so that the room is confirmed but the purchase remains on hold for another reason. The room-status judgment should change; the purchase instruction should still fail. This variation checks whether reviewers can separate two conditions that happened to move together in the first case.

Turn this exercise into a comparison of your own

Use the worksheet below to preserve the same task, prompt, source text, settings and stop condition for both answers. Keep each actual output unchanged beside its score, system name, disclosed version and account tier. The exported JSON stays on your device; the worksheet does not send your source pack to a model. Avoid entering confidential material without authorization.

For a real tool comparison, keep inputs and allowed tools consistent and record the product, date and disclosed settings. Add ordinary, ambiguous, conflicting and unanswered cases. Reserve some examples from prompt development. Repeat cases when variability could change your decision, and report the limits of the sample instead of declaring a universal winner.

Finally, count review time and rejected attempts. If a fictional batch takes 60 minutes and produces eight accepted summaries, the effort is 7.5 minutes per accepted summary. A higher raw score can still be less useful if it requires more checking or produces a critical failure.

USE THE RUBRIC

Compare two answers yourself

Give both answers the same task and source material. Record your reasons before revealing product names. This tool adds your scores; it does not check facts or grade an AI model.

Nothing you enter is sent to us or saved in this page. Download your record before leaving. Avoid entering confidential material.

0 = unusable · 2 = substantial correction needed · 4 = meets the criterion. Use 1 and 3 for intermediate cases.

Answer A
Not fully scored

Incomplete: record the common task, prompt, sources, stop condition, actual output and evidence.

Answer B
Not fully scored

Incomplete: record the common task, prompt, sources, stop condition, actual output and evidence.

Keep incomplete records if useful, but do not report them as completed evaluations. There is no automatic passing score or universal winner.

EVIDENCE & FURTHER READING

Continue with the original sources

These claim-relevant primary and first-party references support the reporting above. Open them for technical detail, current requirements and subsequent updates.

Finished reading? Save that here without waiting for a timer.

Corrections & updates

September 23: expanded the comparison worksheet to retain prompts, source text, settings and full outputs.

See something we should fix or clarify? Read the corrections policy or tell the newsroom. Material changes are noted here.