CANADA / POLICYUnderstanding Canada’s AI transparency consultationIndependent Canadian publication
MODELS / EVALUATION

AI model evaluation: compare answers and check evidence

Model announcements compress many trade-offs into one launch score. A useful comparison separates capability, reliability, speed, cost, control and fit for the actual workload. Start with the answer-comparison scorecard, then work through release checks and document-retrieval failure cases.

Three questions to keep asking
  1. Which representative tasks does the model complete without rescue?
  2. What does an accepted outcome cost after retries and review?
  3. Which controls, deployment choices and evidence are available?
READING ROADMAP

How to use this guide.

01

Read benchmarks as clues

A benchmark shows performance on a defined test. It does not guarantee reliability on your data, workflow or risk level.

02

Test the complete system

Retrieval, prompts, tools, permissions and review steps often determine quality as much as the underlying model.

03

Measure change over time

Keep a stable evaluation set so model, prompt and data updates can be compared against the same acceptance criteria.

CURATED COVERAGE

Read the topic in a useful order.

3 selected guides
Illustration of a comparison checklist between two computer displays.

Compare AI answers: score evidence, not confidence

To compare AI answers, use the same task and sources, then score evidence, completeness, uncertainty, usefulness and permissions. A fluent answer must still fail if it invents authority to act. The 20-point worksheet below includes two fictional answers, source records and an answer key.

AI New Desk4 min readIntermediate how-to
Illustration of test icons, checklists and monitoring screens for evaluating AI.

The one failure that outweighs 99 passing AI checks

An AI release can pass 99 of 100 checks and still be unsafe to launch if the remaining failure breaks a critical boundary. Define blocking failures before testing, report results by failure type and plan a rollback. The 99-of-100 example here is hypothetical, not a measured product result.

AI New Desk3 min readSystem design guide
Illustration of indexed documents in a filing drawer with a search symbol.

Eight document traps for an AI retrieval system

Our synthetic selector passed its original eight document cases, then failed four of five additional probes involving bad or contradictory source records. Inspect both sets of inputs and outputs to see what version, audience and citation checks do—and what they leave untested.

AI New Desk4 min readSystem design guide