CANADA / POLICYUnderstanding Canada’s AI transparency consultationIndependent Canadian publication
DESK / 02
MODELS DESK

Models AI news, guides and analysis.

Practical methods for comparing AI answers and planning evaluations, with limits and failure cases.

Latest models stories

2 articles
Illustration of a comparison checklist between two computer displays.

Compare AI answers: score evidence, not confidence

To compare AI answers, use the same task and sources, then score evidence, completeness, uncertainty, usefulness and permissions. A fluent answer must still fail if it invents authority to act. The 20-point worksheet below includes two fictional answers, source records and an answer key.

AI New Desk4 min readIntermediate how-to
Illustration of test icons, checklists and monitoring screens for evaluating AI.

The one failure that outweighs 99 passing AI checks

An AI release can pass 99 of 100 checks and still be unsafe to launch if the remaining failure breaks a critical boundary. Define blocking failures before testing, report results by failure type and plan a rollback. The 99-of-100 example here is hypothetical, not a measured product result.

AI New Desk3 min readSystem design guide
THE AI NEW LEARNING LAB

Put what you read into practice.

Follow guided tracks, save a focused reading path and test what you understood. Progress stays in this browser.

FREE LEARNING LABBuild a source-led reading path.

Choose a focused track, save useful articles and check what you understood with practical questions and flashcards.

Open the Learning Lab →No signup. Learning progress stays in this browser.