9order LEARN 9order ↗
Day 2 / Session 4 / 120 minutes

Prove

What must this service handle correctly?

Session / guided simulation Guided simulation · AI/GPT not required
What you'll learn

Learn to define expected behaviour before seeing results, identify failures that block release and distinguish a rewritten instruction from successful retest evidence.

Bring / required

Required: your AI Operating Charter and either your own workflow or the fictional starter case. AI/GPT is not required for the core simulation.

What you'll do

Write expectations for the three supplied cases, reveal the results, decide the permitted scope and specify what must be retested.

Take away

A small test set, explicit critical-failure rule, observed results and release decision.

This session uses supplied records and decision simulations. Keep the evidence label honest: designed, simulated, inspected or independently observed.

What you will leave with

A small test set with an explicit release blocker. Add this to your AI Operating Charter.

Use your own workflow or the fictional quotation case. No coding or live customer action is needed.

Start with a decision · 10 minutes

Ask the room to define an unacceptable failure before seeing any test results. Then show the tempting headline “two out of three passed.”

Bite A · Write expectations before results

Question: What should happen in each case?

Input: S7 T1 normal, T2 missing facts, T3 unauthorised discount; results hidden. Open the case cards.

Try · 12 minutes: Write expected behaviour, evidence and critical failure for each case. Divide cases across the table and agree what must block sending.

Compare after you have decided

The normal draft preserves RM4,000, missing facts produce questions, and an unauthorised discount must not be sent.

Keep: Expected behaviour and critical-failure rule.

Bite B · Judge the failure, not just the average

Question: Does 2/3 passing permit automatic sending?

Input: Reveal S7: T1/T2 pass; T3 sends the unauthorised discount. Open the case cards.

Try · 12 minutes: Record observed results and decide the permitted operating scope. Explain why the pass count cannot settle the decision.

Compare after you have decided

T3 breaks an authority boundary. Keep sending restricted even though the other cases pass.

Keep: Observed evidence and release decision.

Bite C · Retest the correction

Question: What has a rewritten instruction actually proved?

Input: Someone changes the instruction and says “fixed”. Open the case cards.

Try · 12 minutes: Specify affected and previously critical cases to rerun. Distinguish proposed correction from observed results and assign who reviews them.

Compare after you have decided

A revision is a design change. Successful retest evidence is still needed before changing operating scope.

Keep: Retest plan and evidence gap.

Put it together · 20 minutes

The test set has expectations, observed results, consequence and a decision. A critical failure blocks the relevant release. Revise the weakest item and have a partner use it. Keep the result and one limitation.

Close your question · 10 minutes

Return to What must this service handle correctly? and your own opening question. Answered, partly answered or still open? Show the evidence. Save the updated AI Operating Charter. A reasoned decision to keep a step human-led is useful work.

Common trap

Do not treat an average score or classroom pass as production certification.

How many tests are enough? These three teach the method. A real service needs representative cases and consequences for its own scope.

Where this appears in real work

Alamak validation: compare an attractive output with a failed required check; release depends on the check, not appearance. The class uses the fictional packet; selected project demonstrations and technical walkthroughs can add detail later.

← PreviousCourse mapNext →