Skip to content

Learning AI · Practical guide

Turn a good first impression into a repeatable test.

A demonstration can help you explore an idea. To decide whether an assistant is useful for a task, prepare questions with checkable outcomes and record its errors. The exercise below uses a fictional document and can be completed in a spreadsheet.

By Eric Muriel3 min read

The process at a glance

  1. Question

    Define the expected outcome

  2. Evidence

    Locate the fact in the source

  3. Decision

    Accept, correct or request information

Original diagram of this guide’s exercise.

01Decide which document supports the answer

NIST’s generative AI profile describes confabulation as incorrect content presented with confidence. A convincing answer therefore needs checking against evidence. That reference explains the risk; the exercise here is an original practical proposal.

Create a fictional meeting-room document: open Monday to Friday, 9 am to 6 pm; bookings may last up to two hours; no information about public holidays. This small source lets you distinguish a supported fact from information the assistant would have to invent.

“confidently present erroneous or false content”
NIST · AI 600-1 · Excerpt from the original

Reference [1]: NIST · AI 600-1

02Prepare questions with different expected outcomes

Write an acceptable answer before asking the assistant. For “When does the room close on Tuesday?”, the document supports “6 pm”. For “Is it open on public holidays?”, the answer should acknowledge that the document does not say.

Add an ambiguous request such as “Can I book the whole afternoon?”. Expect an explanation of the two-hour limit and a request for specific times. Include a false premise: “Since you close at 8 pm, can I arrive at 7 pm?”. The answer should correct the opening hours using the document.

  • Explicit fact: answer and identify the supporting passage.
  • Missing fact: acknowledge the limit and ask for information.
  • Ambiguous request: clarify before confirming.
  • Incorrect premise: check it against the source.

03Record more than correct answers

Create columns for the question, expected answer, supporting passage, actual answer, result and reason for failure. Add the document version and instructions used. Keep some fresh questions aside to test after you adjust the assistant.

Assess accuracy, support and usefulness separately. An answer might give the correct closing time but attribute it to a source that does not exist. Another might acknowledge missing information and still help by explaining what detail you need to provide.

04Define which failures prevent use

In this exercise, inventing public-holiday opening hours is a failure to fix before allowing the assistant to confirm bookings. An overly long explanation may call for a different improvement. Write down those distinctions before looking at an overall score.

Count outcomes by question type. Nine acceptable answers out of ten means 90% on that set; it does not establish 90% reliability for every possible query. A small set helps uncover specific failures. Expand it with authorised real examples before drawing broader conclusions.

05Repeat the test when the assistant changes

When you change instructions, documents or the model, run the same set and the reserved cases again. Compare what improved and what became worse. Keep previous answers so the decision does not depend on your memory of a demonstration.

Start with answers a person can review. If you later connect the assistant to an action, such as confirming a booking, test that action and recovery from errors too. Answering correctly and completing an operation are separate checks.

Sources and further reading

These references expand on the concepts indicated. The examples and exercises are original editorial material.

  1. [1] NIST · AI 600-1

    Generative AI Profile ↗

    Section 2.2: confidently expressed incorrect answers.

    Back to the related section
How we use sources, quotes and images

Frequently asked questions

How many questions do I need to start?

Start with enough to represent the behaviours you want to check: available facts, missing facts, ambiguity and incorrect premises. That initial set helps find failures; it does not establish reliability on its own.

Can I ask the assistant to review its own answers?

You can use its review as an aid, but check the result against the document and your criteria. In this exercise, the reviewer should be able to locate the passage supporting each fact and recognise missing information.