Lesson 11 · Intermediate · 15 min
Comparing two answers with a grid
Compare two Copilot Chat answers with a four-criteria grid: accuracy, relevance, format, sources. Method, example, exercise and quiz.
- Goal
- You will be able to obtain two answers from Copilot Chat, score them on four criteria starting with accuracy, and keep the better one for good reasons rather than for its fluency.
- Skills
- Check

Your first attempt, unaided
Send the same request to Copilot Chat twice, in two separate conversations. Pick the better of the two answers and write in one sentence why. Keep that sentence: the lesson will put it to the test.
Two answers are better than one when you know how to decide between them. The grid holds four criteria: accuracy, relevance, format, sources. You score each answer from 0 to 2 per criterion, starting with accuracy, then you keep the better one or merge the two. The most fluent answer is not the most accurate: the grid protects you against that reflex.
1Four criteria to tell two answers apart
Copilot Chat produces a different text every time, even with the same prompt. This variability is a resource. By comparing two outputs, you see what is stable, probably drawn from the source, and what varies, often filled in by the model. Work on self-consistency, a method where the model compares several of its own answers before deciding, shows a clear gain in accuracy; you apply the same idea by hand.
Three ways to get two answers. In a single prompt: "propose two numbered variants, then state what sets them apart". In two conversations: the same prompt, run twice. Or with two response modes: Copilot Chat offers Auto, Quick response and Think deeper in the selector at the top right. Quick response is fast; Think deeper takes the time to plan and check. The regenerate buttons described in 2026 for Copilot Chat are probably available without a licence, but Microsoft does not say so.
The grid is read in order. Accuracy: does every fact, figure or name appear in the source or behind an opened link? An answer that fails this criterion is eliminated, whatever the rest. Relevance: does the answer deal with the question asked, for the named recipient, with nothing off topic? Format: are the requested structure, numbered length and tone respected? Sources: are the quotations or the passages of the file indicated, and the gaps declared?
Criterion Question to ask 2 points if Accuracy Is every fact in the source? No unverifiable statement Relevance Does it answer the question, for this reader? Nothing off topic Format Structure, length, tone respected? All three instructions met Sources Passages cited, gaps declared? Every point is traceable Score from 0 to 2 per criterion, with one sentence of justification. The total is only a guide: an answer at 7 out of 8 with a wrong figure loses against an accurate answer at 5 out of 8. You can ask Copilot to fill in the grid on its two variants; this is useful for format, but accuracy remains for you to check, against the source.


Diagram "Compare two answers with a grid"Diagram generated by AI and reviewed 2Two committee summaries, only one is accurate
A project manager wants a summary of minutes for his steering committee.
Weak prompt:
Summarise these minutes in five bullets. Minutes: [pasted text]Summarised answer: a single version, fluent, with a bullet "decision: postpone work package 3 to the first quarter" whereas the minutes record a disagreement that was not settled. With no point of comparison, nothing raises the alarm.
Strong prompt:
Summarise the minutes below in five bullets for the steering committee. Propose two numbered variants. In each, one bullet = one decision or one action with owner and due date when they are given; if a topic was not settled, write "not settled". 80 words per variant. End with one sentence on what sets the two variants apart. Minutes: [pasted text]Summarised answer: variant 1 writes "work package 3: not settled, sponsor's decision awaited"; variant 2 adds a due date that the minutes do not contain. The grid gives 2 for accuracy to the first, 0 to the second: the choice is made.
What changes: two variants make the invention visible by contrast. The "not settled" instruction gives the model an honest output for doubt. The grid decides on accuracy before style.
3Score two versions of an announcement email
Take an email announcing an organisational change, written from fictional facts: two departments merged, no job cuts, a discussion meeting on Tuesday at 10 am. Here is the starting prompt, to be improved:
"Write this email in two versions."
Add the recipient, the numbered length, the structure (announcement, reason, next step, invitation to ask questions), the instruction "promise nothing that is not in the context" and the request for two distinct variants with what separates them. Run it, then fill in the grid for each variant and keep one. The composer below guides you on the Format, Length and Verification fields.
Self-assessment grid: (a) all four criteria are scored for each variant; (b) accuracy has been checked against your facts; (c) the choice is justified in one sentence.
4Twelve confident lines against one accurate line
An assistant asks for two variants of a reply to a customer who wants to know the withdrawal period set by the attached terms and conditions. Variant A runs twelve lines, cites three articles, explains the calculation and concludes without reservation on fourteen days. Variant B fits in three lines: "The attached terms provide ten days (article 7). The statutory fourteen-day period only applies to consumers; the customer's status is not specified." She keeps A, fuller and more self-assured. What should have been seen: length and confidence are not criteria of the grid. A cites articles the document does not contain; B is the only accurate variant, and it flags what is missing. Correction: score accuracy first, opening the document for every article cited, before reading the rest. A variant with one unverifiable claim scores 0 and drops out of the comparison, whatever the rest. Add to the prompt: "Cite an article only if you quote it word for word." Rule to remember: in the grid, one correct line beats twelve confident ones.
5Quiz
Three questions, instant feedback. Each option comes with an explanation.
6Proof of mastery
Paste the prompt, the two answers obtained, your completed grid (four criteria scored from 0 to 2 with one sentence of justification) and the version you kept.
This lesson counts towards the Intermediate badgeSee the four badges
Criteria
What you wrote at the start of the lesson
Going further
The grid prepares the hallucination hunt, which trains the eye on planted errors. The lesson on iteration shows how to correct the variant you kept rather than redoing it. The term hallucination is defined in the glossary.
Frequently asked questions
How do I get two answers?
Three ways: ask for two numbered variants in the same prompt; run the same prompt in two conversations; switch response mode between Quick response and Think deeper. The regenerate buttons described in 2026 for Copilot Chat are probably available without a licence, but Microsoft does not say so.
Do I always need all four criteria?
Accuracy is mandatory and is checked first. The other three are weighted according to the task: relevance matters more for a summary, format for a table, sources for a search.
Can I ask Copilot to score its own answers?
Yes, for format and relevance, where it is good at spotting an overrun in length or an off-topic passage. For accuracy, it cannot judge without a source; you check yourself against the document or the links.