Skip to content
QDNALearn AI, from beginner to expert
FR

Lesson 19 · Advanced · 20 min

Writing an evaluation grid and having Copilot apply it

Write an evaluation grid of observable criteria, have Copilot Chat apply it in a separate conversation, then correct and measure the gap.

Goal
You will be able to write a grid of five observable criteria, have Copilot Chat apply it to a deliverable in a separate conversation, apply the corrections proposed and measure the progress between two versions.
Skills
Check
A briefing note placed next to a grid of five criteria scored from 0 to 2, with a second Copilot Chat window playing the reviewer.
Illustration generated by AI

Your first attempt, unaided

Take a deliverable you produced with Copilot Chat and ask it, in a new conversation, to mark it out of ten. Look at the mark, then ask yourself what exactly it measures.

In brief.

Asking Copilot "is it good?" gets you a compliment. Giving it a grid of five observable criteria, in a separate conversation, gets you a verdict per criterion, a justification and applicable corrections. The grid is written before the production prompt, as a success criterion, and replayed after correction to measure progress. It does not replace your proofreading: it makes it faster and more consistent.

  1. 1Why a criterion must be observed, not judged

    Anthropic's prompt engineering guide opens with a piece of advice most users skip: define a success criterion and a way to test it before working on the prompt. The evaluation grid is that way. It lists five observable criteria, each scored from 0 to 2 with a one-sentence justification, and it applies to any version of the deliverable, produced by you or by Copilot. A good criterion can be observed: "the target audience is named in the first sentence", "every figure refers to a source", "no sentence over 25 words", "the recommendation says who does what and when", "the points in disagreement are listed separately from the decisions".

    Why not simply ask Copilot to check its work? Because research is clear on this point: Huang and his co-authors show that language models do not self-correct reliably without external feedback, and that performance can drop after self-correction. What works is checking against something outside the production: an attached source, or a grid whose criteria do not depend on what the model has just written. Hence the two rules of the lesson: the grid is written beforehand, and the evaluation takes place in a separate conversation, with only the text to be judged.

    The set-up has four stages. You write the grid. You produce version 1, with the usual production prompt. In a new conversation, you give Copilot a reviewer role, the grid, the text, and you ask for a score per criterion, a justification, then the three priority corrections phrased as rewriting instructions. You apply those you find justified, you refuse at least one if it is bad, and you replay the grid on version 2. The "Think deeper" mode, documented by Microsoft for analysis and verification, suits the evaluation stage. The booster of this lesson, "self-assessment criteria", adds a self-score at the end of the production prompt's answer: a useful checklist, not a proof.

    Diagram "A grid to judge an answer": Criteria: accurate, complete, sourced, tone; Score 0 to 2 per criterion; Decision: keep, fix, redo; Copilot applies the grid itself. Note: A grid written before the answer avoids judging by impressionDiagram "A grid to judge an answer": Criteria: accurate, complete, sourced, tone; Score 0 to 2 per criterion; Decision: keep, fix, redo; Copilot applies the grid itself. Note: A grid written before the answer avoids judging by impression
    Diagram "A grid to judge an answer"Diagram generated by AI and reviewed
  2. 2A board note that goes from 5 to 9 out of 10

    A manager had Copilot write a one-page note proposing an organisational change to his management.

    Weak prompt, in the same conversation:

    Is it good?
    

    Copilot replies that the note is "clear and well structured", suggests "strengthening the introduction" and stops there. Nothing is measurable, nothing is actionable, and the note goes out as is with two unsourced figures.

    Strong prompt, in a new conversation:

    You are a reviewer. Evaluate the note below against this grid. For each criterion, a score from 0 to 2 and a one-sentence justification:
    1) every figure refers to a named source;
    2) the target audience and the decision requested appear in the first sentence;
    3) the final recommendation says who does what and when;
    4) no sentence exceeds 25 words;
    5) the foreseeable objections are listed with an answer for each.
    End with the three priority corrections, phrased as rewriting instructions. Do not rewrite the note.
    Note: [pasted text]
    

    The grid comes back with a score of 5 out of 10: two figures without a source, a requested decision missing from the opening, four sentences too long. The three corrections are precise. The manager applies two, refuses the third, which removed an objection he wants to keep, and replays the grid: 9 out of 10.

    What changes: observable criteria produce a verdict comparable between the two versions, instead of an opinion. Corrections phrased as instructions can be reapplied without interpretation. The reasoned refusal of one correction shows that the referee remains the manager, not the model.

  3. 3Write five criteria checkable in one minute

    Choose a routine deliverable of your job: client email, briefing note, meeting report, product sheet. Here is the starting prompt, to be improved:

    "Read my note and tell me what you think."

    First write your grid of five observable criteria, specific to that deliverable, scored from 0 to 2. Produce a version 1 with Copilot, or take an existing anonymised text. In a new conversation, rewrite the prompt: reviewer role, grid, score and justification per criterion, three priority corrections as instructions, no rewriting allowed. The composer below structures the evaluation prompt; enable the "List the points a human must check" booster. Apply the corrections you find justified, replay the grid, compare the scores.

    Self-assessment grid: (a) the five criteria can each be checked in under a minute; (b) the evaluation took place in a separate conversation; (c) the before and after scores are recorded and at least one correction was refused with a reason.

    Open the prompt composer

  4. 4The grid that scores 10 out of 10 every time

    A project manager writes his grid with five criteria: "the tone is professional", "the memo is clear", "the structure is logical", "the content is relevant", "the level of detail is appropriate". He asks Copilot to score a briefing memo, improve it, then score it again. Twice in a row: 10 out of 10, with a glowing justification for each criterion. What he should have seen: none of the criteria is observable, so the model, judging its own output, ticks "OK" to please. The grid measures its goodwill, not the memo. Correction: rewrite each criterion so that a busy colleague could check it in under a minute without forming an opinion: "every figure points to a named source", "no sentence exceeds 25 words", "the recommendation says who does what and when". Have the scoring done in a separate conversation, require for each criterion the quoted passage that justifies the score, and check two criteria yourself at random. Rule to remember: a criterion the model cannot contradict measures nothing.

  5. 5Quiz

    Three questions, instant feedback. Each option comes with an explanation.

    1. What makes a grid criterion usable by Copilot and by a colleague?

    2. A support manager wants to evaluate his team's ten standard replies. Which prompt follows the lesson?

    3. Review: Copilot answers a legal question with great confidence. What does that tone tell you?

  6. 6Proof of mastery

    Paste your grid of five criteria, Copilot's evaluation of version 1 with the scores, the corrections you applied, then the scores of version 2.

    Advanced badgeThis lesson counts towards the Advanced badgeSee the four badges

    Criteria

Going further

This grid is prompt C of Breaking a task into three chained prompts, and it belongs in every card of Building your team's prompt library, under "sample output". The glossary recalls what a hallucination is, which the criterion "every figure refers to a source" is meant to catch.

Frequently asked questions

Why a separate conversation for the evaluation?

Because in the same conversation, the model defends what it has just written: it has the whole context of its production and tends to confirm it. A new conversation, with the grid and only the text to evaluate, produces a more independent judgement. It is still not an impartial judge: you remain the referee.

Can I ask Copilot to self-assess in the production prompt?

Yes, that is the "Self-assessment criteria" booster of this lesson, useful as a checklist at the end of an answer. But research shows that self-correction without external feedback does not improve reasoning, and can degrade it. Self-assessment flags; the grid applied separately, then your proofreading, decide.

How many criteria in a grid?

Five, rarely more. Each criterion must be observable in under a minute by a colleague: a length, a presence, a match with the source. "The tone is right" is not a criterion; "no sentence over 25 words" and "every figure refers to a source" are.

Sources