Read: ~10 min · Try it: 15–20 min on one student's feedback (3–5 minutes per student with practice)
There is a natural pull, when reviewing AI-generated feedback, to think in terms of accept or reject. The output is on screen. The cursor is hovering over the post button. The question feels binary.
That framing is too narrow. A more useful question is this: is the feedback accurate, is it fair, is it usable for this student? That is evaluation, not approval. And it is the same kind of judgment that applies to any serious assessment decision a teacher makes.
This article describes a five-part protocol for evaluating AI-generated feedback before it reaches a student. The protocol is structured but not bureaucratic — it gives names to checks that experienced teachers often perform instinctively, and slows the process down enough that those checks happen in fact rather than in intention.
The first time through, the protocol takes about fifteen to twenty minutes on a single student's feedback. After that, it takes three to five minutes per student. The time is real. It is also the difference between AI-assisted feedback and AI-replaced feedback.
The five parts
- Student-Focus Check — does the feedback serve this student, or a generic student?
- Curriculum-Sensitivity Check — does the feedback use the right language for the framework the student is being assessed against?
- Grade Verification — does the suggested score actually fit the descriptor and the work?
- Essential Teacher Additions — what does this feedback need that only the teacher can know?
- Red Flags — anything that should stop the process before the post button.
Each check has a guiding question. None of them require running a checklist on every comment. The point is to look at the whole piece of feedback through one frame at a time.
1. Student-Focus Check
Does this feedback serve THIS student's learning, or any student's?
AI-generated feedback can sound personalized while being generic. The structure is right. The rubric language is present. The next steps are named. But none of it would change if the student were someone else in the class.
Six quick reads catch the most common failures:
- Individual context. Does the feedback reflect this student's current skill level, recent progress, or particular struggles — or could it apply to anyone?
- Actionability for this student. Can this student understand the vocabulary used? Do they have the prerequisite skills to act on the guidance? Does the complexity match their developmental level?
- Tone appropriateness. What does this student need most right now — encouragement, challenge, reassurance, direct critique? Does the feedback's tone match?
- Cognitive load. How many distinct improvement areas does the feedback identify? Is that the right number for this student? Two or three is usually a useful ceiling; more becomes noise.
- Relational warmth. Would the student recognize this as coming from you? Is the voice present, or has it been flattened?
- Growth acknowledgment. Has this student shown progress since the last assignment, tackled something difficult, or taken a real risk? Is that visible in the feedback?
If most of these read as generic, the feedback is not yet ready to send — even if the content is accurate.
2. Curriculum-Sensitivity Check
Does this feedback use the right language for the framework students are being taught in?
Curriculum sensitivity matters because feedback works only when it makes sense inside the student's assessment world. A piece of feedback can be useful in the abstract and still drift from the framework that actually defines success for the student.
The check varies by curriculum, but the questions are the same:
- Terminology accuracy. Is the rubric language correct for the curriculum — Criterion A (not "Category A") for IB, Row A for AP, marking scheme terms for CBSE, AO1–AO4 and command words for IGCSE? Generic phrases like level or band are red flags here.
- Standards alignment. Does the feedback address the criteria or assessment objectives this assignment is actually being marked against? Are any criteria missing? Is any criterion getting more attention than its weighting deserves?
- Assessment preparation. Will this feedback help the student succeed on the curriculum's own assessments — the IB exam, the AP test, the CBSE board paper, the Cambridge assessment? Or is it generic feedback about writing that happens to mention the curriculum?
- Curriculum-specific skills. Are the relevant skills addressed — IB's Approaches to Learning, AP's disciplinary practices, the assessment objectives for CBSE or IGCSE?
- Progression stage. Is this an early-term task, a mid-term proficiency check, or an exam-preparation piece? Do the expectations in the feedback match the stage?
A piece of feedback that passes the student-focus check but fails the curriculum check is feedback the student cannot use — because it is pointing them somewhere their assessment is not going.
3. Grade Verification
Does the suggested score actually fit the descriptor and the work?
This is usually the slowest part of the protocol, and also one of the most important. The system's suggested grade is a starting point, not a verdict.
The check has a structure:
- Initial agreement. Read the suggested grade. Does it match your immediate professional impression? Complete agreement, mostly agreement (within one or two points), or significant disagreement?
- Criterion-by-criterion analysis. Where you disagree, open the rubric descriptor for that level. Compare the descriptor to the student's actual writing. Decide whether the descriptor fits the work better at the suggested level, or at the level above or below.
- Reasons for override. When you disagree, name the reason. Common ones: the system missed context you know about the student; the rubric did not capture quality you can see; the student showed growth that matters; cultural or linguistic factors; creative risk-taking the rubric does not reward.
- Feedback–grade alignment. Does the feedback tone match the final grade? A high grade with mostly critical feedback (or a low grade with mostly positive feedback) is a signal something needs adjusting — usually the feedback, sometimes the grade.
When you override, write down the reason in a short note. Over time, these notes show a pattern. The pattern usually points back to the rubric: a criterion the system reliably misreads, a descriptor that needs sharpening. That is useful information for the next assignment cycle.
4. Essential Teacher Additions
What does this feedback need that only you can know?
Some of what students most need from feedback is genuinely not generatable. The system has the rubric, the assignment, the writing. It does not have the conversation from last week, the family situation, the prior draft, the breakthrough moment, the lingering misconception. Some of that material is essential to the feedback.
Four categories cover most of it:
- Personal knowledge of the student. Recent conversations, previous work patterns, current circumstances, learning style, stated goals. What from this list, if any, must be reflected in the feedback?
- Classroom context. Recent discussions, lessons, activities, peer examples, resources used. What connections to class would make the feedback more useful?
- Relational elements. Specific praise for effort, acknowledgment of difficulty, personal observation, forward-looking encouragement. What is needed to maintain your relationship with this student?
- Strategic guidance. Upcoming lessons that will help, resources to consult, an office-hours invitation, specific practice activities. What does your teaching plan say about what comes next for this student?
Not every piece of feedback needs additions in every category. Some students need almost none. Others need a sentence or two that the system genuinely cannot produce. The check is whether anything is missing — not whether everything has been added.
5. Red Flags
Is anything here that should stop the process?
Some issues are not adjustments — they are stops. If any of the following appear, the feedback is not ready to send, no matter how strong the rest of it is.
- A factual claim about the work that is wrong. The system has stated something about the student's writing that, on inspection, isn't true. A misread quotation. A misattributed argument. A reference to a paragraph that doesn't exist. Factual errors destroy trust faster than anything else.
- Tone that would undermine the relationship. Sharp where it should be supportive. Cold where it should be warm. Dismissive of effort the student clearly invested.
- Any hint of bias or stereotype. Assumptions about the student based on name, language background, or apparent identity. Language that treats the student as a member of a category rather than as a writer.
- Comments focused on the student as a person rather than the work. "You always..." or "You tend to..." patterns. Feedback should address what is in the writing, not what the writing implies about character.
- Wrong curriculum terminology that the student is being assessed on. Naming an AP rubric row that does not exist, an IB criterion that has been removed, a marking scheme element that does not apply.
When a red flag appears, fix it before anything else moves. None of these are stylistic preferences. They are conditions for sending.
For the broader practice of recognizing calibration patterns — including red-flag patterns inside your own feedback workflow over time — see Calibrating Quality: Is Your Feedback Good Enough?
Try it: Run the Protocol on Real Feedback
About 15–20 minutes on one student's feedback. Print the feedback if possible — the protocol is easier with annotation by hand.
The protocol becomes faster with practice. The first time through is slower because every section is unfamiliar. Use this exercise to get one full run under your belt, with a real student's draft feedback.
Step 1 — Read the feedback once, all the way through (about 2 minutes)
Do not annotate yet. Read it the way the student will read it. Notice the first impression.
Step 2 — Run each of the five checks in order (about 10–12 minutes)
For each section, hold the question in mind and mark the feedback in pen.
- Student-Focus. Mark anything that reads as generic. Circle any tone moment that feels off for this student.
- Curriculum-Sensitivity. Underline any rubric or curriculum terminology. Verify each against the framework. Note any missing standard.
- Grade Verification. Open the rubric. Compare the suggested grade against the descriptor and the writing. Write your own grade beside the suggestion if you would mark differently.
- Essential Additions. In the margin, note one or two things that need to be added that the system could not have known.
- Red Flags. Scan once more, looking only for the five conditions above.
Step 3 — Make the adjustments (about 3–5 minutes)
Apply the changes inside TA39's Review & Transform screen. Soften tone where needed. Add the personal note. Correct the factual error if there is one. Override the grade if your judgment differs from the system's.
Step 4 — A short reflection (about 1 minute)
Note one pattern you saw — something the system did well, and something you had to adjust. Over the first five to ten times you run this protocol, those notes start to point at something useful: a rubric that needs sharpening, a template that needs voice, a tone instruction that needs to be tighter.
That reflection is what turns the protocol from a checklist into a professional skill.
A note on speed and practice
The first time through, the protocol feels long. By the fifth or sixth student, it feels closer to natural — the checks become things you notice rather than steps you follow. By the tenth, three to five minutes per student is realistic.
The time investment is not a cost. It is the work of being the evaluator rather than the acceptor. AI-generated feedback that has not been through this protocol is not yet feedback. It is a draft that has not been read with care.
What to read or watch next
Watch Reviewing Feedback in TA39: What to Look For. The product video walks through applying this protocol on a real Review & Transform screen, with the protocol visible as a side reference.
Continue Layer 3 with Module D: Planning Your Practice — My First TA39 Assignment Planner. Once your evaluation muscle is forming, the planner helps you choose the right first real assignment.
For the broader practice of recognizing calibration patterns once you have run several assignments, see Calibrating Quality: Is Your Feedback Good Enough? in Layer 4.
Closing thought
A clear rubric helps. A well-shaped template helps. A good first draft from the system helps. None of those replace the moment when a teacher looks at what is about to go to a student and asks: is this accurate, is it fair, is it usable for this student?
That moment is what makes the difference between feedback students read and feedback students actually use. It is also what makes AI-assisted feedback an extension of teacher judgment, rather than a substitute for it.