All articles

Leading TA39 Implementation

5 min read

Read: ~18 min · Plan it: 45–60 min to plan your first support cycle


In this guide

When a department or school starts using AI-assisted feedback, the leadership question is not simply, "Are teachers using it?"

That question is too thin.

The better question is:

Are teachers using it in ways that protect feedback quality, professional judgment, curriculum alignment, and student trust?

That is the work of implementation leadership. It is not about enforcing use. It is not about checking compliance. It is not about becoming the person everyone comes to for every technical issue.

It is about building the conditions where teachers can use TA39 well: clear standards, safe calibration, practical support, peer learning, and light evidence that helps the school know what is improving and what still needs attention.

If teachers feel watched, quality suffers. If teachers feel supported, practice improves.

That distinction shapes the whole role.

If you arrived here from Assessing Your Own Growth with TA39, you have already done the personal version of this work. This guide extends that same reflective practice to the team, department, or school you support.

For clarity, this article refers to an implementation "lead." In practice, that role may sit with one assessment lead, a department head, a coach, a small committee, or a distributed group of teachers. The structure matters less than the function: someone needs to protect quality, create safe support, and help the work become sustainable.

Your role: what you are and what you are not

An implementation lead is a quality facilitator, not the TA39 police.

That sentence matters. Teachers need to know the difference.

If the role feels evaluative, teachers are more likely to hide uncertainty, send weaker samples to calibration, avoid asking for help, or use the tool in shallow ways because they think usage itself is being judged. That is the opposite of what good implementation needs.

Your role is to help teachers develop AI feedback literacy: the judgment to decide when to use the tool, how to prepare rubrics and templates, how to review output, when to adapt, and when to use less assistance.

You are not:

  • Technical support. Login issues, bugs, and platform problems should go to the appropriate support channel.
  • An enforcer of mandatory use. Teachers retain professional agency. The goal is principled use, not maximum use.
  • A performance evaluator. You may review samples for quality patterns, but the purpose is support and calibration, not judgment of the teacher.
  • Responsible for solving every problem alone. Sustainable implementation distributes expertise over time.

You are:

  • A facilitator of quality calibration. You help teachers develop shared standards for what strong AI-assisted feedback looks like.
  • A monitor of pedagogical quality. You look for patterns in rubrics, templates, feedback, curriculum language, and teacher review.
  • A supporter of teacher growth. You help teachers understand what is not working and choose a practical next step.
  • A connector. You create spaces where teachers learn from one another instead of depending only on you.
  • A capacity builder. Your goal is a system that becomes less dependent on one person over time.

The role is not to catch teachers doing something wrong. It is to make good practice easier to see, easier to discuss, and easier to repeat.

Start with a small set of quality standards

Implementation becomes hard to lead when "quality" stays vague.

Everyone may agree that feedback should be good, specific, accurate, and helpful. But those words need a shared meaning before they can guide practice.

A useful quality standard for TA39-assisted feedback has six dimensions:

  1. Assessment tool quality. Rubrics, prompts, and templates are clear enough to guide useful output.
  2. Pedagogical quality. Feedback helps students understand what to do next.
  3. Voice and relationship authenticity. Feedback sounds like the teacher and respects the teacher-student relationship.
  4. Curriculum alignment. Feedback uses the right terminology and standards for the curriculum being taught.
  5. Student-focus evidence. Feedback responds to this student's actual work, not a generic version of the assignment.
  6. Teacher agency maintenance. Teachers review, adapt, and decide rather than accepting output passively.

These dimensions do not need to become a heavy rubric for every review. They are a shared language.

For a fuller review tool, use the companion AL Quality Indicators Framework.

In a calibration session, they help teachers say more precise things:

  • "The feedback is specific, but it is not yet student-focused."
  • "The tone is warm, but the curriculum language is too generic."
  • "The rubric is creating the problem; the feedback is only revealing it."
  • "This teacher has strong judgment in review, but the template is making them work too hard."

That shared language changes the conversation. Teachers stop talking only about whether the output is "good" and start naming what kind of quality is present or missing.

Build structures that distribute the work

One person cannot be the entire implementation system.

At the beginning, teachers may need visible support from a lead: calibration sessions, office hours, sample review, quick answers, and reassurance. But if every question and every quality concern returns to one person, the system will not hold.

The goal is to build structures that distribute the work.

Calibration sessions

Calibration sessions are the center of quality implementation.

The purpose is simple: teachers look at feedback samples together and develop a shared sense of what strong, usable, teacher-shaped feedback looks like.

A good calibration session does not need to be elaborate. A 60-minute structure is enough:

  1. Set the frame. This is professional learning, not evaluation. Samples are discussed to improve judgment, not to judge the teacher.
  2. Review one sample quietly. Teachers read the student work, rubric, draft feedback, and final feedback if available.
  3. Use the quality dimensions. What is strong? What needs revision? What should not be sent?
  4. Discuss patterns. Is the issue in the rubric, template, student context, curriculum language, or review step?
  5. Name one practice move. What should teachers try before the next session?

Run at least three calibration moments in a term: early after first use, mid-term when patterns are visible, and near the end when teachers can reflect on growth.

Keep the tone safe. The strongest sessions happen when teachers can say, "I sent something like this and now I can see why it was not ready," without feeling exposed.

Peer partnerships

Peer partnerships reduce isolation and make feedback review less dependent on formal sessions.

A simple partner protocol can take 20–30 minutes:

  • Teacher A brings one AI-generated draft and the final adapted version.
  • Teacher B reads both using one quality lens: specificity, voice, curriculum alignment, or student focus.
  • The pair discusses what changed and why.
  • They name one adjustment for the next assignment.

The point is not to grade each other's use. The point is to make teacher judgment visible.

Pair teachers by subject when curriculum language matters. Pair across subjects when the goal is voice, student focus, or workflow habits. Keep the structure light enough that teachers will actually use it.

Subject cohorts

Subject or curriculum cohorts are useful when the main quality issue is curriculum alignment.

In a school using multiple frameworks, teachers may need separate time to calibrate IB criteria, AP rubric rows, CBSE marking schemes, IGCSE command words, or local department standards. A general calibration session can build shared quality language. A subject cohort can make that language precise for the curriculum teachers actually use.

Keep the cohort focused:

  • one curriculum or subject area
  • one common assignment type
  • one sample rubric, template, or feedback set
  • one decision rule to carry forward

The goal is not another meeting. The goal is sharper curriculum language and fewer avoidable alignment errors.

Drop-in support

Drop-in support works best when it is informal and low pressure.

Teachers may arrive with platform questions, but often the real issue underneath is pedagogical:

  • "The feedback is too generic."
  • "It does not sound like me."
  • "The grade looks wrong."
  • "I am not sure this is the right assignment to use."
  • "Students may not trust it."

Respond by bringing the teacher back to the core questions:

  • Is the rubric specific enough?
  • Is the template shaped enough?
  • Does the feedback point to the student's actual work?
  • Does it sound like the teacher?
  • Has the teacher verified the grade?
  • Is this assignment a good fit?

Drop-ins are not workshops. They are moments to help teachers apply the principles in the messy reality of their own assignments.

Resource sharing

A shared library can help, but only if context travels with the resource.

A rubric or template is not self-explanatory. If teachers share only the file, others may reuse it in the wrong context or assume it is universally good.

Ask teachers to attach a short note:

  • Assignment:
  • Grade or course:
  • Curriculum framework:
  • Intended use:
  • Last revised:
  • What worked:
  • What to check before reuse:

That small note protects quality. It reminds everyone that shared materials are starting points, not finished answers.

For a more structured version of this work, use the companion AL Community Structures Planner.

Support teachers with curiosity first

When implementation problems appear, the first move should be curiosity.

Not because quality does not matter. It does. But because most problems have a reason.

A teacher may be sending feedback with too little adaptation because they misunderstood the workflow. Another may be over-relying on suggested grades because they assume disagreement means the tool failed. Another may be avoiding the tool because the first assignment was a poor fit. Another may be trying hard but using a vague rubric that keeps producing generic feedback.

The visible problem is rarely the whole problem.

Use a five-part conversation structure:

  1. Connect. Make the purpose safe and collaborative.
  2. Share the observation. Use specific evidence, not vague judgment.
  3. Listen and understand. Ask what is happening from the teacher's perspective.
  4. Problem-solve together. Offer options and choose a next step that is feasible.
  5. Commit and follow up. Agree on one action and return to it.

For example:

"I looked at three feedback samples from the last assignment. The comments were accurate, but they were very similar across different students. I wanted to understand what the review process looked like from your side."

That opening is very different from:

"You are not personalizing enough."

The first gives evidence and invites the teacher into analysis. The second creates defensiveness.

Good support conversations preserve dignity. They focus on the work, the workflow, and the next step.

For extended language, scenarios, and follow-up structures, use the companion AL Supportive Conversations Guide.

Common patterns and how to respond

Most implementation problems fall into a few recognizable patterns.

Over-reliance

The teacher accepts drafts or grades with little review.

This may come from time pressure, misunderstanding, or overconfidence in the output. The response should be clear but supportive:

  • Re-establish the principle: TA39 drafts; the teacher reviews, adapts, and decides.
  • Review one sample together.
  • Ask the teacher to identify what they would keep, revise, and verify.
  • Set a short-term expectation: for the next assignment, adapt three samples and compare before/after changes.

If grades are involved, be firmer. Suggested scores must be checked against the rubric descriptor and the student work. The teacher remains responsible for the final grade.

Under-use

The teacher avoids the tool after the first try or only uses it for the safest tasks.

This may be appropriate. Not every assignment is a good fit. But if the avoidance comes from uncertainty or a poor first experience, support should focus on a smaller, better-scoped use:

  • Choose one low-stakes formative assignment.
  • Use a rubric that is already reasonably clear.
  • Build or refine one template.
  • Review a small batch carefully.
  • Reflect on what worked before expanding.

The goal is not to push more use. The goal is to help the teacher make a principled decision.

Adaptation fatigue

The teacher is spending so much time editing that the workflow feels worse than before.

This often points back to setup. If the same edits happen repeatedly, the rubric or template should change.

Ask:

  • What do you keep changing?
  • Are next steps too long?
  • Is the tone too generic?
  • Is the rubric producing broad comments?
  • Is the feedback structure wrong for the assignment purpose?

Then turn the repeated edit into a setup improvement. A template that keeps needing the same fix is asking to be revised.

Confidence gaps

The teacher is not sure when to override, adapt, regenerate, or use less.

This is where calibration helps. Teachers need to see examples of decisions, not just hear principles.

Use sample feedback and ask:

  • What is ready?
  • What needs light editing?
  • What needs substantial revision?
  • What should stop the process?
  • What points back to rubric or template setup?

Confidence grows when teachers can name the decision they are making.

Red flags: when support needs to become more direct

A red flag does not mean a teacher has failed. It means the implementation needs attention before students are affected.

At the individual feedback level, Calibrating Quality helps teachers recognize red-flag feedback, and Evaluating AI-Generated Feedback names specific stop conditions before feedback is sent. The patterns below are the system-level signals visible to an implementation lead.

Some red flags should be addressed quickly:

  • Feedback is being sent with little or no teacher review.
  • Suggested grades are accepted without verification.
  • Students report that feedback feels robotic or unlike the teacher.
  • Feedback is generic across very different students.
  • Curriculum terminology is consistently wrong.
  • Feedback quality is worse than the teacher's previous practice.
  • Teachers are more stressed and the workflow is breaking down.
  • Sensitive or personal student writing is being handled too clinically.

Respond proportionally.

For a concerning pattern, a check-in and targeted support may be enough. For a critical pattern — unreviewed output reaching students, repeated grade-verification problems, or clear relationship damage — act quickly and privately.

The response should still begin with understanding, but the expectation should be clear:

"This feedback cannot go to students without teacher review. Let's look at what made that hard this time and decide what needs to change before the next assignment."

Support and standards belong together. Teachers need both.

For a fuller response guide, use the companion AL Red Flags & Response Protocols.

Gather the lightest evidence that still tells the truth

Evidence gathering can easily become too heavy.

If teachers feel that every sample, edit, and timing decision is being monitored, they will become cautious in ways that damage learning. If leaders gather no evidence at all, the school cannot tell whether implementation is improving feedback or merely increasing activity.

The useful middle ground is light, purposeful evidence.

Choose three or four indicators that matter most in your context.

Possible indicators include:

  • Feedback quality. Are comments specific, actionable, aligned, and teacher-shaped?
  • Teacher adaptation. Are teachers reviewing and changing drafts meaningfully?
  • Curriculum alignment. Is feedback using the correct standards, terminology, and assessment priorities?
  • Student uptake. Are students acting on feedback in revisions?
  • Teacher sustainability. Is the workflow saving time, increasing stress, or changing how teachers use their time?
  • Student perception. Do students experience the feedback as useful and connected to their teacher?

Then choose the lightest method that gives meaningful information:

  • three sample feedback reviews per teacher per term
  • one calibration sample per department meeting
  • short teacher reflection after a first run
  • student perception pulse survey once per term
  • before/after comparison of AI draft and teacher-adapted final feedback
  • revision uptake check on one assignment

Do not collect evidence because it is available. Collect evidence because it will change support.

The question is always: what will this tell us, and what will we do differently if we learn it?

For a structured planning tool, use the companion AL Evidence Collection Strategy.

Report patterns, not surveillance

Administrators may want to know whether implementation is working.

That is reasonable. But reporting needs boundaries.

Share patterns that help the school improve:

  • common rubric issues
  • common template needs
  • calibration themes
  • teacher support requests
  • evidence of student uptake
  • time and workflow patterns
  • resource needs

Protect information that would turn support into surveillance:

  • individual teacher mistakes shared without consent
  • raw feedback samples that identify students unnecessarily
  • simplistic usage counts presented as quality evidence
  • comparisons that rank teachers by adoption

Usage is not the same as quality. A teacher who uses TA39 less but reviews carefully may be using it more professionally than a teacher who uses it constantly but accepts output too quickly.

That nuance matters in leadership communication.

A useful report might say:

"Teachers are using the tool most successfully on formative, rubric-driven writing tasks. The strongest examples show clear teacher adaptation and student-specific next steps. The main support need is rubric clarity; several rubrics still contain vague level descriptors that lead to broad feedback. Next month, calibration will focus on rubric optimization and quality thresholds."

That kind of report gives leadership something meaningful without turning teacher learning into a compliance dashboard.

A first support cycle

If you are beginning from scratch, keep the first cycle simple.

Week 1: Set the frame

Clarify that the goal is quality, not compliance. Share the role distinction: facilitator, not evaluator. Name the core principle: teachers review, adapt, and decide.

Weeks 2-3: Support first use

Offer drop-in support. Help teachers choose appropriate first assignments. Encourage low-stakes, rubric-driven, revision-friendly tasks. Watch for setup problems: vague rubrics, generic templates, unclear prompts.

Weeks 3-4: Run the first calibration

Use two or three anonymized feedback samples. Ask teachers to sort them: strong, good but missing something, questionable, red flag. Identify what the samples reveal about rubrics, templates, review, and student context.

Mid-term: Look for patterns

Gather light evidence. Which issues are recurring? Are teachers adapting drafts? Are students using feedback? Are some curricula showing alignment problems? Choose one support focus for the next month.

End of term: Reflect and distribute expertise

Ask teachers what changed in their feedback practice. Identify useful rubrics, templates, and decision rules. Invite teachers who have developed strong practice to support peers next term.

The aim is not to finish implementation in one cycle. The aim is to create a repeatable rhythm: use, review, calibrate, adjust.

What success looks like

Successful implementation does not mean every teacher uses the tool for every assignment.

It looks more like this:

  • Teachers can explain when they use TA39 and when they choose not to.
  • Rubrics and templates improve over time.
  • Teachers adapt drafts rather than sending them untouched.
  • Students receive feedback that is specific, usable, and recognizable as coming from their teacher.
  • Curriculum terminology becomes more accurate, not less.
  • Calibration conversations become more precise.
  • Peer support increases.
  • Evidence is used to improve practice, not to pressure teachers.
  • The system becomes less dependent on one lead because more teachers can name quality and support one another.

That is a healthier sign than a high usage count.

Try it: plan your first support cycle

Set aside 45–60 minutes.

Use one page. Do not build a complex implementation plan yet.

1. Name your role

Write one sentence completing this frame:

My role is to support quality by...

Then write one boundary:

My role is not to...

2. Choose three quality indicators

Pick the three indicators that matter most for the next six to eight weeks:

  • rubric/template quality
  • feedback specificity
  • teacher adaptation
  • curriculum alignment
  • student uptake
  • teacher sustainability
  • student trust

3. Choose two support structures

Do not launch everything at once.

Choose two:

  • one calibration session
  • weekly drop-in support
  • peer partnerships
  • one shared resource library
  • curriculum cohort meeting
  • sample review cycle

4. Decide what evidence you will collect

Choose the lightest evidence that will tell you whether support is working.

Examples:

  • three anonymized feedback samples for calibration
  • teacher reflection after first use
  • one student perception question
  • before/after comparison of draft and final feedback
  • list of repeated rubric/template issues

5. Plan your first communication

Write a short message to teachers that makes the role safe:

This work is about feedback quality and shared learning, not evaluation. TA39 drafts; you review, adapt, and decide. The support structures this term are here to help us develop that judgment together.

If that message feels true, the implementation has a better chance of staying healthy.

For a fuller planning template, use the companion AL Implementation Planner.

Read Calibrating Quality: Is Your Feedback Good Enough? before leading a calibration session.

Read Evaluating AI-Generated Feedback: A Five-Part Protocol if teachers need a shared review language.

Read Assessing Your Own Growth with TA39 if you want teachers to reflect on their development before setting goals.

Read Hard Cases: Where Teacher Judgment Matters Most when teachers are struggling with sensitive, contextual, or curriculum-specific decisions.

Closing thought

Leading implementation is not about making teachers use a tool.

It is about protecting the quality of the professional decisions around that tool.

The healthiest implementation is one where teachers feel more capable, not more watched; where evidence informs support, not surveillance; and where the central promise remains intact:

TA39 drafts. Teachers review, adapt, and decide. Implementation leadership creates the conditions where that happens consistently, so students receive feedback they can actually use.