The educational proposition and its Cycle 2 application
TA39 is designed to make the formative-assessment loop visible and actionable—not simply record whether feedback was generated, viewed or followed by another submission. This is its purpose for teachers, schools and districts wherever the product is used. WISE Cycle 2 provides a structured opportunity to examine how well that educational purpose is being achieved.
Across TA39's established one-submission assessment pathway, the study can examine whether the evidence and final teacher-released feedback credibly represented the student's work and whether that evidence informed a teacher decision. When Revision Rounds is used, the study can also examine whether students identified and carried out a productive next move and whether their revised work demonstrated stronger evidence on the targeted learning criterion.
This creates four consequential educational questions that Cycle 2 can examine:
- Was the feedback educationally sound? Was the final teacher-released feedback accurate, rubric-aligned, specific, actionable, developmentally appropriate and sufficiently demanding without doing the student's thinking?
- Did the student make a meaningful next move? Did the student report clarity, select a focus, seek help where needed and make a traceable response in the next draft?
- Did the revised work demonstrate stronger criterion evidence? Among eligible students with sufficient comparable evidence and a genuine revision opportunity, did independent rating find stronger, similar, weaker or different valid evidence on the targeted criterion?
- Did the evidence help the teacher decide what to do next? Did it confirm, refine, change or reveal an instructional response—or provide no useful or sufficient additional information?
The measures therefore have a deliberate hierarchy:
| Research role | Measures |
|---|---|
| Primary student outcome | Targeted criterion-evidence change within supported revision, among students with sufficient comparable evidence and a genuine revision opportunity |
| Primary teacher outcome | Evidence contribution to the next instructional decision: confirmed, refined, changed, revealed something new, did not affect it, or not enough evidence to decide |
| Quality prerequisite | Independent validation of the final teacher-released feedback for the relevant language, task and model/prompt/rubric configuration |
| Mechanisms | Feedback view, reported clarity, priority/plan, traceability, feedback uptake and teacher review |
| Implementation and context | Opportunity, activity, pipeline, timeliness, participation, rubric/template activity, teacher intervention and assignment context |
| Exploratory later learning | Performance in a separately linked comparable task; deferred until the method and data exist |
Together, these measures form a clear educational and research story:
Credible evidence and feedback → meaningful student response and stronger criterion evidence → actionable evidence for the teacher
The strength of any conclusion depends on the evidence available:
| Evidence level | Responsible claim |
|---|---|
| Product evidence | Students received teacher-released feedback, identified a focus, revised and made a traceable response to the feedback. |
| Independently validated outcome evidence | Within supported revision, independently rated student work showed stronger, similar, weaker or different valid evidence on the targeted criterion, with the eligible cohort and evidence coverage reported. |
| Suitable comparative design | Students or assignments using TA39 demonstrated greater criterion-level change than a suitably comparable condition, if Cycle 2 adopts and successfully executes an appropriate comparative design. |
The proposition is not that every interaction proves learning. It is that TA39 creates a credible evidence chain through which meaningful student change and instructional response can be examined rather than assumed.
1.TA39's assessment foundation and Revision Rounds extension
TA39's established one-submission assessment pathway already supports a central formative-assessment process:
Student submits work → TA39 organizes rubric-aligned evidence of understanding → teacher reviews and adapts it → teacher releases personalised feedback → teacher and student can act on it
This pathway matters even where a teacher does not open another draft. It can show what the submitted work demonstrated, what TA39 prepared, what the teacher retained or changed, what feedback became available and whether the student opened it.
Revision Rounds extends—not replaces—that foundation:
Student responds and chooses a focus → revises → teacher examines the linked change → decides what should happen next
Revision Rounds preserves each draft, lets the teacher deliberately release feedback and open another opportunity, and connects feedback to the student's next action and the teacher's next instructional decision.
This is TA39's central educational contribution: it turns feedback prepared with TA39 and released by the teacher from an endpoint into a traceable formative-assessment process. Teachers can see what students did with the feedback and decide what should happen next. School leaders and researchers can examine whether the feedback was educationally sound, whether students made a meaningful next move, whether subsequent work showed stronger evidence on the targeted criterion, and whether the evidence informed the next instructional decision.
We recommend organizing the measurement approach around three connected questions:
- Credibility of the evidence and feedback: Did the evidence accurately represent what the student’s work showed, and was the final teacher-released feedback accurate, rubric-aligned, actionable and appropriate for the learner and task?
- Student response and learning evidence: Under what conditions did teacher-released, rubric-aligned feedback support students in identifying and carrying out a productive next move without replacing their cognitive work? What did their subsequent work demonstrate on the targeted criterion, including where the loop did not progress?
- Value for the teacher: What credible student and class-level needs became visible, and did teachers report that this evidence confirmed, refined, changed or revealed their next instructional response—or added no useful information?
Understanding the requested action and demonstrating the intended learning are different. Both require more than submission or comparison counts, so we recommend mixed evidence:
- Product records show what happened in the feedback and revision process.
- Student voice explains clarity, usefulness, trust, and why students acted, deferred or sought help.
- Teacher voice and professional evidence explain credibility, judgment, and whether evidence changed what happened next.
Some evidence exists in the established assessment workflow; some requires the Revision Rounds V2 product development or separate study instruments. Product records preserve the work-to-feedback and feedback-to-revision chains but cannot alone show that feedback caused a change, that a student valued it, or why a teacher acted. Professional-development contribution must also remain distinct from software use.
Three related constructs must therefore remain separate:
Feedback uptake ≠ criterion-level evidence change ≠ transfer
Uptake asks how the student responded to the feedback. Criterion-level improvement asks whether the revised work shows stronger evidence of the intended learning criterion, whether or not the student followed the suggested route. Transfer asks whether that learning appears later in a suitably comparable task with less immediate support.
Product-wide safeguards for transparency and authentic learning
These safeguards apply to every TA39 assessment, not only Revision Rounds:
- Students should receive an age-appropriate assignment expectation before work begins. TA39 should preserve what was known about its visibility, never infer missing guidance, and not misclassify approved accommodations as prohibited AI use.
- Students should receive an accurate account of TA39’s and the teacher’s roles. The product must not imply item-level teacher review when only release is known.
- TA39 should make the learning problem actionable without replacing the cognitive work required to solve it. Feedback and planning support must not do the student’s reasoning or writing.
- Any student AI-use declaration is optional, non-punitive, and separate from learning measures. It never blocks feedback, revision, or support.
- When ordinary work is insufficient, the teacher may choose another suitable Show what you know demonstration. It remains separate from draft-to-draft uptake.
- TA39 must not turn changes, declarations, style, or revision behavior into an authorship finding, risk score, or automated penalty. Any concern returns to the work, the student’s explanation, teacher judgment, and school policy.
The purpose is to understand learning and decide what support comes next—not infer wrongdoing from behavioral traces.
2.From activity measures to evidence of learning and instructional action
WISE has identified a useful set of operational measures: when TA39 is used, submission volume, teacher review activity, revision depth, rubric use, subject and curriculum context, and teacher changes to AI-drafted feedback. We recommend retaining all of them and connecting each one to the complementary evidence needed for a sound educational interpretation.
| WISE area | What the product record contributes | What complementary evidence completes the claim |
|---|---|---|
| Timing | How quickly students submit and teachers review or release | Feedback-quality evidence and teacher context for interpreting short or long review periods. |
| Volume | How often TA39 and Revision Rounds were used | Eligible-cohort outcomes and teacher/student explanations of selective or non-use; greater frequency is not automatically better. |
| Feedback pipeline | Whether teachers released, edited, rewrote or rejected drafted feedback | Structured intervention reasons, teacher professional evidence and independent feedback-quality review; intervention may demonstrate good judgment. |
| Revision depth | How many rounds, attempts and comparisons occurred | Reported clarity, traceability, uptake and targeted criterion-evidence change. |
| Rubric use | How rubrics were created, reused, edited or optimized | Review of task-rubric-template fit and the intended learning goal. |
| Subject and curriculum | Where TA39 was used | Assignment-readiness and comparability rules that determine which results may be interpreted together. |
| Quality and trust | Where teachers intervened and how students responded | Independent quality review and routed teacher/student voice about credibility, clarity, relationship and trust. |
For the shared assessment foundation, we therefore recommend nine computational measure families. They retain WISE's requested activity measures while connecting them to an interpretable assessment opportunity:
| Measure family | What it contributes |
|---|---|
| Assessment activity and timing | Where and when the eligible teacher/student workflow occurred, without treating greater frequency as better. |
| Submission opportunity | A defensible roster denominator, including true no-submission and excluded/unavailable states. |
| Initial skill evidence | What the selected student work demonstrated by applicable criterion, with insufficient evidence preserved. |
| Feedback-pipeline outcome | Released unchanged, edited, substantially rewritten, rejected or not released. |
| Feedback timeliness | Submission-to-analysis, teacher-review, release and student-visibility intervals. |
| Feedback view | Whether a student opened feedback that was genuinely visible and renderable. |
| Teacher intervention | Structured reasons for changes/rejection and a properly bounded AI-issue subset, with reason coverage. |
| Rubric and template activity | Creation, adaptation, reuse and AI-optimization outcomes—without claiming quality from activity. |
| Assessment context | Grade, subject, language, task, curriculum/framework and instrument versions needed for valid interpretation. |
Revision Rounds adds eight measures: cycle depth, clarity response, clarification, priority/plan, revision participation, feedback-to-change traceability, uptake and targeted criterion-evidence change. The last is conditional on a validated human-rating method and is not calculated by subtracting AI-generated rubric scores. The teacher loop adds explicit evidence review, student/class need, instructional decision and follow-up fulfillment. Later learning remains a separate deferred measure requiring comparable later work.
The operational dataset should still preserve school-hours versus after-school timing, active teacher-review duration, and students with zero submissions. The last measure requires an explicit denominator of eligible rostered students who had a genuine opportunity to submit; otherwise “zero” is not interpretable.
The additional measures rest on two requirements:
- Credible: distinguish requested action, surface edits, meaningful response, later recurrence, and missing evidence.
- Useful for teaching: reveal needs, task/rubric fit, a teacher response, and a suitable next opportunity.
In short, the evidence should be true enough to trust and useful enough to teach from.
The quality prerequisite is a gate, not another descriptive metric. R-07 feedback uptake and R-08 targeted criterion-evidence change should not be presented as headline learning-impact evidence until the relevant feedback-quality and outcome-classification methods have passed the agreed independent validation. Every R-07 or R-08 result must also show the complete denominator cascade: eligible → feedback available → viewed → revision opportunity → revised → traceable → sufficient evidence → outcome. This prevents a positive result among only the final analyzable cases from being mistaken for an outcome for the whole class.
3.The student process: from feedback to later learning
Receive feedback → Understand it → Choose what to work on → Revise → Check the result → Use the learning again
Where a reliable, versioned method exists, the original draft can provide starting evidence for each relevant skill: demonstrated, partly demonstrated, not yet demonstrated, or insufficient evidence. Otherwise the starting point should be reported as not measured, not inferred from submission. The six stages then provide progressively stronger evidence: access is not understanding, another draft is not necessarily meaningful improvement, and immediate improvement is not transfer.
The central evidence requirement is a traceable connection:
Teacher-released feedback point → student-selected action or plan → relevant change in the next draft → teacher-reviewed outcome
Each stage should preserve that connection. Without it, TA39 can show that the writing changed, but cannot credibly say that the change was linked to the feedback. Where the connection cannot be established, the result should be reported as a change observed—not as feedback uptake.
| Stage | Why it matters | What we recommend measuring | Example of what the evidence could show |
|---|---|---|---|
| Receive feedback | Establishes opportunity; a missing revision cannot be interpreted as failure if feedback was not opened. | View rate: learners with a qualifying view ÷ learners for whom feedback was visible and renderable; state whether the view preceded the next action | “27 of 30 opened it before their next action — 90%.” |
| Understand it | Shows whether the requested action was clear or help was needed. | Reported clarity: know what to do, not sure, or ask teacher; report committed-response coverage over eligible checkpoints. Sample students may explain the action in their own words. | “22 of 25 responded: 15 knew what to do, 5 were unsure and 2 asked their teacher; response coverage was 88%.” |
| Choose what to work on | Turns feedback into a manageable intention. | Priority selection: learners choosing a focus ÷ learners eligible to choose. Plan rate: learners saving a plan ÷ learners eligible to plan; valid no-plan paths remain separate. | “21 of 24 chose a focus; 14 of 18 learners eligible to plan saved one, while 4 explicitly continued without a plan.” |
| Revise | Shows action, not yet effectiveness. | Participation: learners submitting the target draft ÷ the frozen cohort with visible feedback and a real open target-round opportunity | “23 of 26 revised — 88.5%.” |
| Check the result | Separates response to feedback from stronger evidence of the underlying learning. | Uptake distribution: how the revision responded to the feedback. Targeted criterion-evidence change in supported revision: stronger, similar, weaker or different valid evidence on the intended criterion, with insufficient cases separate. The criterion measure requires a validated human-rating method and is not a score subtraction. | “Of 50 sufficient feedback-item cases, 21 were substantive uptake. In the independently rated criterion sample, 18 of 30 showed stronger evidence, 8 similar evidence, 2 weaker evidence and 2 a different valid approach.” |
| Use the learning again | Shows persistence or recurrence. | Future: each later outcome ÷ learners with sufficient evidence in an explicitly linked comparable later task; show no-comparable-work and insufficient evidence separately | “Of 10 with sufficient reviewed later evidence, 5 sustained the skill, 3 showed stronger evidence and 2 showed recurrence.” |
For an individual student, these stages should not be collapsed into an artificial impact score. A teacher should see a short evidence summary:
Meaningful response shown: The student viewed the feedback, asked for clarification, planned two changes and submitted a revision. One priority was substantively addressed and one partly addressed. The revised work showed stronger evidence on one targeted criterion; the other was similar. Later use of the skill has not yet been measured.
4.The teacher process: from evidence to the next teaching decision
Review the evidence → Notice student needs → Decide what to do next → Give students another opportunity → See whether the learning continues
Evidence becomes formative when it changes what happens next. TA39 may organize evidence and suggest patterns, but the teacher reviews, interprets and decides.
| Stage | Why it matters | What we recommend measuring | Example of what the evidence could show |
|---|---|---|---|
| Review the evidence | Keeps AI-assisted interpretation from becoming the final judgment. | Review coverage: report individual review, sampled-batch confirmation, deferred, unreviewed and unavailable cases; reviewed cases ÷ cases with reviewable evidence | “24 of 28 reviewable cases were reviewed; 18 individually and 6 through a sampled batch.” |
| Notice student needs | Shows whether a need is isolated or shared without assigning its cause. | Need prevalence: classified need ÷ enough evidence; report individual review, sampled-batch confirmation and unreviewed TA39 suggestions separately | “12 of 24 still needed support with evidence; review depth is shown beside the result.” |
| Decide what to do next | Shows the decision that followed and whether the teacher says the evidence contributed to it. | Decision rate: reviewed patterns with a decision ÷ patterns reviewed. Separately, eligible teachers report whether the evidence confirmed, refined, changed or revealed a response, did not affect it, or was insufficient for a decision. | “The teacher reported that the evidence revealed a need they had not noticed and chose targeted practice for five students.” |
| Give another opportunity | Connects a decision to student action. | Follow-up fulfillment: linked activities created ÷ intended follow-ups whose cutoff has passed or whose state is terminal; pending work remains separate | “Eight of ten due follow-ups were created, one cancelled and one expired — 80% fulfilled.” |
| See whether learning continues | Shows what happened in linked later work. | Future follow-up outcome: each outcome ÷ students with enough linked evidence | “Of 18 students with sufficient linked later evidence, 14 sustained or showed stronger evidence on the targeted criterion.” |
When a class pattern appears, the teacher considers instruction, feedback clarity, time and support, and whether the task and rubric allowed students to show the skill. TA39 records the chosen response but does not assign the cause.
If stronger individual evidence is needed, the teacher may choose a brief Show what you know demonstration. It is neither an authorship test nor a class outcome measure; any research sample must disclose how students were selected and how many completed it.
A short post-round question can ask whether evidence was credible and affected the teacher’s response; opening a dashboard does not establish value. Nor is shorter review automatically better: where teachers report saved time, ask how it was reinvested and what benefit they observed.
5.Why curriculum, task and rubric context matter
Useful feedback and interpretable revision evidence depend on the assignment, rubric and feedback template fitting the writing and learning goal.
Curriculum variation is not simply a weakness: teacher autonomy supports creativity and responsiveness. Measurement must distinguish purposeful variation from weak task-rubric-template alignment. Professional development can strengthen shared principles without replacing teacher choice.
Consider two Grade 8 classes. In the first, students read a shared article and write an argument using evidence from that text. The task and rubric clearly define a claim, relevant evidence and explanation. In the second, students complete an unrelated writing task using a generic criterion such as “uses evidence effectively.” Both classes may complete two revision rounds and show the same submission rate, but the evidence is not equivalent. In the second class, apparent lack of progress may reflect unclear criteria or poor task-rubric fit rather than a student learning failure.
We therefore recommend preserving the grade, subject, curriculum or framework, source material, task type, intended skill, and rubric and feedback-template versions for each assignment. Results should be compared only where learning contexts are suitably comparable.
Not every assignment should enter the uptake or criterion-evidence analysis. The research team should first complete a six-item, pass/fail Assignment Measurement Readiness Record: the targeted construct is explicit; the task elicits it; the rubric criterion is interpretable; the feedback offers a plausible next move; source and target work are comparable; and the student had a genuine opportunity to revise. This is an eligibility record, not a score, and each failed or unavailable condition remains visible. Validation of the feedback and outcome-classification methods for the relevant language, task and model/prompt/rubric configuration is a separate gate. Other assignments can still contribute implementation measures without being treated as learning-outcome evidence.
For Qatar, validity should be examined explicitly across English and Arabic, grade bands, subjects, rubric structures and relevant accommodations, including independently double-rated English and Arabic samples. Before comparing student outcomes across groups, the protocol should pre-register checks for feedback quality, insufficient evidence, traceability, genuine opportunity, teacher correction and rater reliability. Differences in those measures should first be investigated as measurement behavior—not interpreted automatically as differences in student attainment.
For cross-school research, WISE may separately consider optional shared anchor tasks designed with Qatar educators: grade-appropriate texts leading from comprehension and inference to a short evidence-based response. A research-only near-transfer task could then use a different but comparable text or prompt, with the earlier feedback hidden and the later work independently rated. This would examine less-supported later use under the study protocol; it would not activate TA39's deferred product transfer measure, require TA39 to replace school curricula, or become routine product telemetry.
6.Complementary post-cycle study instruments
Product activity cannot explain whether feedback felt clear, trustworthy or worth acting on, why a teacher accepted or changed a suggestion, or what else influenced a revision. Surveys, interviews and open responses are therefore explanatory evidence, not a satisfaction add-on. Route them by actual experience and use selected post-round, mid-cycle or end-of-cycle points rather than interrupting every interaction.
Student feedback experience and understanding
Begin with neutral screening: Did you open or read the feedback? and Did you revise after receiving it? Ask only questions the student is in a position to answer:
| Student’s position in the process | Appropriate questions | Questions not to ask |
|---|---|---|
| Feedback was not made visible | Do not send the feedback-experience instrument. Count the case in product coverage and examine the reason through implementation records or the teacher study. | Clarity, usefulness, trust, teacher voice or reasons for acting. |
| Feedback was visible but not opened | Awareness, notification, access, time, technical difficulty and why it was not opened. | Whether the content was clear, useful, personal, trustworthy or worth acting on. |
| Feedback was opened but no plan or revision followed | Clarity, relevance, trust, whether support was needed, and why the student did not act or could not act. | Whether a particular revision resulted from the feedback. |
| A plan or revision followed | Which feedback mattered, whether it influenced a specific change, why the student acted, and what other influences contributed. If clarification was requested, ask whether a response arrived and helped. | Claims that the response proves causation. |
For eligible students, examine clarity and manageability of the next step; timing, relevance and usefulness; trust, personal connection and consistency with the teacher's voice; and whether TA39's and the teacher's roles were understood. Two short routed statements should be included for students who opened the feedback: “The feedback gave me enough direction to think for myself” and “The amount of feedback felt manageable.” A small research sample can explain the requested action in its own words and what it is intended to improve, or state what the student would look for to judge whether the revision became better. Where a revision followed, a light-touch routed question can record major additional influences—such as a teacher conversation, peer feedback, tutoring, an example, family help or another AI tool—without introducing surveillance telemetry. Office-hours questions require a teacher log or interview and must not be inferred from silence.
Report eligibility, invitations, responses, period and cohort for each branch. Trust and value are reported among students who opened feedback, not all enrolled students. Anonymous responses are not linked to an individual revision history. Perception evidence explains experience; it does not prove uptake or causation.
Teacher evidence and value
Represent all participating teachers, including selective, low and non-users. Ask about credibility or instructional value only when the teacher reviewed the relevant evidence.
| Teacher’s position in the process | Appropriate questions |
|---|---|
| Did not release feedback or use Revision Rounds | Implementation conditions, fit with the assignment or curriculum, access, time, training, and reasons for selective or non-use. |
| Released feedback but did not review the revision evidence or class view | Feedback-review experience and reasons the later evidence was not opened or usable; do not ask whether that evidence changed instruction. |
| Reviewed evidence or recorded a response | Credibility, sufficiency, professional corrections, whether it affected the intended response, what students received next, and what later evidence would be needed. |
A structured review can examine feedback actionability, curriculum fit, teacher voice, student questions, differences across groups, evidence sufficiency, the teaching decision that followed, and how any saved time was reinvested. For each sampled eligible case, the teacher should record whether the evidence confirmed, refined, changed or revealed a response, did not affect the decision, or provided not enough evidence to decide; nonresponse remains missing rather than being inferred. The review should also ask why the teacher retained, changed or rejected drafted feedback. Appropriate intervention—including substantial rewriting or rejection—may demonstrate the professional judgment TA39 intends to support; a lower edit rate is not inherently better. Private narratives use an agreed qualitative method and are not transmitted as raw product events.
Independent validation of feedback and learning evidence
Before downstream student behavior is interpreted as evidence that the system worked, a sampled set of final teacher-released feedback should be independently rated for evidence accuracy, rubric alignment, instructional soundness, specificity, actionability, prioritization, cognitive demand, developmental appropriateness, task/teacher fit and safety or fairness. The validation sample should include independently double-rated English and Arabic work and relevant grade, subject and accommodation groups. The agreed protocol should define what constitutes a pass for each relevant language, task and model/prompt/rubric configuration; until it passes, R-07 and R-08 remain product suggestions or not measured, not headline learning-impact evidence.
The same AI pipeline must not generate the feedback and serve as its sole validator. Human raters should also validate the uptake and targeted criterion-evidence methods using agreed coding guidance, agreement analysis and error analysis by relevant context. Student work and feedback text remain in the authorized research environment; they are not transmitted through the routine usage-events API.
Teacher development in using AI-assisted feedback
Professional development includes assignment and rubric design, teacher involvement in writing, informed adaptation or rejection of drafted feedback, and using revision evidence to plan what students need next. Editing frequency cannot show that development. Pair baseline and end-of-cycle reflection with an applied exercise in which teachers review sample feedback, decide what to retain, change or reject, and explain why. Analyze product contribution, professional-development contribution, and their combined implementation separately where the study design permits.
7.What is available now—and what would be added
The final plan should distinguish:
- Established TA39 assessment foundation: assignment/rubric context, immutable submissions, rubric-aligned analysis, teacher review/adaptation, personalised feedback and the available ordinary-assessment evidence chain.
- Revision Rounds V2 product development: deliberate versioned release/view, clarity, priorities/plans, stable feedback-to-revision linkage, explicit evidence review, class patterns, teacher decisions and linked follow-ups. Later-work evidence remains deferred until comparable-task requirements are met. Product-wide AI expectations and role disclosure remain context, not learning measures.
- Separate study and PD evidence: routed student/teacher instruments, explanation samples, applied teacher calibration, relevant participation or practice changes, and any approved research-only near-transfer task.
- Independent validation and optional comparison: sampled feedback-quality and criterion-evidence ratings, language/context validity checks and—if causal claims are intended—an agreed comparison such as a within-teacher crossover using suitably comparable assignments.
Where TA39 is introduced with professional development, the analysis should distinguish three related questions:
- Product contribution: What changed when and where teachers and students used TA39 and Revision Rounds?
- Professional-development contribution: What changed in teachers’ knowledge, judgment, assignment design and feedback practice?
- Combined implementation: What changed when the product and professional-development approach operated together?
Unless the design separates causal effects, findings should be associated with TA39 and its accompanying professional-development approach, not software alone. Product records support language such as followed, was linked to, was associated with or was traceable to. Student and teacher reports support reported that it influenced, confirmed, refined, changed or revealed. Claims such as improved, led to or caused require an appropriate comparative design. Student and teacher-loop evidence is reported per assignment/round; usage weekly or monthly; later learning after a comparable task and by term; student/teacher experience at agreed mid- and end-cycle points; and teacher development at baseline and cycle end.
8.What we recommend sharing with EdTech Impact
We recommend privacy-safe records supporting the agreed operational scope—roster and opportunity, timing and submission, feedback pipeline, rubric activity, and subject/curriculum context—and the learning measures: feedback release/view, clarity, plan, revision, uptake, targeted criterion-evidence change and review provenance, student-level evidence from which class need is calculated, teacher decision, and linked follow-up. Stable research identifiers replace names and emails; assignment, round, time and relevant learning context remain attached. Independently coded feedback-quality results and qualitative instruments remain separate study evidence unless WISE and EdTech Impact expressly agree on a structured, privacy-safe research record.
Approved AI-policy and role-disclosure versions may travel as context, but visibility must state only what TA39 can establish. Optional declaration or demonstration categories remain local unless specifically authorized. TA39 should not send student work, feedback or plan text, open responses, private reviews, direct identifiers, recordings, handwriting, or chat—and no record may become an authorship probability or disciplinary trigger.
Any summary includes its educational construct, unit of analysis, eligible cohort, observation window, numerator/category counts, denominator, exclusions, deduplication rule, filters, provenance, interpretation boundary and metric version. Every R-07 or R-08 summary also carries the complete denominator cascade from eligibility through outcome; any inferential estimate includes its uncertainty. The companion TA39 Learning Impact Measurement Dictionary — External Edition v1.1 supplies those controlling specifications and worked examples. EdTech Impact's API should map to that educational contract, not determine it.
TA39 recommends atomic, pseudonymous records rather than invented weekly aggregates. The receiver may calculate the measures only by applying the same cohort, window and exclusion rules. If TA39 also sends a calculated summary, it carries the complete denominator and coverage fields so it can be reconciled with the atomic evidence. There is no single TA39 “impact score.”
9.Recommended agreement
Before implementation, we recommend agreement on:
- Primary outcomes, mechanism and implementation measures; calculations; starting-evidence, uptake and targeted criterion-evidence methods; review provenance; and insufficient-evidence rules.
- Which evidence exists now, requires product development, or belongs to a separate routed study instrument.
- The student/teacher mixed-method and professional-development design, including eligible audiences, reporting periods, major co-interventions, reasonable burden and a small student advisory review of age-appropriate explanations and checkpoint language.
- The companion dictionary, atomic records versus summaries, and privacy, retention and deletion rules.
- Independent feedback-quality and method validation, including English/Arabic and relevant context checks; whether to include a research-only near-transfer task; and whether Cycle 2 is descriptive or includes an appropriate comparison for stronger impact claims.
- Product-wide transparency context and the absolute safeguard that no usage, declaration, revision-behavior or impact measure becomes an authorship detector, automated grading decision or disciplinary trigger. Any separate school-led concern process requires human review and an age-appropriate opportunity for the student to explain their process; TA39 does not decide or recommend the outcome.
This approach retains the operational measures WISE requested while addressing the full Cycle 2 inquiry: Did the evidence and released feedback credibly represent the student’s work, did the student respond meaningfully and show stronger evidence on the targeted learning criterion, and did that evidence contribute to the teacher’s next decision?