Feedback & Assessment

How do you stay in control when using AI to mark student work?

You stay in control by treating AI as a drafting assistant, not a decision-maker. Set the rubric and success criteria yourself, have the tool propose feedback against them, then read every draft alongside the student's work, edit what is off, and approve each comment before a student sees it.

What does human-in-the-loop actually mean when marking?

Human-in-the-loop means the machine drafts and a person decides. Applied to marking it is a narrow claim: no AI output counts for anything until a teacher has read it against the student’s work and chosen to keep it. Everything else here is about making that boundary hold on a Sunday night with thirty scripts left.

The boundary is easy to state and easy to lose. It slips when you approve a batch without opening the submissions. It slips when a fluent comment reads so well that you stop asking whether it is true of this student. It slips when you see a proposed grade before you have formed your own. None of those are technical failures. They are workflow failures, fixed by changing the order in which you do things.

It also helps to be clear about what a language model does when it drafts feedback. It produces text that fits the pattern of good feedback, given the work and the criteria you supplied. It is not measuring anything. Fluency is not evidence of accuracy, and a well-written comment about a paragraph the student never wrote is still wrong.

Why does reading the work before the draft matter so much?

Read the student’s work first, form a rough view, then open the draft. Reversing that order is the most common way teachers lose control without noticing. Once you have seen a proposed grade it becomes the number you adjust away from, rather than the conclusion you reach yourself. It is the same pull that makes the first script in a pile quietly set the standard for the rest.

This costs less time than it sounds. You do not need a full second marking pass. Skim the response against your success criteria, decide roughly where it sits, note the one or two things you would say, then read the draft. Where the two agree, approve and move on quickly. Where they disagree, you have a real disagreement to resolve instead of an unexamined suggestion to wave through.

The disagreements are the useful part. Sometimes the draft has caught something you skimmed past. Sometimes it has credited a strength the student never demonstrated. Either way, go back to the evidence in the script. Treat a mismatch as a prompt to look again, never as a verdict.

What should you check before approving a drafted comment?

Run the same short check on every draft, in the same order, so nothing gets skipped when you are tired. A fixed sequence does not depend on your attention being good that particular evening.

Where something is off, edit in place rather than regenerating. Usually one sentence needs changing, not the whole comment. Regenerating also hides the problem: you never learn what your setup keeps getting wrong, so you go on fixing the same fault every week.

Two failure patterns are worth watching for. The first is invented evidence, where a comment credits an argument or a technique to a student who did not use it. The second is criteria drift, where the comment is sensible advice about writing in general but never touches the standard being assessed. Criteria drift is the worse of the two, because it looks fine and teaches the student nothing.

  • Accuracy: does every claim in the comment describe something the student actually wrote?
  • Alignment: does it name the right criterion, in the wording your class already knows?
  • Grade: is the proposed mark one you would defend to a colleague reading the same script?
  • Specificity: does it point to a place in the work, rather than praising in general terms?
  • Next step: does it give the student one thing to do next, not five?

How do you write marking guides that a draft can actually hit?

Most of your control is exercised before any student work is uploaded. A drafting tool works from the rubric, success criteria and comment banks you give it, so those guides set the ceiling on draft quality. Vague criteria produce vague feedback, and editing afterwards never buys back that time.

Write criteria that describe observable features of the work rather than qualities of the student. “Uses evidence from at least two sources and explains how each supports the claim” can be checked against a script. “Shows good analytical thinking” cannot, by you or by anything else. Where a standard has grade descriptors, put the actual wording in, including the distinctions that separate one grade from the next, because those distinctions are exactly where drafts go wrong.

Comment banks do the other half of the job. Load the phrasings you already use, including the ones you would only say to a student close to giving up. That is what stops drafted feedback drifting into a generic, upbeat register no teacher in your department would recognise. Change a guide once rather than correcting the same thing thirty times.

It also helps to know what the student will do with the comment. If the gap is a one-off slip, a sentence is enough. If it is structural, point them at practice tasks or study guides they can work through before the next assessment, so the next piece of work is different rather than the next comment being longer.

Why does teacher control matter under NCEA?

Internally assessed standards are marked in school, and NZQA runs external moderation over those decisions. Its stated purpose is assurance that assessment decisions, in relation to assessment standards, are consistent nationally. That only works if the decision recorded is one an assessor actually made and can explain.

This is not an argument against drafting with AI. It is an argument for keeping the trail that shows where your judgement entered. If a moderator or your Principal’s Nominee asks how a grade was reached, your annotations and assessment schedule should carry your reasoning, not just a paragraph you accepted.

NZQA already publishes what you need in order to calibrate: internal exemplars, clarifications on standards and National Moderator’s reports on the subject pages, plus an Assessor Practice Tool for practising judgements on full samples of student work. Set your own benchmark with those first, then check that the drafts you approve sit inside it.

Authenticity runs the other way too. Your school’s assessment policy and current NZQA guidance govern what students may use, and those rules are separate from what staff use to mark. Be explicit with classes about which is which.

How do you build a review routine that saves time without losing control?

Batch by task, not by student. Take one question or one criterion across the whole class, review those drafts together, then move to the next. You hold one part of the standard in your head at a time, which makes inconsistency visible and each decision quicker.

Sample before you trust. On a new task, mark the first five scripts yourself with the drafts hidden, then compare. If you agree on four or five, the setup is sound and you can review at pace. If you do not, the fix is almost always in the criteria rather than the individual comments, and making it now saves correcting the same error twenty-five more times.

Keep a note of what you change. If you delete the same over-generous sentence every week, that says something about your comment bank or your grade descriptors, not about that student. Treat your own edits as feedback on the setup. And decide in advance what the saved time is for, because otherwise the workload simply refills.

Where does a tool like Jeddle fit into this workflow?

Everything above applies to any drafting tool. This is the only section about a specific one. Jeddle is an AI marking platform for secondary teachers, and its feedback engine, JeddAI, is built around the proposal-and-approval loop described here: it drafts against the rubric, success criteria and comment banks you supply, and nothing reaches a student until you approve it.

The practical differences from a general chat assistant are mostly about where control sits. Marking guides are stored and reused rather than pasted in each time, so changing a criterion applies to every future draft. Drafts sit beside the submission, which makes reading the work first the default rather than an act of discipline. Grades are proposed and never set. The same review habits apply either way, whether you use a general assistant or something built for expert essay marking.

None of that removes the work that matters. It removes the retyping. You still read the script, you still make the call, you still own the grade. If a tool ever makes it easier to approve without looking than to look, that is a reason to change the workflow, not a feature.

Who decides what at each stage of AI-assisted marking
Stage What the tool does What you decide
Setup Reads the rubric, success criteria and comment banks you supply What counts as Achieved, Merit and Excellence in this task
First read Nothing yet Your own rough judgement of the script, formed before you see a draft
Drafting Proposes comments and a mark against your guides Whether each suggestion is true of this student
Review Presents drafts you can edit in place Tone, accuracy, and the one next step this learner needs
Grade Suggests a grade only The grade you record and the reasoning you could defend at moderation
Release Holds everything until approval When and what the student actually sees

Frequently asked questions

Does AI-drafted feedback go to students automatically?

It should not. Any tool used for marking needs to hold drafts until you approve them. If feedback can reach a student without an explicit approval step, that is a reason not to use the tool for assessed work.

Can a tool override my professional judgement on a grade?

No. A drafted mark is a suggestion. The grade you record is the one you have decided, and it is the one you would have to explain at moderation.

Is using AI to draft feedback allowed under NCEA?

Teachers remain responsible for their own assessment judgements. Check your school's assessment policy and current NZQA guidance, and keep the rules for staff marking tools separate from the rules for what students may use.

What if the drafted feedback does not sound like me?

Load your own comment banks and edit freely. What you keep and what you delete both shape the setup, so the register moves toward yours over a term.

Do I still need to read every student's work?

Yes. Reading the script is what makes the loop trustworthy. If you are approving comments on work you have not read, you no longer have a human in the loop, only a human near it.

Get started with Jeddle

Jeddle gives teachers and students instant, syllabus-aligned feedback powered by JeddAI.

Get started with JeddAI

Looking for study material? Browse Jeddle's Australian-English subject resources, or explore more articles on Feedback & Assessment.

Shopping cart0
There are no products in the cart!