EdTech & Tools

How do I set up an AI grading assistant for the first time?

Start by writing down what good work looks like: your rubric, your success criteria, and the comments you already reuse. Set the subject and grade level, run a calibration batch of five or six pieces you have already graded yourself, then read every draft against the evidence before anything reaches a student.

What does an AI grading assistant actually do?

An AI grading assistant works in the drafting layer of your workflow. You give it the criteria you already grade against and a set of student work. It returns a first pass: a comment on each piece, usually a suggested score or performance level, and a next step for the student. None of that is a decision. It is a draft, in the same sense that a student teacher’s grading is a draft until you have read it.

That framing sets the right expectation for a first session. These tools are good at applying the same standard to the thirtieth paper as to the first, at naming what a response is missing against a stated criterion, and at turning your shorthand judgment into a sentence a student can act on. They are weak on anything they were never told: a student’s history, an accommodation, the class conversation two weeks ago that changed what the task was really asking for. You supply that context, every time.

The failure mode to plan for is not gibberish. It is fluent, plausible feedback that is slightly wrong. A comment praising evidence the student did not give. A score half a level generous. Fluent errors are harder to catch than obvious ones, which is why the review step is not optional and why your first batch should be small enough to read closely. Federal guidance on AI in schools makes the same point, recommending that a person stay in the loop on decisions that affect students.

What do you need ready before your first session?

Set up one class and one assignment. Not your whole course. Subject and grade level matter more than they look, because they set the register: feedback pitched for a Grade 7 lab report reads very differently from feedback on a Grade 11 argumentative essay. Drafts that come back too advanced or too simple are almost always a class setup problem, and that is the most common first-session frustration.

Then gather the material you already grade against. Most teachers have all of it somewhere, in a unit folder or in their head. Pulling it into one place is the real setup task.

  • Your rubric, with the performance levels you score against.
  • Success criteria: the observable things a strong response contains.
  • A comment bank of ten or so comments and next steps you reuse most.
  • The standards or learning goals this assignment reports against.
  • Five or six pieces of student work you have already graded yourself.

How do you write criteria an AI can act on?

Rubrics written for other teachers lean on shared understanding. Levels separated by adverbs, thoroughly explains against adequately explains against partially explains, work fine between colleagues who have normed together, because you all roughly know where the line falls. Software has no staffroom. Given only adverbs it will guess, and in practice it guesses generously.

Rewrite the distinctions as things that are either present in the work or not. Instead of uses evidence effectively, try quotes at least two sources and explains in their own words how each supports the claim. The same specificity that makes study guides useful to students makes criteria usable by software, because a reader can check them against the page. This is the most valuable half hour in the whole process, and it sharpens your own grading at the same time.

Comment banks do the other half of the job. Load the comments you actually reuse, in your own wording, including the next steps you give most often. Without them a tool fills the silence with generic praise. With them, drafts come back sounding closer to you and the edits get shorter.

  • Name the observable move, not the quality judgment.
  • Say how many, how long, or how often wherever you can.
  • Describe each level boundary by what appears in the work, not by degree adverbs.
  • Write next steps a student could act on before the next lesson.
  • Keep your own phrasing. A bank is for your voice, not a house style.

How do you run a calibration batch?

Run five or six pieces first, chosen deliberately: one you consider strong, two in the middle, one weak, and one you found hard to place. Grade them yourself first, or at least write down the level you expect. Without your own judgment recorded in advance you will read the drafts and quietly agree with them, which tells you nothing.

Then compare on three questions. Does it land on the level you gave? Does it cite evidence that is actually in the work? Is the next step one you would have written? Consistent drift, such as every draft sitting half a level above yours, is a criteria problem, and you fix it by sharpening the boundary language. A single odd draft among six is noise. Edit it and move on.

If several drafts miss the same thing, that points at a vague success criterion or a missing comment-bank entry. Every adjustment you make carries forward to every batch after it, which is why the first session is worth doing slowly and the tenth is not.

What should you check before feedback reaches a student?

Read every draft before release. Read it the way you would read expert essay grading handed to you by a colleague: does the comment match what is on the page, and would you defend the level to a parent? Four checks cover most of it, and they take less time than they sound.

Evidence first, because it is the check people skip. If a comment refers to a quotation, a method, or a paragraph, confirm it is really there. Then calibration: is the level right, and right for the same reason you would give? Then tone, because a student who has just failed twice needs a different opening from one who is coasting. Then the next step, which should be one thing, specific, and doable before the next task.

Once your criteria are tight, most edits are small, and that is where the time saving comes from. Measure it honestly. The number that matters is minutes per piece including your review, not minutes saved on generation. Compare your first batch against your fourth before you decide whether any of this has earned a place in your week.

How do you turn one session into a routine?

Keep the sequence the same every time: set criteria, upload, generate, review, deliver. Consistency is the point. The same standard applied the same way is what makes grades comparable across a class and across a semester, and it is the first thing to slip at 10pm on a Sunday with twenty-eight papers left.

Widen the batch only once the drafts hold up at the size you are already running. Revisit criteria between units rather than partway through a batch, so your standard stays stable while a class set is being graded. Once a semester, norm with a colleague: take a few pieces you graded with the tool and a few you graded by hand, and check whether the levels agree. Disagreement is useful either way.

Two things are easier to settle before this becomes routine than after. Check your district or school policy on uploading student work to an outside service, and confirm what the provider does with that work. Then decide what you will tell students and families about which part of the process is machine-drafted. Both conversations go better with one class than after a semester of grading nobody knew about.

Where does a tool like this fit into the week?

Everything above holds whatever software you use, and it is worth trying more than one. Jeddle is built around that sequence: you load your rubric, success criteria, and comment banks, set the class and grade level, and its engine drafts feedback and a suggested level for each piece, which you edit on a review screen before anything reaches a student.

The features worth checking in any product you trial are the same four. Can you bring your own rubric instead of picking from presets? Are the comment banks yours, in your wording? Is the review step a real editing surface rather than an approve button? And does it show you what in the student’s work drove each comment, so the evidence check takes seconds instead of a re-read? Jeddle anchors comments to the passage they refer to for that reason. If a tool fails those last two, the time it saves on generation comes straight back on review.

A fair trial of any of them looks the same: one class, one assignment, thirty minutes writing criteria, six pieces you have already graded, and an honest count of the minutes. That is enough to tell whether JeddAI, or anything else, has earned a place in your routine.

Reading your calibration batch
What you notice across the batch What it usually means What to change
Every draft sits about half a level above yours Level boundaries are described in adverbs, so the model resolves ambiguity upward Rewrite the top two levels as features that are present or absent
A comment praises evidence that is not in the work The model filled a gap with plausible text Name the required evidence in the criterion, and check every quotation on review
Comments are accurate but generic No comment bank, or one that is too short Add ten of your own most-used strengths and next steps
Feedback is pitched too high or too low for the class Subject or grade level is set wrong Fix the class setup before running another batch
One draft is odd, the other five are fine Noise, not a pattern Edit it and move on

Frequently asked questions

Do I need a finished rubric before I start?

No. A lightweight rubric or your existing success criteria is enough to begin, and you can sharpen it after seeing the first batch. Be aware that vague level descriptors are the single biggest cause of generous, generic drafts.

Does an AI assistant set the final grade?

It should not. Treat every score and comment as a draft you review, edit, and approve. Keeping a teacher in the loop on decisions that affect students is also what federal guidance on AI in schools recommends.

How much student work should be in the first batch?

Five or six pieces, spread across the ability range, and ideally ones you have already graded yourself so you have something to compare the drafts against.

Will the feedback sound like mine?

Only if you give it your language. Comment banks built from the phrases and next steps you already use are what pull drafts toward your voice. Without them you get feedback that is correct but generic.

Does this work for standards-based grading?

Yes, if your rubric language is aligned to the standards or learning goals you report against. Write the level descriptors in the same terms that appear on the report, and the drafted comments will reference those expectations.

Get started with Jeddle

Jeddle gives teachers and students instant, syllabus-aligned feedback powered by JeddAI.

Get started with JeddAI

Looking for study material? Browse Jeddle's Australian-English subject resources, or explore more articles on EdTech & Tools.

Shopping cart0
There are no products in the cart!