Design a rubric by naming one observable skill per criterion, tying each to a grade-level standard, and writing performance-level descriptors that point at evidence rather than quality. Keep the levels non-overlapping so two graders cannot both defend different scores. Then calibrate: score the same anchor papers as a team and reconcile the gaps before grading live.
What actually makes two graders give the same essay the same score?
Two graders land on the same score when the rubric tells them exactly what to look at and what counts as evidence of each level. Most disagreement between colleagues is not a difference of taste. It traces back to three fixable faults in the document: criteria that overlap, so one strength gets counted twice; levels written in comparative adjectives instead of observable evidence; and a team that has never scored the same paper side by side.
A rubric divides an assignment into component parts and describes what work looks like at each level of mastery. Carnegie Mellon’s Eberly Center notes that rubrics are especially valuable in courses with multiple graders, because they help ensure consistency across graders and reduce the systematic bias that can appear between them. The same holds for a Grade 9 English team of five. The rubric is the only piece of professional judgment all five share.
So treat consistency as a design property of the document, not a quality of the teachers using it. If your department argues about scores every October, the argument is usually already written into the wording of level 3.
How do I tie each criterion to a standard without turning the rubric into a checklist?
Start from the standards your district has adopted rather than from a list of adjectives. In Common Core states, the College and Career Readiness anchor standards form the backbone of the ELA and literacy standards, and the grade-specific standards supply the detail for each grade. That split maps onto rubric design: the anchor standard gives a criterion its name, and the grade-level standard sets what the top band demands of a Grade 8 writer as against a Grade 11 one.
Hold to one standard per criterion. A row that cites three standards is really three criteria wearing one label, and each grader will quietly weight whichever of the three they care about most.
Put the code beside the criterion name and keep the descriptor in plain language. The code is what you show a parent or an administrator who asks how a score was reached. The descriptor is what a student uses to revise. A row written as a citation does neither job well.
How do I keep criteria from overlapping?
Overlap is the most common defect in a homemade rubric, and it is easy to test for. Take one sentence of student work and ask how many scores it can move. If a sharp thesis statement lifts the score for argument, for organization, and for analysis, those three rows are measuring one decision, and the student is rewarded three times for it. That is how a rubric quietly turns a two-point gap into a six-point gap.
Published rubrics are disciplined about this. The College Board’s scoring rubrics for the AP English Language and Composition free-response questions, effective Fall 2019, score each essay out of six points across three separate rows: Row A for thesis (0 to 1 point), Row B for evidence and commentary (0 to 4 points), and Row C for sophistication (0 to 1 point). A defensible thesis earns its point whether or not the evidence underneath it holds up. The weighting lives in the point spread, not in extra rows.
Fewer criteria helps too. Four to six rows are enough for most extended writing tasks in Grades 7 to 12. Any two rows that always move together are a sign you have split one skill in half.
How should I write performance-level descriptors?
Describe evidence, not quality. Analysis is thorough asks every grader to supply a private definition of thorough. Explains how each piece of evidence supports the claim it is attached to points at something two graders can find, or fail to find, in the same paragraph. If you cannot underline the proof in a real student paper, the descriptor is not finished.
Keep the levels parallel. Down a single row, the same dimension should vary from top to bottom: how consistently the behavior appears, how much of the text it covers, how independently the student manages it. Trouble starts when level 4 talks about evidence and level 2 talks about effort. Those are two scales stacked in one column, and each grader will apply whichever one the paper in front of them satisfies.
Two habits to avoid. Do not use counting as a proxy for quality, because a rule like three pieces of evidence scores a list higher than an argument. And do not write the bottom levels only as absences, which leaves the student with no next step and the grader with no positive evidence to look for.
Three to five levels per criterion is usually the honest range. If you cannot write a findable difference between level 5 and level 6, that boundary gets settled by mood, and two colleagues will settle it differently.
How does a department calibrate so its scores actually match?
Calibration is a meeting, not a document, and it is the step most teams skip. National assessment programs treat it as unskippable. NAEP trains its scorers with anchor sets containing three or four clear examples of each score category, each shown with its score and a written annotation explaining the reasoning, and scorers keep those papers beside them as they work.
Run a small version of that. Pull six to eight real papers spanning the range, have every teacher score them independently, then compare row by row. Where two teachers differ by more than one level, the fault is almost always in the descriptor rather than in the paper, so edit the wording in the room while the disagreement is fresh. Keep the agreed papers and their annotations as next year’s anchor set.
Set a realistic bar. NAEP’s own quality checks treat roughly 60 percent exact agreement as acceptable on complex six-point writing items, and look for a Cohen’s kappa above .6 on four- to six-point items. Those are trained scorers working full time on a single prompt. A department that demands perfect agreement will get polite silence instead of the disagreements that improve the rubric.
Re-anchor at the start of a long grading run and again halfway through. Grading against an explicit, descriptive set of criteria also stops one teacher’s standard from drifting between the first stack and the last, so this pays off even for a teacher grading alone.
What should I check before the rubric grades anything that counts?
Dry-run it on three papers you have already graded: one strong, one middling, one struggling. You are checking whether the rubric reproduces judgments you already trust. Where it does not, you find out which row is wrong before it touches a report card.
Publish the student version at the same time. The same descriptors rewritten in the second person work as revision checklists, as the criteria for peer review, and as study guides for the next task of that type. Students who can see the difference between level 3 and level 4 in concrete terms will write toward it.
- Can a genuinely excellent paper reach the top level, or is the ceiling written so tightly that nobody gets there?
- Does an honest weak attempt score above zero, so the bottom of the scale still separates students?
- Do the point values add to the total you told students, and does the weighting match what you actually care about?
- Would a substitute grader, given only your rubric and your anchor papers, arrive at the same score?
Where does software fit into a rubric built this way?
Nothing automated repairs a vague rubric. A tool applies the criteria and level descriptors you give it, so good analysis produces inconsistent machine scores for the same reason it produces inconsistent human ones. The design work above is what makes any assistance worth having.
With a rubric that has survived calibration, a grading platform can apply it evenly across a whole class and draft comments tied to the row that earned them. Jeddle works this way: you enter your own criteria, levels, descriptors and point values, and JeddAI drafts against exactly those without rewriting your wording or reassigning your points. The teacher reviews and approves every score before a student sees it. If you want to see what expert essay grading looks like when it is driven by your own rubric, that is the place to start.
| Vague descriptor | Observable rewrite |
|---|---|
| Uses evidence effectively | Each quotation is followed by an explanation of how it supports the claim in that paragraph |
| Well organized | Every body paragraph opens with a claim that connects to the thesis, and transitions name the relationship between paragraphs |
| Shows strong understanding of the text | References at least two moments from different parts of the text and reads them consistently with the whole |
| Few errors | Sentence errors do not obscure meaning at any point in the response |
Frequently asked questions
How many criteria should a grading rubric have?
As few as clearly capture what you are grading. Four to six criteria for an extended piece of writing is a workable range. A short rubric with well-defined rows scores more consistently than a long one with overlapping or vague categories, because every extra row is another place two graders can diverge.
Should I use an analytic rubric or a single holistic scale?
Both work, and the choice depends on what you owe the student. A holistic scale is faster and fine when one overall judgment is what gets reported. Analytic rows tell a student which skill to work on next, and they give a team something specific to argue about during calibration, which is how the wording improves.
What do I do when two teachers score the same paper differently?
Reconcile the descriptor, not the paper. Ask both teachers to point at the text that earned their level. If they are looking at different evidence, the criterion is doing two jobs. If they are looking at the same evidence and reading it differently, the boundary between levels needs sharper wording. Fix it before the next batch.
Do I have to rebuild the rubric when our standards are updated?
Usually not. Standards updates tend to change the codes and the expectations at the top of each grade band, so the work is re-checking the standard reference on each criterion and the ceiling descriptor. The evidence you look for in a paragraph rarely changes as much as the paperwork around it.
Can I change a rubric after I have already graded with it?
Yes, and you should, as long as you do not re-score work already returned under the old version. Note which rows you kept overriding, sharpen those descriptors, and apply the improved rubric to the next task. A rubric that never changes is one nobody has tested against real student work.
Get started with Jeddle
Jeddle gives teachers and students instant, syllabus-aligned feedback powered by JeddAI.
Looking for study material? Browse Jeddle's Australian-English subject resources, or explore more articles on Feedback & Assessment.



