Standards Based Grading Gradebook: How to Build One That Survives a Semester

Title card for a standards based grading gradebook guide listing four points: rows are standards, five decisions, mixed research, and that grading alone does not teach

A standards based grading gradebook stores evidence by learning target instead of by assignment. Each row is a standard. Each entry is a piece of evidence a student produced for that standard. That one structural change is what makes a gradebook readable to a student, a parent, and the teacher who inherits the class next year. It is also what creates every hard decision you will spend the semester making.

Most of what gets written about standards-based grading is about philosophy. This is not that. This page is about the file — what goes in the columns, what happens when a kid takes a quiz three times, what you do when the office needs a letter grade by Friday, and what the research says happens when you get those choices wrong.

I will say up front that the evidence here is thinner and more mixed than the professional development usually admits. That is covered honestly below, because a teacher about to rebuild a semester of records deserves to know what they are buying.

Key Takeaways

  • The structure is the whole idea. Rows are standards, not assignments. If your gradebook still reads as a chronological list of tasks with points attached, you have changed the labels and not the system.
  • One randomized trial found a real gain. A cluster RCT across 29 schools of a proficiency-plus-reassessment program in ninth-grade algebra and geometry reported a 0.33 standard deviation improvement on end-of-course mathematics tests — roughly the 50th to the 63rd percentile.
  • But no grading system improves learning by itself. Guskey and Link put it plainly: grading does not change curriculum or instruction, and there is no evidence that standards-based grading on its own raises achievement. What it reliably does is make the grade mean something specific.
  • Taking practice work out of the grade has a measurable cost. A two-year study of a policy change in secondary mathematics found that when practice stopped counting, completion of practice fell and achievement declined on several standards.
  • Students object on fairness grounds, and their objections are specific. In a survey of 478 high school students, the complaints were not about rigor. They were about inconsistent reassessment rules between teachers and about losing credit for work they had done.
  • The five decisions below are where implementations live or die: how many standards, what scale, how repeated attempts combine, what counts as evidence, and where behavior goes.

What is a standards based grading gradebook?

It is a gradebook whose primary axis is the standard rather than the assignment. A traditional gradebook answers “what did this student turn in and what was it worth?” A standards-based gradebook answers “what can this student do, and how do we know?”

In practice that means a column per learning target instead of a column per task. A single test might feed six different columns, because a single test usually assesses six different things. Three separate assignments over a month might all feed one column, because they were all evidence for the same skill. The assignment does not disappear — it becomes the container for the evidence rather than the thing being measured.

If you want the underlying rationale and the arguments for and against the approach as a whole, that is covered in the main standards-based grading guide. This page assumes you have already decided, or been told, and now have to build the thing.

Does a standards based grading gradebook actually improve learning?

The honest answer: one good randomized trial says a specific version of it helped, a large review says grading systems do not improve learning on their own, and at least one comparison found worse outcomes. Anyone who tells you the evidence is settled has not read it.

The positive result

The strongest evidence for this approach comes from a cluster randomized controlled trial of the PARLO program run by Kramer and colleagues and published in the Journal of Research on Educational Effectiveness. Twenty-nine schools were randomized — 14 treatment, 15 control — with the intervention in ninth-grade algebra and geometry classrooms. Students were rated not-yet-proficient, proficient, or high-performance on each learning outcome, and could reassess any outcome for full credit after further study. The program produced a 0.33 standard deviation improvement on end-of-course mathematics tests, moving the average student from roughly the 50th to the 63rd percentile. The effect was strongest among students who were already more motivated.

Two caveats on that one. First, PARLO is a defined program with mastery definitions and reassessment built in, not a gradebook layout — the trial tested the whole package. Second, the full published article was not accessible to me, so the figures above come from the study summary published by the Society for Research on Educational Effectiveness rather than from the paper itself. I am citing what I actually read.

The sobering results

Laura Link and Thomas Guskey, writing in Theory Into Practice, argue that standards-based grading is best understood as a communication tool and not as an achievement intervention. Their sentence is the one worth keeping: no grading system by itself improves student learning, because grading does not alter curriculum or instruction. They note that the practice remains largely unstudied, that emerging research shows it produces a stronger relationship between grades and external measures, and that no evidence indicates it improves achievement.

Four research summaries on standards-based grading: Kramer 2024 cluster randomized trial, Link and Guskey 2022, Townsley and Varga 2017, and Huey 2022
Five peer-reviewed sources, and they do not agree with each other.

Matt Townsley and Matt Varga compared two demographically similar Midwestern high schools of about 450 students each, one using standards-based grading and one traditional, across the graduating classes of 2015 and 2016. They found no significant difference in math, English, or cumulative GPA — and ACT scores roughly 2.2 to 2.7 points higher at the traditional school. That is a quasi-experimental comparison of two schools with all the confounding that implies, and the authors flag the rural, low-diversity sample themselves. It is not proof that the gradebook caused it. It is also not a result anyone should skip past.

The broadest context comes from A Century of Grading Research in the Review of Educational Research, an eight-author review of the whole field. Two findings from it matter for anyone building a gradebook. Teachers mix achievement with work habits whether or not their system says they should — even teachers who report supporting standards-based reform report using practices that blend effort and improvement into academic marks. And grades, messy as they are, consistently predict educational persistence and completion better than standardized tests do. The multidimensionality that standards-based grading tries to strip out may be part of what makes a grade useful.

The five decisions every standards-based gradebook has to make

Make these five on purpose, write them down, and do not change them midyear. Every failed implementation I have read about failed at one of these, usually by leaving it to be decided case by case.

DecisionThe optionsWhat to weigh
1. GranularityEvery state standard, or 8–15 bundled targets per courseToo many rows and nobody reads the gradebook. Too few and the grade stops being diagnostic. Most workable secondary courses land between eight and fifteen per semester.
2. Scale4-point, 3-level proficiency, or letter equivalentsGuskey argues for a limited number of performance categories rather than percentage scales, to cut down on false precision. The 1–4 scale and its conversion problems are worth reading before you commit.
3. How attempts combineMean, most recent, highest, or professional judgmentThis is the one students notice. Averaging punishes early struggle. “Most recent” can drop a student for one bad day. Write the rule down and publish it.
4. What counts as evidenceAssessments only, or assessments plus observed performanceNarrow it too far and you are grading test-taking. Widen it without criteria and you are back to impressions.
5. Where behavior goesA separate reported strand, or nowhereSeparating academic marks from work habits is one of the three criteria Guskey names. It only works if the second strand is actually reported, not deleted.

On decision three specifically: students in one large survey complained that a reassessment score replaced a higher earlier score even when it was worse. That is a defensible rule and an indefensible surprise. The rule is not the problem; discovering it after the fact is.

Five numbered gradebook design decisions: granularity, scale, how attempts combine, what counts as evidence, and where behavior is reported
Decide these on purpose, write them down, and hold them all year.

Should practice work count in the grade?

The evidence says be careful, and it points the opposite direction from the standard advice.

The usual argument is that practice is formative, so it should not be graded — students will do it because it helps them. A two-year study by Huey and colleagues in Studies in Educational Evaluation tested exactly that policy change with 122 eighth graders and 123 ninth graders in year-long geometry courses. When practice work was removed from grade computations, completion of practice work decreased overall, and student achievement declined on several of the standards being assessed. The authors concluded that process grades are likely essential to maintaining achievement in that population, and noted their result contradicts reform advocates who claim engagement holds regardless.

That is one study of high-achieving students in one subject, and it should not be read as a verdict on the whole question. But it is a direct test of a specific gradebook decision, and it came out against the conventional answer. If you take practice out of the grade, put something else in its place that makes the practice visibly matter, and watch your completion rates rather than assuming.

Students see the same thing from the other side. In a study of 478 students at one high school, the most common complaint about standards-based grading was that homework should count — not because they wanted easy points, though some did, but because with few graded items the whole grade rested on two or three assessments. “There are so few points available on quizzes and tests” is a fairness objection, and it is a reasonable one.

How do you convert to a letter grade when the office needs one?

Publish the conversion rule before the first grade goes in, and use the same one every time.

Most secondary teachers running this system do not control the report card, so the standards-based gradebook feeds a traditional one. That translation is where trust gets lost, because a student who sees a row of 3s and receives a B wants to know how. Three workable approaches:

  1. A published crosswalk. A fixed table — all 4s and 3s with no 2s is an A, and so on. Simple, transparent, and slightly blunt at the boundaries.
  2. A proficiency threshold. The grade is determined by the proportion of standards at proficient or above, with a stated floor of standards that must be met regardless.
  3. A weighted composite with named priority standards. More precise, harder to explain, and worth it only if you will actually explain it.

Whichever you choose, the parent-facing version has to be one paragraph long. Brookhart and colleagues found that many parents attempt to interpret standards-based labels by translating them back into letter grades regardless of what the school intends. Give them the translation rather than letting them invent one.

What breaks, and when

Reassessment load, teacher-to-teacher inconsistency, and midyear rule changes — in that order.

  • Unlimited reassessment. Guskey and Link warn directly that unlimited retakes create burdensome workloads, and that retaking a poorly aligned assessment does not produce learning regardless of how many attempts you allow. Cap attempts, require evidence of new study before a retake, and put reassessment in a fixed window rather than on demand.
  • Different rules in different rooms. Students in the 478-student survey said it repeatedly: nobody communicates the same, there is no standardization, some teachers appear to make reassessment deliberately difficult. Within a department, the five decisions above should have one answer, not six.
  • Changing the rule in January. A grading rule changed midyear is functionally retroactive, because the evidence already in the book was produced under the old one. If a decision is wrong, finish the term and change it at the semester line.
  • A gradebook nobody opens. Fifteen standards with one entry each in November is not a standards-based gradebook, it is a spreadsheet. The system only pays off if there is enough evidence per row to make the row mean something.
Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.

Before you do this alone: what it costs beyond your own room

A standards-based gradebook inside one classroom in a traditional building is a bigger undertaking than it looks, and four things will bite you.

  • The transcript is not yours. Whatever you do internally, a letter grade and a GPA leave your room and go on a document that affects class rank, scholarship eligibility, and athletic eligibility. Confirm with your counseling office how your marks will be recorded before you change how you produce them, not in May.
  • Your software may not support it. Most district gradebooks are built around assignments with point values. Some can be bent into standards; some can only fake it with categories. Find out which you have before you design a system it cannot hold, because a parallel spreadsheet you maintain by hand is a system that dies in November.
  • Students transfer. A student who arrives in week twelve arrives with points, and a student who leaves takes your proficiency levels into a building that will not know what they mean. Have a stated rule for both directions.
  • Reassessment is the real labor cost. Be concrete about it before you promise anything. If 20% of 140 students reassess one target a week, that is 28 additional items to write, administer, and score every week, on top of everything else. That number is why reassessment windows and attempt caps exist, and it is why Guskey and Link warn that unlimited retakes become burdensome.

There is an instructional point buried in that last one that a coach would raise immediately. Reassessment only works if you have more than one assessment item per target that genuinely measures the same thing. If the retake is the same quiz, you are measuring memory of a quiz. If the retake is harder, you are punishing the student for needing it. Building a second and third form for each priority standard is the unglamorous prerequisite, and it is most of the actual work of this system.

Four failure points for a standards-based gradebook: unlimited reassessment workload, inconsistent rules between teachers, midyear rule changes, and too little evidence per standard
In roughly this order, and usually by midyear.

None of this is an argument against doing it. It is an argument for doing it with the department and the counseling office rather than in private, and for starting with one course rather than a schedule.

A setup checklist you can work through in an afternoon

  1. List the course’s learning targets and bundle them down to 8–15 per semester. Write each one in student-facing language. If you cannot say it in a sentence a fifteen-year-old understands, it is still a standards document and not yet a target.
  2. Pick the scale and write the descriptors. Not just the numbers — what a 3 actually looks like in this course. Rubric templates are the fastest way in.
  3. Write the five decisions on one page. Granularity, scale, how attempts combine, what counts as evidence, where behavior goes. Date it.
  4. Build the columns before the semester, not during it. Retrofitting a gradebook in week six means hand-remapping every entry already in it.
  5. Decide the reassessment window and the entry requirement now. Which days, how many attempts, what a student must show to earn one.
  6. Write the one-paragraph explanation for families. What the numbers mean, how they become a letter, and when. Send it in the first two weeks rather than after the first complaint — the same front-loaded contact that works for everything else works here.
  7. Pick your check date. A day in week eight when you open the gradebook and ask whether any row has too little evidence to defend, and whether your conversion rule still produces grades you believe.

What to tell students, and what to tell families

Tell students the rules before the first entry, and show them the gradebook. Most of the recorded student resistance to this system is not resistance to rigor. It is the experience of being graded by a machine whose rules nobody explained.

Grades 6–12 students can handle the actual explanation, and they deserve it. Say what the scale means. Say how repeated attempts combine, including the case where a retake scores lower. Say which standards carry the most weight. Say what happens to homework. Then hold to it, because the fastest way to lose a room under this system is to make an exception for one student and get caught.

For families, lead with the translation and the one thing that changed. If your building is also making the shift, the arguments against are worth knowing as well as the arguments for — the case that standards-based grading does not work is an honest summary of the objections, and a teacher who can state the objection fairly is more persuasive than one who cannot.

What to do next

Before you rebuild anything, write the five decisions on one page and show it to one colleague who teaches the same course. Most of the damage in a standards based grading gradebook happens because a decision was never made explicitly — it just emerged from whatever the software defaulted to and whatever felt fair in the moment.

Then build one unit, not one semester. Run it, look at how many entries each row actually accumulated, and see whether your conversion rule produced a grade you would defend to a parent. Adjust once, at the unit boundary. That is a slower start than a summer rebuild and it is the version that is still running in April. When the unit is done, a short structured reflection on what the rows told you is worth more than another round of reorganizing the columns.

Frequently Asked Questions

What is a standards based grading gradebook?

A gradebook organized by learning target rather than by assignment. Each row is a standard and each entry is evidence a student produced for that standard, so a single test can feed six different rows and three assignments over a month can feed one. The practical test is whether the gradebook answers “what can this student do” rather than “what did this student turn in.” If the rows are still tasks with point values, the labels have changed and the system has not.

How many standards should be in the gradebook?

For most secondary courses, eight to fifteen bundled targets per semester. Listing every state standard produces a gradebook nobody opens, including you. Bundling too aggressively produces a grade that is no more diagnostic than a single letter. The working check is whether each row will accumulate enough evidence by the end of the term to defend a judgment — a row with one entry in November is not measuring anything.

Should homework and practice work count toward the grade?

The evidence is more cautious than the usual advice. A two-year study of a policy change in secondary mathematics found that when practice work was removed from grade computations, practice completion fell and achievement declined on several standards. Students in a separate survey said the same thing from their side: with few graded items, removing practice makes the entire grade rest on two or three assessments. If you take practice out, replace it with something that makes practice visibly matter and watch your completion rates.

How should multiple attempts at the same standard be combined?

Pick one rule, publish it before the first entry, and do not change it midyear. Averaging punishes early struggle, which is the thing the system is supposed to fix. “Most recent” can drop a student for one bad day. Highest score removes the incentive to prepare. Many teachers use most-recent with professional judgment for outliers. Whichever you choose, students need to know in advance what happens when a retake scores lower than the original — that specific surprise is one of the most common student complaints on record.

Does standards-based grading raise test scores?

The evidence is mixed and should be described that way. A cluster randomized trial across 29 schools of a proficiency-plus-reassessment program in ninth-grade math reported a 0.33 standard deviation gain on end-of-course tests. A comparison of two Midwestern high schools found no GPA difference and ACT scores roughly 2.2 to 2.7 points higher at the traditional school. Guskey and Link argue no grading system improves learning on its own, because grading changes neither curriculum nor instruction. What it reliably changes is what the grade means.

How do you turn standards-based marks into a letter grade?

With a published conversion rule applied identically every time. The three workable approaches are a fixed crosswalk table, a proficiency threshold based on the proportion of standards met, and a weighted composite with named priority standards. Write the parent-facing version in one paragraph and send it in the first two weeks. Research on grading found that many parents translate standards-based labels back into letter grades regardless of what the school intends, so it is better to hand them the translation than to let them invent one.

What is the biggest implementation mistake?

Leaving the rules implicit and letting them emerge from whatever the software defaults to. The five decisions — granularity, scale, how attempts combine, what counts as evidence, and where behavior goes — need explicit written answers that are the same across a department. Students surveyed about standards-based grading complained most about inconsistency between teachers, not about difficulty. Unlimited reassessment is a close second; it produces a workload that collapses by midyear.

Can one teacher run this if the rest of the school does not?

Yes, with three conversations first. Confirm with the counseling office how your marks will be recorded on the transcript, since class rank, scholarships, and athletic eligibility depend on a document you do not control. Confirm that your district gradebook can actually hold standards rather than categories pretending to be standards. And agree on a rule for students who transfer in or out mid-term. Start with one course rather than a full schedule.

Sources

  • Link, Laura J., and Thomas R. Guskey. “Is Standards-Based Grading Effective?” Theory Into Practice, 2022. Full text (PDF).
  • Brookhart, Susan M., Thomas R. Guskey, Alex J. Bowers, James H. McMillan, Jeffrey K. Smith, Lisa F. Smith, Michael T. Stevens, and Megan E. Welsh. “A Century of Grading Research: Meaning and Value in the Most Common Educational Measure.” Review of Educational Research, vol. 86, no. 4, 2016, pp. 803–848. Full text (PDF).
  • Townsley, Matt, and Matthew Varga. “Getting High School Students Ready for College: A Quantitative Study of Standards-Based Grading Practices.” Journal of Research in Education, vol. 28, no. 1, 2017, pp. 93–112. Full text (PDF).
  • Huey, Maryann E., Peter R. Silvey, Amanda G. Vaughan, and Anne L. Fisher. “Assessing the Impact of Standards-Based Grading Policy Changes on Student Performance and Practice Work Completion in Secondary Mathematics.” Studies in Educational Evaluation, vol. 75, 2022, article 101211. Journal listing.
  • Peters, Randal, Jerrid Kruse, Tom Buckmiller, and Matt Townsley. “‘It’s Just Not Fair!’ Making Sense of Secondary Students’ Resistance to a Standards-Based Grading.” American Secondary Education, vol. 45, no. 3, 2017, pp. 9–28. Full text (PDF).
  • Kramer, Karen, Jill Posner, Alexander Browman, Rebecca Lawrence, Jennifer Roem, and Kathryn Krier. “The Impact of a Standards-Based Grading Intervention on Ninth Graders’ Mathematics Learning.” Journal of Research on Educational Effectiveness, 2024. Figures cited here come from the Society for Research on Educational Effectiveness study summary, which is what was read; the full article was not accessible.

A note on the evidence. Five of the six sources above are peer-reviewed, and they do not agree with each other. That is the accurate state of this field, and a gradebook decision made on the assumption that the research is settled is a decision made on something that is not true.


About Clay Shumate

By Clay Shumate — Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full.

Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.