Tag: education research

  • Test Corrections: How to Make Them Teach Instead of Hand Back Points

    Test Corrections: How to Make Them Teach Instead of Hand Back Points

    By Clay Shumate

    Test corrections are a structured second pass in which a student identifies why an answer was wrong and produces a correct one. The research on learning from errors is strong and supports the practice. The research on what students actually do with error feedback is much less flattering, and it is the part that determines whether your version works.

    Done well, a test correction is the most efficient reteaching you will ever get: the student already knows what they missed. Done badly, it is twenty minutes of copying answers off a neighbor’s paper for half the points back.

    Key Takeaways

    • Making errors and then correcting them beats avoiding errors. A review in the Annual Review of Psychology concluded that error avoidance “appears to be the rule in American classrooms” and that errorful learning followed by corrective feedback produces better retention.
    • Feedback is not optional — it is the whole mechanism. In one lab study, errors were corrected on a later test about 70% of the time with feedback and roughly 4% of the time without it.
    • Your most confident wrong answers are the ones most likely to get fixed. That is the hypercorrection effect, and it means the student who argues with you about question 14 is the student most likely to remember the right answer in May.
    • When corrections are optional, most students skip them. Across 20,058 assessments from 2,826 students in grades 5–11, students opened the detailed error feedback in only 44% of cases — and the students with the lowest scores were the least likely to look.
    • A reflection form on its own does nothing. A randomized study of “exam wrappers” found no effect on exam scores, final grades, or measured metacognition. The structure has to force the thinking, not just ask for it.

    Free Download · PDF

    The Test Correction Sheet and Policy Card (2 pages)

    Page 1 is the student sheet — three blocks of the four-box correction, with the reasoning box that makes copying an answer useless. Page 2 is a policy card to fill in and staple to the first test of the term, plus a sorting sheet for reading a class set in five minutes.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Two-column graphic on test corrections evidence: 70 percent versus 4 percent correction rates with and without feedback and roughly 82 percent for high-confidence errors on one side, a 44 percent feedback open rate and a median of one error reviewed on the other
    The case for corrections and the complication, side by side.

    What Are Test Corrections?

    A test correction is a required second pass over a returned assessment in which the student states what the right answer is and why their first one was wrong. That second clause is the whole practice. Without it you have a transcription exercise.

    Two versions run in American schools under the same name, and they produce opposite results.

    Version one: hand back the test, let students fix wrong answers, give half credit back. Most look at the answer key, write the right letter, hand it in. Nobody learns anything, the grade goes up, and the next test looks exactly like the last one.

    Version two: the student has to name the error — not the answer, the error. What did I think was true that was not? That version is slower, harder to grade, and is the one with research behind it.

    If you only take one thing from this article: the credit is not the intervention. The credit is the thing that gets students to do the intervention. Confusing the two is how a good practice turns into grade inflation with extra steps.

    Does the Research Actually Support Learning From Mistakes?

    Yes, and more strongly than most teachers assume. Janet Metcalfe’s review “Learning from Errors,” published in the Annual Review of Psychology in 2017, surveys the laboratory evidence and reaches a blunt conclusion: error avoidance “appears to be the rule in American classrooms,” and it is the wrong rule. Errorful learning followed by corrective feedback produces better retention than carefully steering students around mistakes.

    Her recommendation is that teachers should “allow and even encourage students to commit and correct errors while they are in low-stakes learning situations rather than to assiduously avoid errors at all costs,” precisely because the goal is performance later, when the stakes are high.

    The crucial qualifier is the phrase followed by corrective feedback. Metcalfe is explicit that the feedback, “including analysis of the reasoning leading up to the mistake,” is what makes the error productive. A wrong answer left wrong is just a wrong answer. Same logic that makes formative assessment worth the class time: information is only worth collecting if something happens next.

    Why Your Most Confident Students Gain the Most

    Students correct the errors they were most sure about more reliably than the ones they guessed at. That is the hypercorrection effect, and it is counterintuitive enough that it is worth stating twice.

    Janet Metcalfe and Bridgid Finn tested it directly in a 2011 paper in the Journal of Experimental Psychology: Learning, Memory, and Cognition. Participants answered general-knowledge questions, rated their confidence, received corrective feedback on their errors, and were tested again later. Errors held with high confidence were corrected on the final test around 82% of the time. Overall recall after feedback was about 70%. Without feedback, correction rates fell to roughly 4%.

    Two things follow. First, the gap between 70% and 4% is the clearest number in this entire literature, and it says the thing teachers most need to hear: returning a graded test without a correction process is close to doing nothing. Second, the student who comes up after class and argues that question 14 was unfair is not being difficult. High confidence plus a wrong answer is the configuration most likely to produce durable learning. Argue back, with evidence.

    Stated honestly: these were college undergraduates, in a lab, answering trivia. Samples ran from 25 to 45 people per experiment. The mechanism is about confidence and attention rather than age, so it is reasonable to expect it to transfer to a fifteen-year-old and a unit test. The exact percentages are not yours to quote as classroom results.

    Four numbered boxes of a test correction form: what I put, why I put it, the right answer and where it came from, and what would make me miss this again
    Box 2 is the one that makes copying an answer off a neighbour useless.

    So Why Don’t Test Corrections Always Work?

    Because when looking at the feedback is optional, most students don’t — and the ones who need it most are the least likely to. This is the finding that should change what you build, and it comes from the largest secondary-school sample in this article.

    Ulrich Maier and Christian Klotz analyzed log data from a digital formative assessment system used in German schools, published in Contemporary Educational Psychology in 2025. The scale is unusual: 2,826 students across 182 secondary classrooms in grades 5 through 11, covering 20,058 formative assessment cases collected between 2020 and 2024. After each assessment the system offered a detailed error feedback page explaining what went wrong.

    What students did with it:

    • They opened the error feedback page in only 44% of cases. In the majority of assessments, nobody looked at the explanation at all.
    • Among those who opened it, the median number of error items reviewed was one. The average was 1.74, roughly half of the errors available to them.
    • Prior knowledge was by far the strongest predictor of whether a student looked. Higher scorers were dramatically more likely to seek feedback. The authors state the paradox plainly: low-achieving students tend to ignore elaborated feedback, despite being the group most likely to benefit from it.
    • Students who thought they had passed were more likely to open the feedback than students who thought they had failed — even when the confident ones had actually failed.

    That is a digital platform rather than a paper test, and German secondary schools rather than American ones, so the exact rates will not be yours. But the shape of it will be. Optional reflection is self-selecting, and it selects for the students who already understand the material. Any correction process you design has to assume the student who most needs it will skip it unless the structure does not permit skipping.

    The Reflection Sheet Is Not the Intervention

    A form that asks students to reflect does not, on its own, produce reflection. There is a direct test of this.

    Raechel Soicher and Regan Gurung studied “exam wrappers” — short reflection sheets students complete after a returned exam about how they studied and what they will change. Published in Psychology Learning & Teaching in 2017, the study randomly assigned 86 students to three conditions: real exam wrappers with metacognitive instruction, sham wrappers with no instruction, or a control group. There were no improvements in exam performance, final grades, or measured metacognitive ability. Scores on the Metacognitive Awareness Inventory rose over the semester in every condition, including the control, which is a reminder of what happens to an uncontrolled before-and-after comparison.

    The authors suggest the effect may require use across multiple courses rather than one. Fair. But the practical lesson for a secondary teacher is immediate: if your test correction process is a sheet of reflection prompts stapled to the front, you have bought the packaging and not the thing. The questions that produce learning are about the specific item — this question, this error, this misconception — not about study habits in general.

    A Four-Part Test Correction That Produces Thinking

    Each wrong answer gets four boxes. Not three, and not six. This is the smallest structure that makes transcription impossible.

    BoxWhat the student writesWhy it is there
    1. What I putThe original answer, copied over.Makes them look at the error rather than skipping straight to the key. Takes five seconds.
    2. Why I put it“I thought the Senate confirmed treaties on its own.” A sentence naming the belief, not “I didn’t study.”This is the box that does the work. It is also the one students resist, because it requires admitting what they believed.
    3. The right answer, and the evidenceThe correct answer plus where it came from — page, slide, notes, a worked line.Forces a source. A student who cannot find it does not understand it yet, and now you both know.
    4. What would make me miss this again“Any question that uses ‘ratify.’” A trigger, not a resolution.Transfer. The point is the next question of this type, not this question.

    Box 2 is the one people cut when they are short on time, and it is the only one that distinguishes this from an answer key. “Careless mistake” is not an acceptable entry in box 2 more than once per test; if a student writes it three times, the pattern is the finding and it is worth a two-minute conversation.

    For a free-response or math item, box 3 becomes “the first line where it went wrong,” which is more useful than reworking the whole problem and usually faster.

    List of five ways test corrections go wrong, including becoming a points economy and only failing students completing them, each with a fix
    Each of these turns a correction into paperwork.

    How Much Credit Should Test Corrections Be Worth?

    Enough that students do them, little enough that the grade still means something. Half the missed points back is the common answer and it is defensible. So is a flat cap — corrections can move a score up to a 79 and no further.

    Three things to settle before you announce it:

    1. Check your district’s grading policy first. Many boards have adopted language about reassessment, grade replacement, and minimum scores. A teacher-invented points-back scheme that conflicts with board policy is a problem that has nothing to do with pedagogy, and you will lose that argument in March rather than in September.
    2. Write it down and hand it out. Parents are reasonable about a policy that exists and furious about one that appears to change by student. “Up to half the missed points, corrections due within one week, box 2 must be completed” fits on an index card.
    3. Decide whether a correction can raise an A. The honest answer is yes. A student who got a 97 has three errors worth understanding, and excluding the strongest students from the only reteaching structure in the class is backwards. If credit is the only incentive you have, offer the strongest students the thing they actually want: the correction counts, and it gets read.

    A note on what this does to your gradebook. If corrections replace scores rather than adding points, you are most of the way to standards-based grading already, and you should read the honest case against it before you commit, because the implementation problems there are real.

    Test Corrections or a Retake — Which One Do You Need?

    A correction analyzes the test the student already took. A retake is a new assessment of the same material. They answer different questions and they are not substitutes.

    • Use corrections when the errors are scattered. A 71 made of eleven different small mistakes is a diagnosis problem. Corrections surface the pattern.
    • Use a retake when the student did not know the material. A 44 does not have a pattern to find. It has a hole, and corrections on a test you did not understand is just copying.
    • Use both, in order, when the stakes are real. Corrections first, as the price of admission to the retake. It is a reasonable gate and it stops the retake from being a free second roll of the dice.

    Sequencing them that way also solves a fairness problem administrators raise: if a retake is available to anyone who asks, the students who ask are the students whose families know to ask. Requiring the correction makes the path the same for everybody.

    What About the Student Who Just Copies the Right Answer?

    Assume this will happen and build so it does not pay. Box 2 is most of the defence — you cannot copy someone else’s misconception off their paper, because it was not your misconception.

    Beyond that the answers are ordinary. Do corrections in class, not as homework, at least the first few times — so you see the thinking happen and so students without a quiet place to work are not disadvantaged. Spend two of those minutes circulating and asking one student per row to explain box 2 out loud.

    The Grading Burden, Honestly

    If you teach 150 students and read four boxes on every missed item, you will stop doing this by October. Anyone who tells you otherwise has not graded a stack of 150.

    Three ways to make it survivable:

    • Cap it. Students correct their five worst items, not all of them. Five is plenty for a pattern and it bounds your reading.
    • Score box 2 only, and score it pass/fail. A complete, specific sentence gets the credit. A blank or “I didn’t study” does not. You can do that at a glance.
    • Read them as a class set, not as individual papers. Sort the box-2 responses into piles by misconception. Five minutes of sorting tells you what to reteach tomorrow, which is the only reason to collect them at all.

    That last one is also the answer to a coach’s question: the evidence of learning is not the corrected test. It is what you teach differently on the next item of that type, and whether the same misconception shows up again on the unit after this one. If it does, the corrections were decoration. The habit of asking that question after an assessment is the same one behind any serious approach to checking for understanding.

    What Does This Look Like From a Student’s Seat?

    It looks like being asked to write down, on paper, in your own handwriting, the thing you were wrong about. That is harder than it sounds at fifteen, and it is worth naming out loud the first time you assign it.

    Say the thing directly: this is not a punishment, everyone does it including the people who scored highest, and nobody is reading box 2 to laugh at you. Then make that true. A correction process where the teacher reads a student’s misconception back to them in front of the class is a process that will be filled in with nothing real by Thanksgiving.

    Students should also be able to argue. If a student’s box 2 says “I put B because the question says ‘primarily,’ and both A and B are true,” they may be right and the item may be bad. Give points back when they are. A correction process in which the teacher is never wrong is one the students will correctly read as theatre.

    And handle accommodations through the plan, not through the policy. A student with extended time or a scribe on the test has the same entitlement on the correction. A student with a processing or writing accommodation can answer box 2 out loud to you in ninety seconds; the requirement is the thinking, not the handwriting. The plan governs, and designing the correction so that honoring it looks ordinary is easier than making an exception every time.

    Where Test Corrections Go Wrong

    • They become a points economy. Students negotiate credit instead of analyzing errors. Fix: cap the recovery and never bargain over it mid-conversation.
    • Only failing students do them. Which teaches that corrections are what happens when you mess up, rather than what everyone does with a returned test. Fix: everyone corrects, including the 97.
    • They are assigned as homework on day one. The students least likely to complete unsupervised work are the students whose errors you most need to see — the same self-selection Maier and Klotz measured.
    • Nobody reads box 2. If the teacher never responds to the reasoning, students learn within two tests that the reasoning is ceremonial.
    • They substitute for reteaching. A correction is the student’s work. If eleven students missed the same item, that is your work.

    What to Do Next

    Take the next test you are about to hand back and do three things. Require corrections from every student, not just the ones who failed. Make box 2 — why I put it — mandatory for credit, and refuse “careless mistake” as a repeat answer. Then sort the box-2 responses into piles before you plan tomorrow, because that pile is the most honest data you will get all unit.

    The deeper argument here is not about points. It is that a test handed back and filed away teaches a teenager that a wrong answer is a verdict, and a test corrected teaches them it is information. That is the same case the site makes about learning from mistakes everywhere else, and it is the reason a correction belongs in the student’s hands rather than the gradebook — which is also the argument for student self-assessment as a routine rather than an event.

    Before you go: grab the free The Test Correction Sheet and Policy Card (2 pages) (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Do test corrections actually work?

    The underlying mechanism is well supported. Metcalfe’s 2017 review in the Annual Review of Psychology concludes that errorful learning followed by corrective feedback beats error avoidance, and in one lab study errors were corrected on a later test about 70% of the time with feedback against roughly 4% without it. The complication is compliance. In a study of 20,058 assessments by 2,826 students in grades 5–11, the error feedback page was opened in only 44% of cases. Corrections work. Optional corrections mostly do not.

    How much credit should test corrections be worth?

    Enough that students do them and little enough that the grade still means something. Half the missed points back is the common answer and it is defensible; so is a flat cap that lets corrections raise a score to a 79 and no further. Check your district’s grading and reassessment policy before you announce anything, write the policy down, hand it out, and do not vary it by student.

    What is the difference between test corrections and a retake?

    A correction analyzes the test the student already took. A retake is a new assessment of the same material. Use corrections when the errors are scattered — a 71 made of eleven small mistakes is a diagnosis problem. Use a retake when the student did not know the material, because corrections on a test you did not understand is just copying. When the stakes are real, use both in order and make the completed correction the price of admission to the retake.

    How do I stop students from just copying the right answer?

    Require a box that asks why they put what they put — the belief, not "I didn’t study." You cannot copy somebody else’s misconception off their paper, because it was not your misconception. Beyond that, do the first few rounds in class rather than as homework so you can see the thinking happen, and ask one student per row to explain that box out loud.

    Should students who got an A do test corrections?

    Yes. A student who scored 97 has three errors worth understanding, and excluding your strongest students from the only reteaching structure in the class is backwards. It also fixes a dignity problem: if only failing students correct, corrections become what happens when you mess up rather than what everybody does with a returned test.

    How do I grade test corrections for 150 students?

    Cap it at the five worst items per student, score only the reasoning box and score it pass/fail at a glance, and read the set as a class rather than as individual papers — sort the responses into piles by misconception. Five minutes of sorting tells you what to reteach tomorrow, which is the only real reason to collect them. If you are reading four boxes on every missed item for 150 students, you will stop doing this by October.

    Is there a free test corrections template?

    Yes — the two-page PDF linked on this page. Page 1 is the student sheet with three blocks of the four-box correction. Page 2 is a policy card to fill in and staple to the first test of the term, plus a sorting sheet for reading the set as a class. Free, printable, no email address required.

    What should a student actually write in the "why I put it" box?

    A sentence naming the belief that produced the wrong answer — "I thought the Senate confirmed treaties on its own" rather than "I rushed" or "careless mistake." Careless mistake is acceptable once per test. Written three times it is itself the finding, and worth a two-minute conversation. Students should also be allowed to argue: if the box says the item was ambiguous and they are right, give the points back. A correction process in which the teacher is never wrong is one students will correctly read as theatre.

    Sources

    1. Metcalfe, Janet. “Learning from Errors.” Annual Review of Psychology, vol. 68, 2017, pp. 465–489. Review concluding that error avoidance “appears to be the rule in American classrooms” and that errorful learning followed by corrective feedback is beneficial; that corrective feedback “including analysis of the reasoning leading up to the mistake” is crucial; and recommending that educators allow students to commit and correct errors in low-stakes situations. A review of laboratory evidence, not a classroom trial. https://www.annualreviews.org/content/journals/10.1146/annurev-psych-010416-044022
    2. Metcalfe, Janet, and Bridgid Finn. “People’s Hypercorrection of High-Confidence Errors: Did They Know It All Along?” Journal of Experimental Psychology: Learning, Memory, and Cognition, vol. 37, no. 2, 2011, pp. 437–448. Three experiments, 25–45 participants each. High-confidence errors corrected at roughly 82% on the final test; overall recall with feedback about 70%; correction without feedback roughly 4%. College undergraduates answering general-knowledge questions in a laboratory, not secondary students on a unit test. https://pmc.ncbi.nlm.nih.gov/articles/PMC3079415
    3. Maier, Ulrich, and Christian Klotz. “Students Ignore Their Mistakes: Elaborated Error Feedback Processing in a Digital Learning System.” Contemporary Educational Psychology, vol. 82, 2025, article 102395. Observational log analysis of 20,058 formative assessment cases from 2,826 students across 182 secondary classrooms, grades 5–11, 2020–2024. Error feedback opened in 44% of cases; median one error item reviewed; 51% of available errors reviewed on average; prior knowledge the strongest predictor of feedback seeking. Observational, German secondary schools, a digital grammar application rather than a paper test. https://www.sciencedirect.com/science/article/pii/S0361476X25000608
    4. Soicher, Raechel N., and Regan A. R. Gurung. “Do Exam Wrappers Increase Metacognition and Performance? A Single Course Intervention.” Psychology Learning & Teaching, 2017, pp. 64–73. 86 students randomly assigned to exam wrappers, sham wrappers, or control. No improvement in exam performance, final grades, or Metacognitive Awareness Inventory scores; MAI rose in all conditions including control. University students, single course. https://liberalarts.oregonstate.edu/biblio/do-exam-wrappers-increase-metacognition-and-performance-single-course-intervention
    5. Rice, Bethany S. “How Extra Credit Quizzes and Test Corrections Improve Student Learning While Reducing Stress.” ASEE 127th Annual Conference & Exposition, 2020. Average exam scores rose 5% (2018) and 3% (2019) when corrections were offered; 80% of students participated; survey responses strongly favourable. Cited as a descriptive classroom report, not a controlled study: no comparison group, 31 survey respondents, university engineering technology students. https://peer.asee.org/how-extra-credit-quizzes-and-test-corrections-improve-student-learning-while-reducing-stress.pdf
    6. McDade, Margaret. “Using Test Corrections as a Learning Tool.” Edutopia, George Lucas Educational Foundation. A high school math teacher’s small-group correction routine, with retesting for full credit rather than partial credit on corrections. Cited as practice knowledge: the article reports a teacher’s own method and cites no studies. https://www.edutopia.org/article/test-corrections-high-school-math/

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Late Work Policy: What the Evidence Supports After the Famous Study Was Retracted

    Late Work Policy: What the Evidence Supports After the Famous Study Was Retracted

    By Clay Shumate

    A late work policy decides two separate things: what happens to the grade, and what happens to the student. Most policies collapse those into one number and then stop working. The evidence does not hand you a clean answer, and the study most often quoted in these arguments was retracted in September 2026 — which is the first thing anybody writing a policy this year should know.

    Below is what the research actually supports, what it does not, and a six-part policy you can write on one page and still defend in a parent meeting.

    Key Takeaways

    • The famous deadline study is gone. Ariely and Wertenbroch’s 2002 paper on spaced deadlines was retracted on 2 September 2026. A 124-person replication found no effect of deadline condition on any outcome measure.
    • Loosening grading has not been shown to help students. Students assigned to stricter-grading teachers scored higher in math — in that class and in later ones — across every subgroup studied.
    • Removing late penalties has a measured cost. In one high school chemistry class, homework completion fell by more than a third.
    • The zero is a math problem before it is a policy problem. On a 100-point scale, the gap between passing grades is 10 points and the gap from D to F is 60.
    • The deeper issue is the scale and the averaging, not the zero. Research supports grading scales with four to seven levels for reliability; the 100-point scale invents precision that is not there.
    • Separate the grade from the behavior. Report lateness as conduct and let the grade report what the student knows. That one move resolves most of the argument.

    Free Download · PDF

    The One-Page Late Work Policy and Window Log (2 pages)

    Page 1 is the six-part policy to fill in, the practice-versus-assessment split, and the sentence to give a parent. Page 2 is the window log that turns the policy into information, plus the honest cost of all three common policies.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Why Is a Late Work Policy So Hard to Get Right?

    Because a grade is being asked to do two jobs at once, and they conflict.

    Job one is reporting what a student knows and can do. Job two is enforcing a deadline. A 10-percent-per-day deduction does job two by corrupting job one: after five days the number on the report card is half knowledge and half calendar, and nobody reading it — not the next teacher, not a parent, not the student — can tell which half is which.

    A zero does job two harder and job one worse. And the people on the other side of the argument are not making a soft case. The claim that deadlines do not matter has a measured cost attached, which is the part that usually goes missing in articles about this.

    So the honest version of the problem is not “should kids face consequences.” It is: how do you keep the deadline real without making the grade lie?

    Four evidence findings on late work policy: stricter grading linked to higher math scores and a one-third drop in homework completion when late penalties were removed, a small undergraduate trial where an early-bonus plus late-penalty policy produced work 1.45 days early, the arithmetic problem of a zero on a 100-point scale, and the September 2026 retraction of the Ariely and Wertenbroch deadline study
    Four findings that do not all point the same way. That is the honest picture.

    What Does the Research Actually Say About Deadlines?

    Less than you have been told, and one widely quoted finding has been formally withdrawn.

    For twenty years, the standard citation for “students do better with evenly spaced interim deadlines” has been Ariely and Wertenbroch (2002) in Psychological Science. You have probably read it quoted in a PD slide deck. It was retracted on 2 September 2026.

    The retraction is not a technicality. Data Colada’s analysis of Study 2 found eighteen of twenty participants in one condition had exact duplicates across all three tasks, correlations that should have been strong were absent, and self-reported times showed almost none of the rounding that real human responses show. Their conclusion was that the data were “severely tampered with or fabricated.” Coauthor Klaus Wertenbroch stated that “much or all of the data — and therefore the results — are false.”

    A pre-registered replication by Hyndman and Bisin, published in 2025 with 124 participants across three deadline conditions, found no statistical evidence that performance was influenced by the deadline condition on any of three measures — errors found, days late, or payment, all p > 0.1. The original had reported all differences significant at p < 0.01. The replication authors concluded that the received wisdom about spaced deadlines limiting procrastination is “possibly false.”

    Two things follow. First, if your department’s late work policy was built on that study, it needs a different foundation. Second — and this is the part worth holding onto — the retraction says nothing about whether deadlines matter in a classroom. It says one famous laboratory result about self-imposed versus imposed deadlines cannot be relied on. Interim checkpoints on a long project may still be good practice. They just are not evidence-backed in the way everyone has been saying.

    Does Removing Late Penalties Hurt Students?

    There is real evidence that looser grading costs something, and it deserves to be stated as plainly as the case for reform usually is.

    Gershenson’s work on high school math found that students with tougher-grading teachers scored higher in math, both in that teacher’s class and in subsequent math courses, and that this held across every student subgroup. Figlio and Lucas found the same direction at elementary level: students assigned to stricter graders showed greater test-score growth in reading and math. In one high school chemistry class where late penalties were removed, homework completion dropped by more than a third.

    The strongest single trial on late policies specifically is small and not from a high school. Korpusik, Freitas and Dionisio compared four late policies across 248 lab submissions in an introductory programming course. Work came in 2.43 days late under no policy, 4.71 days late under an early-completion bonus alone, 0.44 days late under a late penalty, and 1.45 days early under a combined early bonus and late penalty — which also produced the highest grades. Their read was that late penalties externally regulate students who have not yet regulated themselves.

    Size that evidence honestly before you use it: 31 students who consented to the analysis, mostly first-years, taught online during the pandemic. It is a signal, not a mandate. But it points the same direction as the grading-strictness work, and a teacher who ignores both because the conclusion is unfashionable is doing the thing this site complains about when it goes the other way.

    Does It Matter What the Assignment Was For?

    More than any other question here, and most policies never ask it.

    If the work was practice — a problem set, a draft, a reading check whose job was to tell you what to reteach — then late is a timing failure with a real instructional cost, because the information arrived after you needed it. The honest response is to collect it, use it, and record the lateness as conduct. Scoring it down does not recover the information.

    If the work was the assessment — the essay, the project, the unit test — then the deadline is doing something different. It is the point at which you are claiming to know what the student can do. Here a window still makes sense, but a narrow one, and after it closes the student demonstrates the standard some other way rather than handing in the same artifact in December.

    Running one blanket rule over both is why so many late work policies feel wrong in practice. A missing formative check and a missing final project are not the same event and should not get the same sentence.

    So Why Not Just Give Zeros?

    Because of arithmetic, not sentiment.

    Reeves laid this out in Phi Delta Kappan. On a standard 100-point scale with letter grades at 10-point intervals, every passing grade sits 10 points from its neighbor — and the interval between D and F is not 10 points but 60. A single zero therefore carries roughly six times the weight of any other failing mark in an average. Reeves’s point is that if an F is one interval below a D, the mathematically consistent value is 50, not 0. On a 4-point scale nobody hesitates: missing work gets a 0, exactly one point below a 1. The same logic on a 100-point scale would require a –6, which no one would give.

    Guskey, Fisher and Frey take that further in Educational Leadership, and their version is the one to carry into a faculty meeting: “The true problem is not the zero; it’s the use of the 100-point percentage grading scale and the practice of averaging scores.” They point out that the problems with percentage scales have been documented since Starch and Elliott in 1913, and that research supports scales with four to seven levels for optimal reliability and discrimination. A 101-level scale manufactures precision nobody can actually defend.

    That reframing matters for a late work policy because it tells you where the real fix is. Arguing about whether a missed assignment is a 0 or a 50 is arguing about a symptom. If your gradebook averages percentages across a semester, a single missed assignment distorts the picture no matter which number you put in the box.

    If this is live at your school, the fuller version of that argument is in the standards-based grading guide, and the honest case against it is in why standards-based grading doesn’t work. Both are worth reading before anyone proposes a schoolwide change.

    Comparison of three common late work policies with honest costs: a zero after the due date, ten percent off per day late, and full credit with no deadline, each listed with what it is honest about and what it costs
    None of these is free. The question is which cost you can live with.

    What Does a Workable Late Work Policy Look Like?

    Six parts. It fits on one page, and every part of it survives a parent asking why.

    Six numbered components of a workable late work policy: state what the deadline is for, set a hard floor instead of a zero, separate the grade from the behavior, publish one window and hold it, make the make-up cost time rather than points, and record which students keep using the window
    Six parts, and the sixth one turns the policy into information.

    1. Say what the deadline is for

    “This is due Friday because on Monday we build on it” is a reason. “Because I said Friday” is not, and teenagers are unusually good at telling the difference. A deadline with a downstream purpose gets taken seriously by more students than a deadline without one, and it also tells you which deadlines you should actually defend. If nothing depends on Friday, Friday was arbitrary and you should stop pretending otherwise.

    2. Set a hard floor, not a zero

    A missed assignment scores the bottom of the scale, not the bottom of the number line. On a four-point scale that is a 0. On a 100-point scale it is a 50. The point is not generosity; it is that the floor should be one interval below the lowest passing mark, which is what every other grade boundary already is.

    3. Separate the grade from the behavior

    This is the move that resolves most of the argument, and it costs nothing. Lateness is a conduct fact. Report it as one — a comment on the report, a contact home that happens the same week, a scheduled session — and let the grade report what the student knows. A parent who is told “she understands this material and she has turned in four assignments late” has two usable pieces of information. A parent told “she has a 61” has none.

    4. Publish one window and hold it

    “Late work is accepted until the unit assessment” is a rule a fourteen-year-old can plan against. “Depends when you ask me” is not a policy, it is a mood, and students read inconsistency as unfairness faster than they read strictness as unfairness. Pick a window, write it in the syllabus, and then do exactly what you said — which is the whole of respecting students in practice rather than on a poster.

    5. Make the make-up cost time, not points

    Being late should cost something. Time is the honest currency: a scheduled session at lunch, before school, or during an intervention block. It is a real cost, students feel it, and it leaves the grade intact. A ten-percent-per-day deduction costs them something too — the accuracy of the transcript.

    6. Write down who keeps using the window

    Keep a list. Not to punish with — to read. Three students using the late window every single time is not a late work problem, it is a signal about workload, home, organization, or something nobody has asked about yet. A policy that generates that list is doing a second job for free. A policy that just applies a deduction tells you nothing you did not already know. This is the same logic as treating attendance data as a referral system rather than a compliance record. The same read applies to arrival times, where a tiered response sorts the one-off from the pattern in a way a blanket deduction never will.

    One cost this policy does carry, and it should be named rather than buried: an open window produces a flood at the end of it. If the window closes at the unit assessment, expect a stack the night before, and expect it to arrive in the same week you are marking the assessment itself. Two things keep that survivable. Cap what comes back — late work gets a score and a one-line comment, not the full written feedback a punctual draft gets, and say so in advance. And put the make-up session mid-unit rather than at the end, so the work trickles in instead of arriving at once. A policy that quietly doubles your marking in the last week of a unit is a policy you will abandon by Thanksgiving, which is its own kind of unfairness. The workload side of this is not a side issue; it decides whether the policy still exists in March.

    What Does This Look Like From the Student’s Side?

    Clear, and askable in advance. Those are the two things students actually want from a late work policy, and neither is the same as lenient.

    Clear means they can find it. Not buried in a syllabus they signed in August — posted, in nine words, where the due dates are. A student should be able to answer “what happens if I turn this in Monday” without asking you.

    Askable in advance means the policy has a front door. A fifteen-year-old working a closing shift, or watching younger siblings, or without reliable internet at home, usually knows on Tuesday that Friday is not going to happen. Right now most classrooms give that student nothing to do with that knowledge except apologize on Friday. Say out loud, more than once, that asking before a deadline is a different conversation than explaining after one — and then make it true, which means the student who asks on Tuesday gets a straight answer rather than a lecture.

    That is not softness. It is the difference between treating a teenager as someone managing competing obligations, which they are, and treating them as someone who needs catching out. The deadline does not move for the student who never asks. It is just that the student who plans ahead gets something for planning ahead, which is the behavior the whole policy claims to be teaching.

    How Do You Explain This to a Parent Who Thinks It Is Too Soft?

    Lead with what did not change, because something did not.

    The work is still required. The deadline is still real. There is still a consequence, and it is time rather than points. What changed is that the report card now tells you what your child knows instead of telling you a blend of what they know and when they handed it in.

    That sentence survives most objections because it is not a concession, it is a clarification. And it is worth having ready in writing — in the syllabus and in the first message home — rather than improvised on the phone in October. Parents who object to no-zero policies are usually objecting to the version where nothing is required; the fastest way to settle it is to show this is not that version.

    There is a fair version of the objection too, and it should not be waved off. If a student can turn anything in whenever they like, some will discover that and use it, and a policy that pretends otherwise is not being honest. That is precisely why parts 4, 5 and 6 exist: a published window, a real cost in time, and somebody actually watching the list.

    What Does This Mean for a Department or a School?

    Consistency across a hallway matters more than the specific policy chosen.

    A student with six teachers running six different late policies cannot plan, and the student least able to absorb that is the one the policy was supposed to help. If a department can agree on a window and a floor, that is worth more than any individual teacher’s preferred version of either.

    Two cautions for anyone proposing this above the classroom level. A policy that requires teachers to accept unlimited late work without additional marking time is a workload decision disguised as a grading decision, and it will be abandoned by March. And a schoolwide floor applied on top of a 100-point averaging gradebook is, by Guskey, Fisher and Frey’s argument, treating the symptom — worth doing, but not worth calling a reform.

    Two constraints sit above everything on this page, and both of them outrank it.

    Your district may already have a grading policy with the force of board approval. A teacher who unilaterally sets a 50 floor in a district whose policy says otherwise has a problem that has nothing to do with pedagogy. Read the policy before writing yours, and if the two conflict, that is a conversation with an administrator rather than a decision to make quietly in a gradebook.

    And an IEP or 504 plan governs. If a student’s plan includes extended time or modified deadlines, the plan is the policy for that student, full stop — no classroom rule and no research finding on this page overrides it. Build your policy so that honoring a plan looks like the ordinary case rather than a visible exception, which mostly means making the window generous enough that nobody has to be singled out to use it.

    What to Do Next

    Write the policy on one page before the next unit starts. One window. One floor. One sentence saying lateness is reported as conduct, not deducted from the grade. One line about what the make-up session is and when it runs.

    Then put it in the syllabus and in a message home the first week, so that the first time a family hears about it is not after a missed assignment.

    And keep the list. At the end of the first unit, look at who used the window and how often. That list will tell you more about your classroom than the policy itself does.

    Before you go: grab the free The One-Page Late Work Policy and Window Log (2 pages) (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Is it true that the study everyone quotes about deadlines was retracted?

    Yes. Ariely and Wertenbroch’s 2002 paper in Psychological Science on spaced deadlines and procrastination was retracted on 2 September 2026, after analyses found duplicated observations, missing expected correlations and implausible response patterns in Study 2. One of the coauthors stated that “much or all of the data — and therefore the results — are false.” A 2025 replication with 124 participants found no significant effect of deadline condition on any outcome. If a PD session or a department policy still cites that study, it needs a new basis. It does not mean deadlines are useless — it means that one famous laboratory result cannot be used as evidence.

    Is giving a 50 for work that was never done just a gift?

    It is a scale decision, not a generosity decision. On a 100-point scale every passing grade sits 10 points from the next, and the gap from D to F is 60 — so a zero carries about six times the weight of any other failing mark in an average. A floor at 50 makes the F one interval below the D, which is what every other boundary already is. The work is still missing, the student still has not demonstrated the standard, and the grade still shows a failure. What changes is that one missing assignment can no longer outweigh several completed ones.

    Can a teacher set a 50 floor if the district grading policy says otherwise?

    No, and this should be checked before anything on this page is implemented. Many districts have a board-approved grading policy, and a teacher who quietly overrides it in a gradebook has created a problem that is not pedagogical. Read the policy first. If it conflicts with what you believe is right, that is a conversation to have with an administrator, and the argument from grading-scale reliability is a strong one to bring to it — but it is a conversation, not a unilateral change.

    How does a late work policy interact with an IEP or 504 plan?

    The plan governs, without exception. If a student’s plan provides extended time or modified deadlines, that is the policy for that student and no classroom rule overrides it. The practical design point is to build the general policy so that honoring a plan looks ordinary rather than exceptional — a window generous enough that a student using it is not visibly marked out. If you are unsure how a plan applies to a specific assignment, ask the case manager before the deadline rather than after.

    Should practice work and assessments have the same late policy?

    No, and running one rule over both is why many late policies feel wrong in practice. Practice work — problem sets, drafts, reading checks — exists to tell you what to reteach, so when it arrives late the instructional cost is already paid and scoring it down recovers nothing. Collect it, use it, record the lateness as conduct. An assessment is different: the deadline is the point at which you claim to know what a student can do. Keep a window there too, but a narrow one, and after it closes have the student demonstrate the standard another way rather than submitting the same artifact months later.

    Won’t accepting late work bury me in grading?

    It can, and a policy that ignores that will not survive to March. Two things keep it manageable. Cap what comes back: late work gets a score and a one-line comment rather than the full written feedback a punctual draft earns, and say so in advance so it reads as a stated rule and not as neglect. And schedule the make-up session mid-unit rather than at the window’s close, so work trickles in instead of landing as a stack the night before the unit assessment.

    Does removing late penalties actually hurt students?

    There is real evidence that it costs something, and it should not be waved away. In one high school chemistry class, homework completion fell by more than a third when late penalties were removed. Separately, students assigned to tougher-grading teachers scored higher in math — in that class and in later courses, across every subgroup studied. None of this proves zeros are correct, and none of it is a randomized trial of a specific late policy. What it does say is that “no deadline, no consequence” is a position with measured costs, which is why the policy described here keeps a real cost and moves it from points to time.

    What should a student do if they know in advance they cannot meet a deadline?

    Ask before it, not explain after it — and the teacher’s job is to make that worth doing. A student working a closing shift or caring for siblings usually knows on Tuesday that Friday will not happen. Most classrooms give them nothing to do with that information. Say out loud, more than once, that a request made before a deadline gets a different conversation than an apology made after one, then honor it: a straight answer rather than a lecture. The deadline does not move for the student who never asks. Planning ahead is the behavior the policy claims to be teaching, so it should get something.

    Sources

    1. Retraction notice: Ariely, D., & Wertenbroch, K. (2002), “Procrastination, Deadlines, and Performance: Self-Control by Precommitment,” Psychological Science. Retracted 2 September 2026. The notice cites a replication study and Data Colada analyses that “raised questions about the underlying data” and “called into question the veracity of the overall findings.” Coauthor Wertenbroch is quoted: “much or all of the data — and therefore the results — are false.” https://retractionwatch.com/?p=135929
    2. Data Colada, post 138, on Study 2 of Ariely & Wertenbroch (2002). Eighteen of twenty participants in the Last Day Deadline condition had exact duplicates across all three proofreading tasks, with ID numbers ten positions apart; expected correlations were absent (original task-performance correlations +.03 to +.27 against +.74 to +.90 in replication); self-reported times showed 11.7% rounding against 85% in the replication. Conclusion: the data “were severely tampered with or fabricated.” https://datacolada.org/138
    3. Hyndman, K., & Bisin, A. (2025). Replication of Ariely & Wertenbroch (2002). 124 participants, three randomly assigned deadline conditions (none, evenly spaced, self-imposed), three proofreading tasks over three weeks. No statistically significant effect of deadline condition on errors found (F = 0.181), days late (F = 0.353) or payment (F = 0.215), all p > 0.1. https://bpb-us-e1.wpmucdn.com/wp.nyu.edu/dist/c/16384/files/2025/09/Ariely_Replication-1.pdf
    4. Korpusik, M., Freitas, J., & Dionisio, J. D. N. (2022). “Impact of Late Policies on Submission Behavior and Grades.” ASEE Annual Conference. Loyola Marymount University, introductory programming lab, 248 submissions, 31 students consenting to analysis, mostly first-years, taught online during the pandemic. Average submission timing: no policy +2.43 days, early incentive +4.71, late penalty +0.44, combined −1.45 days (early). Grades 94.2% / 96.6% / 97.1% / 99.5% respectively. Undergraduates, not grades 6–12 — the mechanism may transfer, the numbers do not. https://people.csail.mit.edu/korpusik/asee22.pdf
    5. Guskey, T. R., Fisher, D., & Frey, N. “The Unwinnable Battle Over Minimum Grades.” Educational Leadership (ASCD). Argues minimum-grade floors treat a symptom: “The true problem is not the zero; it’s the use of the 100-point percentage grading scale and the practice of averaging scores.” Cites Starch and Elliott (1913) on percentage-scale problems and Lozano et al. (2008) and Preston and Colman (2000) on four-to-seven-level scales producing optimal discrimination, validity and reliability. https://www.ascd.org/el/articles/the-unwinnable-battle-over-minimum-grades
    6. Reeves, D. B. (2004). “The Case Against the Zero.” Phi Delta Kappan, 86(4). The arithmetic argument: on a 100-point scale the interval between passing grades is 10 points while “the interval between the D and F is not 10 points but 60 points,” making the mathematically consistent value of an F 50 rather than 0. https://www.researchgate.net/publication/285846142_The_Case_against_the_Zero
    7. Thomas B. Fordham Institute. “Think Again: Does ‘equitable’ grading benefit students?” Cited as a review, not as the primary studies. This is where the grading-strictness findings above come from: Gershenson (2020) on high school math students with tougher-grading teachers scoring higher in that class and in later courses across all subgroups; Figlio and Lucas (2004) on elementary test-score growth; and the high school chemistry case in which removing late penalties cut homework completion “by more than one-third.” The review’s own position is that there is no hard evidence that more lenient grading benefits students long term — a contested claim, and presented here as the strongest version of the case against loosening a late work policy. https://fordhaminstitute.org/national/research/think-again-does-equitable-grading-benefit-students

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Collective Teacher Efficacy: What the 1.57 Effect Size Actually Means

    Collective Teacher Efficacy: What the 1.57 Effect Size Actually Means

    By Clay Shumate

    Collective teacher efficacy is the shared belief among a staff that, working together, they can move student learning — and it carries the largest effect size in John Hattie’s rankings, 1.57. That number is real, it is correlational, and it is almost certainly too large to mean what your next professional development day will tell you it means.

    Both halves of that sentence matter. Below is where 1.57 came from, why the researcher who published it alongside Hattie thinks numbers like it get misread, and what a secondary department can actually do with the idea once the slide deck is over.

    Key Takeaways

    • Collective teacher efficacy is a belief about the group, not about you. Not “I can teach this kid” — “we can teach these kids.” That distinction is the whole construct.
    • The 1.57 figure is correlational. It describes schools where staff belief and student achievement travel together. It does not establish which one moved first, and the honest answer is that success probably builds belief at least as much as belief builds success.
    • The same journal published the rebuttal. Matthew Kraft, writing with Hattie in Educational Leadership, calls effect size “a misleading term” and reports that only 13 percent of nearly 2,000 effect sizes from randomized trials reached 0.40 at all.
    • A big effect size is not a plan. Nothing about 1.57 tells you what to do Monday, which is why it shows up on so many slides and changes so few schools.
    • What the research does name is conditions — evidence that the work is landing, structures where teachers actually talk, and enough safety that a teacher can say “that lesson did not work” out loud.
    • The cheapest version of this is one conversation. Ask a colleague how they teach the thing you teach badly. That is the whole mechanism, and it does not require a framework.

    Free Download · PDF

    The Effect Size Sanity Check

    Five questions to run any effect size through before it reaches a slide or a board packet — plus Hattie’s hinge point set beside Kraft’s causal benchmarks, the 1.57 figure worked all the way through, and a blank sheet for the next number somebody hands you.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    What Is Collective Teacher Efficacy?

    It is the shared belief among the teachers in a building that their combined effort can move student learning, including for the students who are hardest to move. The unit is the group. That is what separates it from ordinary teacher confidence.

    The difference sounds academic until you hear it in a hallway. “I can get through to that kid” is individual efficacy. “The kids who come to us behind, leave us less behind” is collective. The second one is a claim about what a staff can do together, and it is the kind of sentence you either hear in a building or you do not.

    You can hear the absence of it just as clearly. A department where the honest summary is “there is only so much we can do with these kids” has low collective efficacy, and everybody in it knows, and nobody says it in those words.

    Two things are worth naming before the numbers. Collective efficacy is measured by asking teachers what they believe, usually on a survey. And the outcome it is correlated with is student achievement, usually on a test. Both of those are real measurements and both have the obvious limits of a survey and a test.

    Where Does the 1.57 Effect Size Come From?

    From Donohoo, Hattie and Eells, writing in ASCD’s Educational Leadership in March 2018. They place collective teacher efficacy at an effect size of 1.57, and they compare it directly against three things schools usually treat as fixed.

    Here is the comparison as they present it:

    • Collective teacher efficacy — 1.57
    • Prior achievement — 0.65
    • Socioeconomic status — 0.52
    • Home environment — 0.52

    Their framing is that collective efficacy is “greater than three times more powerful and predictive of student achievement than socioeconomic status,” and more than double the effect of prior achievement. You can see why it travels. It is a research finding that says the thing schools control matters more than the things they do not, and every educator in the room wants that to be true. I want it to be true.

    One caution about that comparison before anyone puts it in a board packet. “Beats socioeconomic status three times over” is one short step from “poverty is not the problem, attitude is,” and that step gets taken in rooms where budgets are decided. The study does not support it. Comparing a belief measured by survey against a demographic measured by household income tells you those two numbers relate differently to test scores — not that one can be swapped for the other. A staff with tremendous collective efficacy and no counselor is still a staff with no counselor.

    That is exactly why the number deserves a harder look than it usually gets.

    Why the 1.57 Is Almost Certainly Too Big

    Because effect sizes drawn from correlational research are inflated relative to effect sizes drawn from experiments, and 1.57 is the correlational kind. The clearest statement of this problem was published by ASCD too, three years later, with Hattie himself on the byline.

    Two-column comparison of correlational and experimental research, noting that only 13 percent of nearly 2,000 randomized-trial effect sizes reached 0.40
    A correlation says two things travel together. It does not say which one moved first.

    In that 2021 piece, Brown University economist Matthew Kraft argues that “effect size is such a misleading term,” because an effect size “tells us nothing about whether the underlying relationship represents cause and effect or a simple correlation.” It is a translation of a relationship onto a common scale. It is not a measure of whether one thing caused the other.

    Kraft goes further, and this is the part that should change how you read every effect size you are shown. Looking at nearly 2,000 effect sizes from randomized controlled trials — studies built to establish cause — he found that only 13 percent reached 0.40 or larger. Hattie’s widely quoted 0.40 “hinge point,” the line above which an intervention is supposedly worth having, sits near the top of what carefully controlled education research actually produces.

    So Kraft proposes different benchmarks for causal studies: below 0.05 is small, 0.05 to 0.20 is medium, and above 0.20 is large. Against that scale, an effect size of 0.20 is not a disappointment. It is a genuine win.

    Hattie’s reply in the same article is fair and worth stating: 0.40 was always an average across all the studies in his database, offered as a starting point for a conversation rather than a universal pass mark. That is a reasonable defense of what he wrote. It is not much of a defense of how the number gets used in a faculty meeting.

    The specific problem with 1.57 is direction. Schools where teachers believe they can collectively succeed also tend to post higher achievement. Nothing in a correlation tells you which came first — and the common-sense reading runs at least partly backwards. A staff whose students have been improving for three years has excellent reasons to believe in itself. Belief is downstream of evidence, not only upstream of it.

    The honest version of the claim is smaller and still worth having: collective efficacy and achievement travel together, the relationship is strong, and the arrow probably runs both directions. That is defensible. “Believe harder as a staff and test scores follow” is not.

    This is not a new standard for this site. We ran the same check on formative assessment’s headline numbers, where a widely repeated 0.4 to 0.7 range comes from a 1998 review and a later meta-analysis landed nearer 0.20.

    Does That Mean Collective Efficacy Doesn’t Matter?

    No. It means it is a condition, not an intervention. Those are different things and schools keep buying the second when the research describes the first.

    An intervention is something you do to a class on Tuesday. A condition is the state of the building that determines whether Tuesday’s thing has a chance. Collective efficacy is the second kind. You cannot run it in fourth period. You also cannot run much of anything well without it.

    Which is why “we are going to build collective teacher efficacy this year” is a sentence that means nothing on its own. It is like announcing that the department will have good morale. The belief is the readout, not the lever.

    That has a direct purchasing consequence, and it is the one a district leader should take from this. You cannot buy a readout. A program promising to raise collective teacher efficacy is selling you the dial rather than the engine, and the survey score will move whether or not anything underneath it did. The spend that actually maps to the research is duller and harder to put in a press release: protected common time, and the data work required to show a staff whether last year’s push moved anything.

    What Actually Builds Collective Efficacy?

    Evidence that the work is landing, structures where teachers genuinely talk, and enough safety to say “that did not work” out loud. Donohoo, Hattie and Eells name a set of enabling conditions, and none of them is a belief you can adopt by deciding to.

    Numbered list of five enabling conditions for collective teacher efficacy: evidence of impact, high-trust collaborative structures, non-threatening evidence-based environments, social sensitivity and empathy, and leaders who model psychological safety
    Four of the five belong to whoever controls the schedule and the tone.

    The conditions they list:

    • Evidence of impact on student outcomes. Teachers believe the group can move learning when they have watched the group move learning. This is first for a reason, and it is the one that explains the direction problem above.
    • High-trust collaborative structures. Time to work together that is protected, regular, and not consumed by announcements.
    • Non-threatening, evidence-based environments. The conversation is about what the student work shows, not about whose class it came from.
    • Social sensitivity and empathy among team members. Teams where people read each other well outperform teams stacked with individual talent.
    • Leaders who model psychological safety. The principal who says “I got that wrong” in front of the staff has done more for this than the principal who buys the book.

    Read that list again and notice what is not on it. No slogan. No banner. No survey. Four of the five are structural and belong to whoever controls the schedule and the tone, which in most buildings is not the classroom teacher.

    None of those four is free, and a protected common period is the most expensive thing on the list — it is a master-schedule decision that costs staffing, and pretending otherwise is how this advice gets dismissed by the person who builds the schedule.

    That is not a reason for a teacher to disengage. It is a reason to be precise about the ask. If your department wants this, the request is not “more collaboration.” It is a protected common period, an agenda that is about student work, and an administrator who does not treat a failed lesson as a performance problem. Those are three specific things an administrator can say yes or no to — which is the same reason a specific, answerable ask beats a general one every time.

    The Effect Size Sanity Check

    Five questions to run any effect size through before you repeat it. None of them requires statistics training, and together they catch most of what goes wrong between a study and a slide. It is the same scepticism worth bringing to any claim made for a professional development program.

    Numbered list of five questions to ask about any effect size: correlational or experimental, which direction common sense would run, who the participants were, what it was compared against, and what would have to be true for it to be wrong
    1. Correlational or experimental? If nobody was assigned to anything, the number describes a relationship, not a result. Read it as “these travel together,” never as “this causes that.”
    2. Which direction would common sense run? Ask it out loud. If the reverse story is at least as plausible — success building belief rather than belief building success — say so when you cite the number.
    3. Who were the participants? A finding from nine-year-olds is not a finding about your sophomores. This one disqualifies more education research than any other question on the list.
    4. Compared against what? An effect size is always a comparison. Against a good alternative, 0.20 is excellent. Against nothing at all, 1.00 is unremarkable.
    5. What would have to be true for this to be wrong? If you cannot answer, you have not finished reading the study — you have finished reading a summary of it.

    Run the 1.57 through those five. It fails the first, wobbles on the second, passes the third, and the fourth and fifth are rarely asked at all. That does not make it worthless. It makes it a strong correlation reported honestly, which is a perfectly respectable thing for a number to be.

    What Kills Collective Efficacy Faster Than Anything Else

    Being asked to collaborate while being evaluated on the result. Every other failure mode on this list is a version of that one.

    • Data meetings that are really accountability meetings. The moment a common assessment becomes evidence about the teacher rather than evidence about the learning, honest conversation ends and everyone brings their best class.
    • Collaboration time that gets eaten. A protected period consumed by announcements three weeks running teaches the staff exactly what the period is worth.
    • Mandated agendas. A team handed its problem of practice is a team doing compliance. The problem has to be one the teachers actually have.
    • Treating a failed lesson as a performance issue. One instance of this and nobody in the building admits to a failed lesson again. This is the single most expensive thing an administrator can do here.
    • Celebrating belief instead of evidence. Posters about believing in kids, in a building where nobody can show whether last year’s reading push worked, produce cynicism rather than efficacy.

    What This Looks Like in a Grades 6–12 Department

    Small, specific, and about a unit you are both teaching. The research operates at the level of a building. The version a teacher can start does not.

    The cheapest possible version is one conversation. Pick the thing you teach worst — everyone has one, and if you cannot name yours you have not thought about it hard enough — and ask a colleague who teaches it better how they run it. That is it. That is the mechanism the whole construct is describing, stripped of the vocabulary.

    I would not have written that sentence in my first year. I learned a review format from another teacher on my hall, and the only reason I learned it is that he noticed I was short on ways to review and offered. Keeping your ears open and your mouth closed sometimes pays off. Keeping your ears open and asking questions pays off more, and it is faster.

    The next version up is a shared artifact. Two teachers, one common assessment or one set of student work, one question: what did students do here that we did not expect? Not “how did your class do.” What did the work show. The best department-level version of this I have been part of started with an instructional coach, not an administrator, which is worth saying plainly — a coach is a partner in this, not an evaluator.

    Three things make a two-person version work where a mandated team often does not:

    • You picked the problem. Not the school improvement plan.
    • The evidence is student work, not scores. Work can be looked at together without anyone being graded — and it rarely needs a name on it. Take the name off when the name is not the point.
    • It repeats. Once is a chat. Every two weeks for a semester is a practice, and the same frequency argument runs under every structure worth building — the first run is about the structure, and only the fourth or fifth is about the content. A signal that took six weeks of holding the line before it worked is the same lesson one classroom down.

    None of this needs permission, a budget, or a framework. It needs one colleague and a recurring thirty minutes, which is also roughly what an honest reflection practice costs.

    What Do Students Experience When a Staff Has This?

    Consistency — and they notice its absence long before any adult names it. This is the part of the research that never makes the slide, and it is the part that matters most on a site about treating teenagers as developing young adults.

    A student moving through six classrooms in a day is running an experiment on whether the adults in that building talk to each other. When late work means three different things in three rooms, when one teacher’s “prepared” is another’s bare minimum, when nobody can tell them why the rule changed at the door — that is what low collective efficacy feels like from a desk. Students do not have the vocabulary for it. They describe it as the school not making sense, and they are describing it accurately.

    The reverse is just as visible. In a building where teachers genuinely work together, a student gets the same answer twice, and the second teacher already knows what the first one tried. Teenagers read that instantly, and they read it as being taken seriously. It is the difference between a set of adults and a staff.

    Worth saying plainly: none of that requires students to be told about collective efficacy, and no fourteen-year-old should sit through a presentation about it. They will judge the staff by whether the answers match.

    What to Do Next

    If you are a teacher: name the unit you teach worst, find the person who teaches it better, and ask them this week. One conversation. You do not need to call it anything — and if you are early enough in the job that asking feels like admitting something, that is exactly backwards.

    If you lead a department or a building: stop asking for belief and start supplying the two things it grows from — evidence that the work is landing, and a room where a teacher can say a lesson failed without it costing them. Protect the common period, put student work on the table instead of a spreadsheet, and go first when something does not work.

    And if somebody puts 1.57 on a slide in front of your staff, the useful question is not whether the number is right. It is the one that follows: what would we have to see in our own students’ work to believe it about us? That question has an answer, and chasing it is the actual work.

    Before you go: grab the free The Effect Size Sanity Check (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    What is collective teacher efficacy?

    It is the shared belief among the teachers in a school that, working together, they can move student learning — including for the students who are hardest to reach. The unit is the group, which is what separates it from individual teacher confidence. “I can get through to that student” is individual efficacy; “students who arrive behind leave us less behind” is collective.

    What is the effect size of collective teacher efficacy?

    Donohoo, Hattie and Eells put it at 1.57 in ASCD’s Educational Leadership in March 2018, compared against 0.65 for prior achievement and 0.52 for both socioeconomic status and home environment. That figure is correlational, which matters a great deal for how it should be read — it describes belief and achievement travelling together, not belief causing achievement.

    Is the 1.57 effect size reliable?

    It is a real figure from real studies, and it is almost certainly larger than any causal effect would be. Effect sizes from correlational research run higher than those from randomized trials. Writing in the same journal in 2021, Matthew Kraft found that only 13 percent of nearly 2,000 effect sizes from randomized controlled trials reached 0.40 at all. A correlational 1.57 and an experimental 0.20 are not the same kind of number and should not be compared as though they were.

    Does collective teacher efficacy cause higher achievement, or the other way around?

    Nobody can answer that from a correlation, and the reverse story is at least as plausible. A staff whose students have improved for three running years has excellent evidence for believing in itself. That is probably why “evidence of impact on student outcomes” is the first enabling condition the researchers list — belief appears to be downstream of results as much as upstream of them. The defensible claim is that the two reinforce each other.

    What actually builds collective teacher efficacy?

    Five conditions, per Donohoo, Hattie and Eells: evidence that the work is moving student outcomes, high-trust collaborative structures, non-threatening evidence-based environments, social sensitivity among team members, and leaders who model psychological safety. Four of the five are structural, which means they belong to whoever controls the schedule and the tone rather than to an individual teacher.

    Can one teacher do anything about this, or is it an administrator’s job?

    Most of the conditions are structural, but the smallest working version needs no permission: pick the unit you teach worst, find a colleague who teaches it better, and ask them how they run it. Add a shared piece of student work and a recurring thirty minutes and you have a two-person version of the whole thing. What a teacher should not do is carry the structural part alone — that is a specific, answerable request to make of an administrator.

    Why does collective efficacy work fail in so many schools?

    Usually because staff are asked to collaborate openly while being evaluated on what the collaboration reveals. Once a common assessment becomes evidence about the teacher rather than about the learning, everyone brings their best class and the honest conversation stops. Collaboration time that gets eaten by announcements, mandated problems of practice nobody actually has, and treating a failed lesson as a performance issue all do the same damage.

    What is Hattie’s 0.40 hinge point and should I use it?

    It is the line in Hattie’s rankings above which an influence is said to be worth having, and it was calculated as an average across his whole database. Use it carefully. Kraft’s analysis suggests 0.40 sits near the top of what well-controlled education research produces, and he proposes separate benchmarks for causal studies: below 0.05 small, 0.05 to 0.20 medium, above 0.20 large. Hattie’s own reply is that 0.40 was meant as a starting point for discussion rather than a pass mark, which is a fair description of what he wrote and a poor description of how it usually gets used.

    Sources

    1. Donohoo, J., Hattie, J., & Eells, R. (2018). “The Power of Collective Efficacy.” Educational Leadership, 75(6). Source of the 1.57 effect size and the comparison figures for prior achievement (0.65), socioeconomic status (0.52) and home environment (0.52), and of the five enabling conditions. https://www.ascd.org/el/articles/the-power-of-collective-efficacy
    2. Kraft, M. A., & Hattie, J. (2021). “Interpreting Education Research and Effect Sizes.” Educational Leadership. Kraft: effect size is “such a misleading term” and “tells us nothing about whether the underlying relationship represents cause and effect or a simple correlation”; only 13 percent of nearly 2,000 effect sizes from randomized controlled trials reached 0.40; proposed causal benchmarks of <0.05 small, 0.05–0.20 medium, >0.20 large. Hattie’s reply that 0.40 is an average and a starting point appears in the same article. https://www.ascd.org/el/articles/interpreting-education-research-and-effect-sizes

    A note on what is not cited here. Learning Forward’s 2015 article on collaborative expertise is frequently recommended alongside this topic; it sits behind a membership wall and could not be read, so it is not cited. The 1.57 figure is reported as its authors report it, and the limits described in this article are the limits its own co-author published three years later — not an outside attack on it.


    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.