Tag: secondary teaching

  • Questions to Ask at a Parent Teacher Conference in Middle and High School

    Questions to Ask at a Parent Teacher Conference in Middle and High School

    By Clay Shumate

    The best questions to ask at a parent teacher conference in middle or high school are the ones that produce something your kid can act on by Monday. Not “how is he doing,” which gets you “he’s doing fine.” Ask what specifically is going wrong, what it would look like fixed, and who is responsible for which part.

    Most conference question lists online were written for the parent of an eight-year-old with one teacher. At thirteen and sixteen the event is different, and the research on what parent involvement does at this age points somewhere those lists do not go.

    Key Takeaways

    • The involvement that correlates best with achievement in middle school is not the kind most conferences produce. In a meta-analysis of 50 studies, academic socialization — helping a teenager connect schoolwork to where they are going — had the strongest positive association. Homework help had the strongest negative one.
    • Ask for specifics, because vague questions get vague answers. “What does a 4 look like on the next project?” is a different question from “how’s she doing?” and it is the only one of the two a teacher can answer usefully in fifteen minutes.
    • Bring your kid. In a four-school study of student-led conferences, parent attendance was at least 92% at every site. Whatever else changes, more families show up when the student is running the meeting.
    • What happens after the conference is what decides whether it mattered. A randomized trial of weekly one-sentence teacher-to-parent messages cut the share of high school students failing to earn credit by 41%. Fifteen minutes in October does not do that on its own.
    • Specific, improvement-focused messages beat warm general ones. In the same trial, messages about what a student needed to improve produced a significant gain. Messages about what the student was doing well did not.

    Free Download · PDF

    The Parent-Teacher Conference Question Card (2 pages)

    Page 1 is the ten questions with what each one is for, the four to ask when the news is bad, and the five that waste the slot. Page 2 is a notes grid for up to seven teachers, with a column for the follow-up date you should not leave without.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Two-column graphic comparing forms of parent involvement in middle school: academic socialization and school-based involvement listed as positive, home-based involvement as not statistically significant, and homework help as the strongest negative association
    Hill and Tyson’s meta-analysis of 50 studies, sorted by what actually tracks with achievement.

    What Should You Actually Ask at a Parent Teacher Conference?

    Ask four things: what the specific problem is, what fixed looks like, what the student does next, and how you will both know whether it worked. Everything else is a variation on those four.

    The reason to be that blunt is arithmetic. A secondary conference is ten to fifteen minutes, and a teacher with five sections sees north of a hundred students a day. Teachers are not being evasive when they say “she’s doing great.” They are answering the question you asked.

    Why the Middle and High School Conference Is a Different Event

    You now have six or seven teachers instead of one, less time with each, and a student old enough to have a position on being talked about. Those three changes break most of the standard advice.

    In elementary school one adult sees your child all day and can describe them whole. In seventh grade nobody has that view. Each teacher sees your kid for fifty minutes in one subject, often in one mood. That is the structure, not a flaw in the teachers — but it means six teachers asked “how is she doing” will give you six fragments and no picture.

    So decide in advance what you are trying to find out and ask the same focused question in every room. Patterns show up fast. Behind in one class is a problem with that class. Behind in five is a problem with time, sleep, or something nobody has named yet.

    What Does the Research Say Parent Involvement Should Look Like at This Age?

    It says the most effective thing a parent can do with an adolescent is help them see the point. Nancy Hill and Diana Tyson’s meta-analysis of 50 studies of parental involvement in middle school, published in Developmental Psychology in 2009, found an average correlation of r = .18 between involvement overall and achievement. Underneath that average, the types split sharply.

    Academic socialization had the strongest positive association. That is the clumsy research name for an ordinary thing: talking with your kid about why school matters, what they want, what their plans require, and how the work in front of them connects to it. Communicating expectations, discussing learning strategies, linking coursework to goals.

    School-based involvement — volunteering, showing up, attending events — was positive and significant but smaller. Home-based involvement was not statistically significant. And helping with homework had the strongest negative association of anything they measured.

    Read that last one carefully, because it is easy to turn into a slogan it does not support. This is correlational, and parents help most with homework when a student is already struggling, so some of that negative number is parents responding to a problem rather than causing one. The honest reading is not “helping with homework hurts your child.” It is that there is no evidence more homework help is the lever, and a conference spent negotiating who checks the math packet is spent on the form of involvement with the least support behind it.

    Which gives you a sorting rule. A question that ends with you doing more of your teenager’s work is a weak question. A question that ends with your teenager understanding what the work is for is a strong one.

    Numbered list of six questions to ask at a parent teacher conference, each with a one-line explanation of what that question is for
    Pick four or five and ask the same ones in every room so the answers can be compared.

    The Ten Questions, and What Each One Is For

    Pick four or five — you will not get through ten — and ask the same ones in every room so you can compare.

    Ask thisWhat it is actually for
    “What is the one thing you’d change about how she works in your class?”Forces one specific answer instead of a summary. Teachers almost always have one.
    “Is this a skill problem or a habit problem?”Separates “doesn’t understand it” from “isn’t doing it.” The fixes have nothing in common.
    “What does strong work look like on the next big thing?”Gets the standard in concrete terms your kid can aim at.
    “What would you want him to do differently tomorrow?”Converts a judgment into an action with a date on it.
    “When is the best time for her to come see you?”Hands the relationship to the student, which is the point at this age.
    “What is he like in class when I’m not there?”You know your kid at home, not at 10:40 in a room of thirty.
    “Is anything I’m doing at home making this harder?”Rarely asked, often answerable, and it signals you are not here to litigate.
    “How will we both know in six weeks whether this worked?”Sets a measure before anyone has reason to argue about one.
    “What’s the fastest way to reach you, and how fast is realistic?”Prevents the two-week silence that turns into a grievance.
    “What should she be doing now if she wants that next year?”The academic socialization question — the one with research behind it.

    Write the answers down in the room. Not because you will forget — because a teenager reads a parent taking notes as evidence the meeting was real.

    What About the Homework Question?

    Ask about homework, but ask about the system rather than the assignment. “Where does she find out what’s due?” and “what happens when something comes in late?” are system questions. “Can you send me the list every week so I can check it” is an assignment question, and you have just volunteered to be your kid’s secretary for the rest of the year.

    The distinction tracks the research. Knowing how the class works is information your teenager can be handed. Doing the tracking yourself is the form of involvement with the weakest evidence behind it and the highest cost to the thing you are supposedly building, which is a kid who can run their own week — and every piece of that you take over is a piece they do not get to practice.

    Should Your Teenager Be in the Room?

    Yes, and it is the single highest-leverage change available to most families. Conferences held about a student who is standing in the hallway teach that student that their education is a negotiation between adults.

    The clearest evidence comes from a four-school program evaluation by Christine Tuinstra and Diana Hiatt-Michael, published in The School Community Journal in 2005, covering 524 middle school students and 523 parents across California, Oregon, Texas and Washington. Parent attendance reached at least 92% at every one of the four schools. Over 94% of students reported revising work and 90% reported setting academic goals.

    Be careful with the rest of it. This was a program evaluation with surveys and interviews, not a controlled trial, and there was no comparison group. Administrators reported higher test scores and fewer discipline problems, but those are reports rather than measured outcomes. The attendance figure is the hard number here. The satisfaction data says people liked it, which is worth something and is not proof it raised achievement.

    Even so, 92% is a result most schools would take — and bringing your own kid to your own conference requires no program at all. If your school runs student-led conferences, go. If not, bring your kid anyway and let them answer first. Hand them the opening question — “tell Mr. Reed what you think is going on in here” — and watch what the teacher does with it.

    Two exceptions, and they are real. Anything touching a disability, a mental health concern, or a formal plan belongs in a different meeting with different people, not a crowded cafeteria at 6pm — ask for it separately. And if a parent genuinely needs to say something without their kid hearing it, say it in the last two minutes after the student steps out, openly, rather than arranging to have the real conversation behind them.

    Give the student the right of reply. If a teacher describes something and the student thinks it is wrong, they get to say so in the room. Not to win — to be corrected, or to correct the record, which is what we claim to want from them everywhere else. A meeting where a sixteen-year-old sits silently while two adults characterize them is not teaching accountability. It is teaching that accountability happens to you.

    List of five parent teacher conference questions that waste time, each with the reason it produces no useful answer
    Each of these gets you a sentence you already knew.

    What Do You Ask When the News Is Bad?

    Ask what the recovery path is, and ask it before you ask anything about blame.

    Bad news usually arrives as a number — a 58, four missing labs, a referral. The number is not the useful part. These four questions are, in this order:

    1. “What is recoverable and what isn’t?” Some of it can be fixed; some is arithmetic that has already happened. Know which before you plan anything.
    2. “What is the first thing, and when is it due?” One task with a date. A list of eleven is a list a struggling teenager will not start.
    3. “What do you need from me, specifically?” Sometimes it is a quiet place and a bedtime. Sometimes it is nothing. Let the teacher say.
    4. “What do you need from him?” Asked with the student in the room, out loud, so the agreement has a witness.

    Then stop. Spending the last six minutes on whose fault it is has never once changed a grade. And if the real issue is that your kid has decided the class is pointless, that belongs to the longer conversation about a student who appears not to care, not to fifteen minutes in a cafeteria.

    What If You Can’t Get There?

    Then it happens on the phone, and that is not a lesser version of the meeting. Conference night is scheduled for the convenience of the building, not of a family working second shift, and a parent missing from the sign-in sheet has not declined to be involved.

    Email the two teachers who matter most this term, ask for ten minutes by phone, and send the same focused question you would have asked in person.

    If English is not the language you want this conversation in, ask the school for an interpreter, in writing. Districts receiving federal funds are required to communicate with families in a language they understand. A bilingual student translating a conversation about their own grades is not meaningful access and puts a teenager in an impossible position. Schools generally arrange it when asked; the asking is the part that often does not happen.

    For the school side: if attendance at your conference night is 30%, the answer is not a sterner reminder letter. It is finding out what the other 70% are doing between four and eight on a Thursday.

    What Teachers Should Do With Fifteen Minutes

    First, a reality check. Plenty of secondary schools run conference night arena-style — every teacher at a table in the gym, a line at each one, ninety seconds a family. The structure below assumes a scheduled slot. For the arena, cut it to three sentences: one specific thing going well, one specific problem with a number attached, one next step and a date. Then hand over your email and mean it.

    Second, for whoever runs the building: if a conference in October is the first a family has heard about a problem that started in August, the conference is not the failure. The communication system is. A conference should confirm what a parent already knows, not break news.

    Open with one specific observation, not a summary. “She’s a pleasure to have in class” costs you thirty seconds and tells a parent nothing. “She writes a better argument than almost anyone in that period and she has turned in four of seven” tells them everything, and it takes the same thirty seconds.

    A workable structure for a short conference:

    • Minute 1 — one thing genuinely going well, named specifically. Not flattery. Something you could show them.
    • Minutes 2–4 — the one problem, with evidence in front of you. A gradebook screen, a piece of work, a count. One problem, not five.
    • Minutes 5–9 — the student talks. Ask what they think is happening and let the silence sit.
    • Minutes 10–13 — the plan. One action for the student, one for you, one for the family if there is one. Dates on all of them.
    • Minutes 14–15 — how you will follow up, and when. Then actually do it.

    The preparation that makes this work across a whole evening is a student self-assessment filled out before the conference. It gives the student something to say and means the first voice in the room is not yours.

    Two lines not to cross. Do not discuss another student, even when the parent in front of you has clearly been told a name at home — “I can’t talk about anyone else’s child” is a complete sentence and parents accept it. And do not improvise a plan that belongs to a formal process. If a parent asks for something requiring a 504, an IEP, or an administrator’s signature, say it is a real question that needs the right meeting, then make sure that meeting gets requested. Promising something at a table in a gym that you cannot deliver in a classroom does more damage than saying you do not know.

    What Happens After the Conference Decides Whether It Mattered

    The follow-up is where the measurable effect is, and there is a randomized trial that says so.

    Matthew Kraft and Todd Rogers ran a blocked randomized experiment on 435 high school students in a five-week summer credit recovery program in a large urban Northeastern district, published in Economics of Education Review in 2015. Teachers sent parents a brief individualized message each week. Students whose parents got messages were 6.5 percentage points more likely to earn course credit — 84.2% to 90.7% — a 41% reduction in the share failing to earn credit. The gain came mostly from preventing dropout rather than from improving work.

    The part worth sitting with: they tested two kinds of message. Messages about what the student needed to improve produced an 8.8 percentage point gain and were statistically significant. Messages about what the student was doing well produced 4.5 points and were not. The authors attribute the difference to actionability — an improvement message tells a parent something specific to do.

    Two limits, plainly. This was a summer credit recovery population — students already behind, 58% African-American and 32% Hispanic, over 80% free or reduced-lunch eligible. Lots of room to move. Do not expect a 41% reduction in your honors section. And it is evidence for weekly contact, not for one conference. The conference’s job is to set up the contact that follows it, which is the argument for an actual parent-teacher communication plan rather than improvising one family at a time.

    So the last thing said in the room, by either party, is a date.

    Questions That Waste the Fifteen Minutes

    • “How is he doing?” You will get “good.” Every time.
    • “Does he behave?” If he did not, you would know. Ask what he is like when the work gets hard.
    • “Is she one of your better students?” Invites a comparison that helps nobody and that the teacher cannot answer honestly anyway.
    • “Can he get extra credit?” Asked before the real work is finished, this substitutes volume for learning. Ask what unfinished work can still be completed first.
    • Anything that is really about a different teacher. Take it to that teacher or to an administrator.

    What to Do Next

    Before the next conference: pick the one question you most need answered and write it on your phone, ask your teenager what they think each teacher will say, and bring them with you.

    If you are the teacher, use the structure above, put a gradebook on the screen, and let the student talk in the middle of it. Set the follow-up date out loud, because the evidence says the five minutes a week after October does more work than the fifteen minutes in it. The wider version of that argument is in the guide to effective parent-teacher communication, and what a conference should hand a student at the end is a goal they wrote themselves — which is what student goal setting is for.

    Before you go: grab the free The Parent-Teacher Conference Question Card (2 pages) (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    What are the best questions to ask at a parent teacher conference?

    The four that do the most work are: what is the specific problem, what does fixed look like, what does my student do next, and how will we both know in six weeks whether it worked. Everything else is a variation. Vague questions get vague answers — "how is he doing?" reliably returns "good." Pick four or five specific ones and ask the same ones of every teacher so you can compare instead of collecting impressions.

    How many questions can I realistically ask in a middle or high school conference?

    Four or five. A secondary conference is usually ten to fifteen minutes, and plenty of schools run them arena-style in a gym with ninety seconds per family and a line behind you. Decide in advance what you are trying to find out and ask that. If you need more time, ask for a phone call rather than trying to extend the slot while five families wait.

    Should my teenager come to the parent teacher conference?

    Yes, and it is probably the single highest-leverage change available to a family. A four-school evaluation of student-led conferences covering 524 middle school students found parent attendance of at least 92% at every site. Even without a formal program, bringing your kid and letting them answer first changes the meeting from something adults negotiate about them into something they participate in. Two exceptions: anything touching a disability, a mental health concern, or a formal plan belongs in a separate meeting.

    What should I ask if my child is failing a class?

    Four questions, in this order. What is recoverable and what is not? What is the first thing and when is it due? What do you need from me, specifically? And — asked with the student in the room — what do you need from him? Then stop. Spending the last six minutes on whose fault it is has never changed a grade.

    Is it true that helping with homework hurts my child’s grades?

    That overstates it. In Hill and Tyson’s 2009 meta-analysis of 50 studies of middle school parent involvement, homework assistance showed the strongest negative association with achievement of anything they measured — but the design is correlational, and parents help most with homework when a student is already struggling. The defensible reading is not that helping hurts. It is that there is no evidence more homework help is the lever, and that the involvement with the strongest positive association was academic socialization: talking with your teenager about why the work matters and what their plans require.

    What if I can’t attend conference night?

    Then it happens by phone, and that is not a lesser version of the meeting. Email the two teachers who matter most this term, ask for ten minutes, and send the same focused question you would have asked in person. If English is not the language you want this conversation in, ask the school for an interpreter in writing. Districts receiving federal funds are required to communicate with families in a language they understand, and a bilingual student translating a conversation about their own grades is not meaningful access.

    What should I ask when everything seems fine?

    Ask the forward-looking one: what should she be doing now if she wants that next year? That is the academic socialization question, and it is the form of involvement with the strongest research behind it at this age. Also worth asking: what is he like in class when I am not there? You know your kid at home. You do not know your kid at 10:40 in a room of thirty.

    What should teachers do differently in a fifteen-minute conference?

    Open with one specific observation rather than a summary — "she writes a better argument than almost anyone in that period and she has turned in four of seven" costs the same thirty seconds as "a pleasure to have in class" and tells a parent everything. Then one problem with evidence on the screen, the student talking in the middle, a plan with dates, and a stated follow-up. And if a conference in October is the first a family has heard of a problem that started in August, the conference is not the failure. The communication system is.

    Sources

    1. Hill, Nancy E., and Diana F. Tyson. “Parental Involvement in Middle School: A Meta-Analytic Assessment of the Strategies That Promote Achievement.” Developmental Psychology, vol. 45, no. 3, 2009, pp. 740–763. Meta-analysis of 50 empirical studies. Overall involvement r = .18. Academic socialization showed the strongest positive association; school-based involvement positive and significant; home-based involvement not significant; homework assistance showed the strongest negative association. The authors note the correlational design limits causal inference. https://courses.edx.org/assets/courseware/v1/afb13fe949b60e70e83a2d18c32dacc0/asset-v1:HarvardX+GSE4x+1T2020+type@asset+block/ParentalInvolvementinMiddleSchoolAMeta-Analysis.pdf
    2. Kraft, Matthew A., and Todd Rogers. “The Underutilized Potential of Teacher-to-Parent Communication: Evidence from a Field Experiment.” Economics of Education Review, vol. 47, 2015, pp. 49–63. Blocked randomized trial, 435 students in a five-week high school summer credit recovery program, large urban district in the Northeastern United States. Credit earned rose from 84.2% to 90.7%, +6.5 percentage points, a 41% reduction in the share failing to earn credit, driven primarily by reduced dropout. Improvement-focused messages +8.8 points (significant); positive messages +4.5 points (not significant). Population was students already behind; 58% African-American, 32% Hispanic, over 80% free/reduced-lunch eligible. https://scholar.harvard.edu/files/todd_rogers/files/empirical_in_press.kraft_rogers.pdf
    3. Tuinstra, Christine, and Diana Hiatt-Michael. “Student-Led Parent Conferences in Middle Schools.” The School Community Journal, vol. 15, no. 1, 2005, pp. 59–80. Mixed-methods program evaluation at four middle schools in California, Oregon, Texas and Washington; 524 students (mostly 7th grade), 523 parents, 30 teachers, 7 administrators. Parent attendance at least 92% at all four sites. Over 94% of students reported revising work; 90% reported setting goals. No control group; the achievement and discipline claims are administrator reports rather than measured outcomes. https://www.adi.org/journal/ss04/Tuinstra & Hiatt-Michael.pdf
    4. Harvard Family Research Project. “Parent–Teacher Conferences: A Tip Sheet for Parents.” October 2010. Distributed by state education agencies including the Texas Education Agency and the Nebraska Department of Education. Practical guidance, not research: review work and progress reports beforehand, bring a written list of questions, treat the meeting as a two-way conversation, record action items for both sides, schedule a follow-up, and share the outcome with the child — “don’t forget to include him or her.” https://education.ne.gov/wp-content/uploads/2017/07/Parent-Teacher-ConferenceTipSheet-100610-2.pdf
    5. Weiss, Heather B., quoted in “A New Look at the Parent-Teacher Conference.” Usable Knowledge, Harvard Graduate School of Education, 2014. Argues conferences should be “accessible, understandable, and actionable” and should cover both strengths and difficulties using student data. Cited here as professional guidance: the piece reports an interview and does not present original research. https://www.gse.harvard.edu/ideas/usable-knowledge/14/09/new-look-parent-teacher-conference

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Late Work Policy: What the Evidence Supports After the Famous Study Was Retracted

    Late Work Policy: What the Evidence Supports After the Famous Study Was Retracted

    By Clay Shumate

    A late work policy decides two separate things: what happens to the grade, and what happens to the student. Most policies collapse those into one number and then stop working. The evidence does not hand you a clean answer, and the study most often quoted in these arguments was retracted in September 2026 — which is the first thing anybody writing a policy this year should know.

    Below is what the research actually supports, what it does not, and a six-part policy you can write on one page and still defend in a parent meeting.

    Key Takeaways

    • The famous deadline study is gone. Ariely and Wertenbroch’s 2002 paper on spaced deadlines was retracted on 2 September 2026. A 124-person replication found no effect of deadline condition on any outcome measure.
    • Loosening grading has not been shown to help students. Students assigned to stricter-grading teachers scored higher in math — in that class and in later ones — across every subgroup studied.
    • Removing late penalties has a measured cost. In one high school chemistry class, homework completion fell by more than a third.
    • The zero is a math problem before it is a policy problem. On a 100-point scale, the gap between passing grades is 10 points and the gap from D to F is 60.
    • The deeper issue is the scale and the averaging, not the zero. Research supports grading scales with four to seven levels for reliability; the 100-point scale invents precision that is not there.
    • Separate the grade from the behavior. Report lateness as conduct and let the grade report what the student knows. That one move resolves most of the argument.

    Free Download · PDF

    The One-Page Late Work Policy and Window Log (2 pages)

    Page 1 is the six-part policy to fill in, the practice-versus-assessment split, and the sentence to give a parent. Page 2 is the window log that turns the policy into information, plus the honest cost of all three common policies.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Why Is a Late Work Policy So Hard to Get Right?

    Because a grade is being asked to do two jobs at once, and they conflict.

    Job one is reporting what a student knows and can do. Job two is enforcing a deadline. A 10-percent-per-day deduction does job two by corrupting job one: after five days the number on the report card is half knowledge and half calendar, and nobody reading it — not the next teacher, not a parent, not the student — can tell which half is which.

    A zero does job two harder and job one worse. And the people on the other side of the argument are not making a soft case. The claim that deadlines do not matter has a measured cost attached, which is the part that usually goes missing in articles about this.

    So the honest version of the problem is not “should kids face consequences.” It is: how do you keep the deadline real without making the grade lie?

    Four evidence findings on late work policy: stricter grading linked to higher math scores and a one-third drop in homework completion when late penalties were removed, a small undergraduate trial where an early-bonus plus late-penalty policy produced work 1.45 days early, the arithmetic problem of a zero on a 100-point scale, and the September 2026 retraction of the Ariely and Wertenbroch deadline study
    Four findings that do not all point the same way. That is the honest picture.

    What Does the Research Actually Say About Deadlines?

    Less than you have been told, and one widely quoted finding has been formally withdrawn.

    For twenty years, the standard citation for “students do better with evenly spaced interim deadlines” has been Ariely and Wertenbroch (2002) in Psychological Science. You have probably read it quoted in a PD slide deck. It was retracted on 2 September 2026.

    The retraction is not a technicality. Data Colada’s analysis of Study 2 found eighteen of twenty participants in one condition had exact duplicates across all three tasks, correlations that should have been strong were absent, and self-reported times showed almost none of the rounding that real human responses show. Their conclusion was that the data were “severely tampered with or fabricated.” Coauthor Klaus Wertenbroch stated that “much or all of the data — and therefore the results — are false.”

    A pre-registered replication by Hyndman and Bisin, published in 2025 with 124 participants across three deadline conditions, found no statistical evidence that performance was influenced by the deadline condition on any of three measures — errors found, days late, or payment, all p > 0.1. The original had reported all differences significant at p < 0.01. The replication authors concluded that the received wisdom about spaced deadlines limiting procrastination is “possibly false.”

    Two things follow. First, if your department’s late work policy was built on that study, it needs a different foundation. Second — and this is the part worth holding onto — the retraction says nothing about whether deadlines matter in a classroom. It says one famous laboratory result about self-imposed versus imposed deadlines cannot be relied on. Interim checkpoints on a long project may still be good practice. They just are not evidence-backed in the way everyone has been saying.

    Does Removing Late Penalties Hurt Students?

    There is real evidence that looser grading costs something, and it deserves to be stated as plainly as the case for reform usually is.

    Gershenson’s work on high school math found that students with tougher-grading teachers scored higher in math, both in that teacher’s class and in subsequent math courses, and that this held across every student subgroup. Figlio and Lucas found the same direction at elementary level: students assigned to stricter graders showed greater test-score growth in reading and math. In one high school chemistry class where late penalties were removed, homework completion dropped by more than a third.

    The strongest single trial on late policies specifically is small and not from a high school. Korpusik, Freitas and Dionisio compared four late policies across 248 lab submissions in an introductory programming course. Work came in 2.43 days late under no policy, 4.71 days late under an early-completion bonus alone, 0.44 days late under a late penalty, and 1.45 days early under a combined early bonus and late penalty — which also produced the highest grades. Their read was that late penalties externally regulate students who have not yet regulated themselves.

    Size that evidence honestly before you use it: 31 students who consented to the analysis, mostly first-years, taught online during the pandemic. It is a signal, not a mandate. But it points the same direction as the grading-strictness work, and a teacher who ignores both because the conclusion is unfashionable is doing the thing this site complains about when it goes the other way.

    Does It Matter What the Assignment Was For?

    More than any other question here, and most policies never ask it.

    If the work was practice — a problem set, a draft, a reading check whose job was to tell you what to reteach — then late is a timing failure with a real instructional cost, because the information arrived after you needed it. The honest response is to collect it, use it, and record the lateness as conduct. Scoring it down does not recover the information.

    If the work was the assessment — the essay, the project, the unit test — then the deadline is doing something different. It is the point at which you are claiming to know what the student can do. Here a window still makes sense, but a narrow one, and after it closes the student demonstrates the standard some other way rather than handing in the same artifact in December.

    Running one blanket rule over both is why so many late work policies feel wrong in practice. A missing formative check and a missing final project are not the same event and should not get the same sentence.

    So Why Not Just Give Zeros?

    Because of arithmetic, not sentiment.

    Reeves laid this out in Phi Delta Kappan. On a standard 100-point scale with letter grades at 10-point intervals, every passing grade sits 10 points from its neighbor — and the interval between D and F is not 10 points but 60. A single zero therefore carries roughly six times the weight of any other failing mark in an average. Reeves’s point is that if an F is one interval below a D, the mathematically consistent value is 50, not 0. On a 4-point scale nobody hesitates: missing work gets a 0, exactly one point below a 1. The same logic on a 100-point scale would require a –6, which no one would give.

    Guskey, Fisher and Frey take that further in Educational Leadership, and their version is the one to carry into a faculty meeting: “The true problem is not the zero; it’s the use of the 100-point percentage grading scale and the practice of averaging scores.” They point out that the problems with percentage scales have been documented since Starch and Elliott in 1913, and that research supports scales with four to seven levels for optimal reliability and discrimination. A 101-level scale manufactures precision nobody can actually defend.

    That reframing matters for a late work policy because it tells you where the real fix is. Arguing about whether a missed assignment is a 0 or a 50 is arguing about a symptom. If your gradebook averages percentages across a semester, a single missed assignment distorts the picture no matter which number you put in the box.

    If this is live at your school, the fuller version of that argument is in the standards-based grading guide, and the honest case against it is in why standards-based grading doesn’t work. Both are worth reading before anyone proposes a schoolwide change.

    Comparison of three common late work policies with honest costs: a zero after the due date, ten percent off per day late, and full credit with no deadline, each listed with what it is honest about and what it costs
    None of these is free. The question is which cost you can live with.

    What Does a Workable Late Work Policy Look Like?

    Six parts. It fits on one page, and every part of it survives a parent asking why.

    Six numbered components of a workable late work policy: state what the deadline is for, set a hard floor instead of a zero, separate the grade from the behavior, publish one window and hold it, make the make-up cost time rather than points, and record which students keep using the window
    Six parts, and the sixth one turns the policy into information.

    1. Say what the deadline is for

    “This is due Friday because on Monday we build on it” is a reason. “Because I said Friday” is not, and teenagers are unusually good at telling the difference. A deadline with a downstream purpose gets taken seriously by more students than a deadline without one, and it also tells you which deadlines you should actually defend. If nothing depends on Friday, Friday was arbitrary and you should stop pretending otherwise.

    2. Set a hard floor, not a zero

    A missed assignment scores the bottom of the scale, not the bottom of the number line. On a four-point scale that is a 0. On a 100-point scale it is a 50. The point is not generosity; it is that the floor should be one interval below the lowest passing mark, which is what every other grade boundary already is.

    3. Separate the grade from the behavior

    This is the move that resolves most of the argument, and it costs nothing. Lateness is a conduct fact. Report it as one — a comment on the report, a contact home that happens the same week, a scheduled session — and let the grade report what the student knows. A parent who is told “she understands this material and she has turned in four assignments late” has two usable pieces of information. A parent told “she has a 61” has none.

    4. Publish one window and hold it

    “Late work is accepted until the unit assessment” is a rule a fourteen-year-old can plan against. “Depends when you ask me” is not a policy, it is a mood, and students read inconsistency as unfairness faster than they read strictness as unfairness. Pick a window, write it in the syllabus, and then do exactly what you said — which is the whole of respecting students in practice rather than on a poster.

    5. Make the make-up cost time, not points

    Being late should cost something. Time is the honest currency: a scheduled session at lunch, before school, or during an intervention block. It is a real cost, students feel it, and it leaves the grade intact. A ten-percent-per-day deduction costs them something too — the accuracy of the transcript.

    6. Write down who keeps using the window

    Keep a list. Not to punish with — to read. Three students using the late window every single time is not a late work problem, it is a signal about workload, home, organization, or something nobody has asked about yet. A policy that generates that list is doing a second job for free. A policy that just applies a deduction tells you nothing you did not already know. This is the same logic as treating attendance data as a referral system rather than a compliance record. The same read applies to arrival times, where a tiered response sorts the one-off from the pattern in a way a blanket deduction never will.

    One cost this policy does carry, and it should be named rather than buried: an open window produces a flood at the end of it. If the window closes at the unit assessment, expect a stack the night before, and expect it to arrive in the same week you are marking the assessment itself. Two things keep that survivable. Cap what comes back — late work gets a score and a one-line comment, not the full written feedback a punctual draft gets, and say so in advance. And put the make-up session mid-unit rather than at the end, so the work trickles in instead of arriving at once. A policy that quietly doubles your marking in the last week of a unit is a policy you will abandon by Thanksgiving, which is its own kind of unfairness. The workload side of this is not a side issue; it decides whether the policy still exists in March.

    What Does This Look Like From the Student’s Side?

    Clear, and askable in advance. Those are the two things students actually want from a late work policy, and neither is the same as lenient.

    Clear means they can find it. Not buried in a syllabus they signed in August — posted, in nine words, where the due dates are. A student should be able to answer “what happens if I turn this in Monday” without asking you.

    Askable in advance means the policy has a front door. A fifteen-year-old working a closing shift, or watching younger siblings, or without reliable internet at home, usually knows on Tuesday that Friday is not going to happen. Right now most classrooms give that student nothing to do with that knowledge except apologize on Friday. Say out loud, more than once, that asking before a deadline is a different conversation than explaining after one — and then make it true, which means the student who asks on Tuesday gets a straight answer rather than a lecture.

    That is not softness. It is the difference between treating a teenager as someone managing competing obligations, which they are, and treating them as someone who needs catching out. The deadline does not move for the student who never asks. It is just that the student who plans ahead gets something for planning ahead, which is the behavior the whole policy claims to be teaching.

    How Do You Explain This to a Parent Who Thinks It Is Too Soft?

    Lead with what did not change, because something did not.

    The work is still required. The deadline is still real. There is still a consequence, and it is time rather than points. What changed is that the report card now tells you what your child knows instead of telling you a blend of what they know and when they handed it in.

    That sentence survives most objections because it is not a concession, it is a clarification. And it is worth having ready in writing — in the syllabus and in the first message home — rather than improvised on the phone in October. Parents who object to no-zero policies are usually objecting to the version where nothing is required; the fastest way to settle it is to show this is not that version.

    There is a fair version of the objection too, and it should not be waved off. If a student can turn anything in whenever they like, some will discover that and use it, and a policy that pretends otherwise is not being honest. That is precisely why parts 4, 5 and 6 exist: a published window, a real cost in time, and somebody actually watching the list.

    What Does This Mean for a Department or a School?

    Consistency across a hallway matters more than the specific policy chosen.

    A student with six teachers running six different late policies cannot plan, and the student least able to absorb that is the one the policy was supposed to help. If a department can agree on a window and a floor, that is worth more than any individual teacher’s preferred version of either.

    Two cautions for anyone proposing this above the classroom level. A policy that requires teachers to accept unlimited late work without additional marking time is a workload decision disguised as a grading decision, and it will be abandoned by March. And a schoolwide floor applied on top of a 100-point averaging gradebook is, by Guskey, Fisher and Frey’s argument, treating the symptom — worth doing, but not worth calling a reform.

    Two constraints sit above everything on this page, and both of them outrank it.

    Your district may already have a grading policy with the force of board approval. A teacher who unilaterally sets a 50 floor in a district whose policy says otherwise has a problem that has nothing to do with pedagogy. Read the policy before writing yours, and if the two conflict, that is a conversation with an administrator rather than a decision to make quietly in a gradebook.

    And an IEP or 504 plan governs. If a student’s plan includes extended time or modified deadlines, the plan is the policy for that student, full stop — no classroom rule and no research finding on this page overrides it. Build your policy so that honoring a plan looks like the ordinary case rather than a visible exception, which mostly means making the window generous enough that nobody has to be singled out to use it.

    What to Do Next

    Write the policy on one page before the next unit starts. One window. One floor. One sentence saying lateness is reported as conduct, not deducted from the grade. One line about what the make-up session is and when it runs.

    Then put it in the syllabus and in a message home the first week, so that the first time a family hears about it is not after a missed assignment.

    And keep the list. At the end of the first unit, look at who used the window and how often. That list will tell you more about your classroom than the policy itself does.

    Before you go: grab the free The One-Page Late Work Policy and Window Log (2 pages) (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Is it true that the study everyone quotes about deadlines was retracted?

    Yes. Ariely and Wertenbroch’s 2002 paper in Psychological Science on spaced deadlines and procrastination was retracted on 2 September 2026, after analyses found duplicated observations, missing expected correlations and implausible response patterns in Study 2. One of the coauthors stated that “much or all of the data — and therefore the results — are false.” A 2025 replication with 124 participants found no significant effect of deadline condition on any outcome. If a PD session or a department policy still cites that study, it needs a new basis. It does not mean deadlines are useless — it means that one famous laboratory result cannot be used as evidence.

    Is giving a 50 for work that was never done just a gift?

    It is a scale decision, not a generosity decision. On a 100-point scale every passing grade sits 10 points from the next, and the gap from D to F is 60 — so a zero carries about six times the weight of any other failing mark in an average. A floor at 50 makes the F one interval below the D, which is what every other boundary already is. The work is still missing, the student still has not demonstrated the standard, and the grade still shows a failure. What changes is that one missing assignment can no longer outweigh several completed ones.

    Can a teacher set a 50 floor if the district grading policy says otherwise?

    No, and this should be checked before anything on this page is implemented. Many districts have a board-approved grading policy, and a teacher who quietly overrides it in a gradebook has created a problem that is not pedagogical. Read the policy first. If it conflicts with what you believe is right, that is a conversation to have with an administrator, and the argument from grading-scale reliability is a strong one to bring to it — but it is a conversation, not a unilateral change.

    How does a late work policy interact with an IEP or 504 plan?

    The plan governs, without exception. If a student’s plan provides extended time or modified deadlines, that is the policy for that student and no classroom rule overrides it. The practical design point is to build the general policy so that honoring a plan looks ordinary rather than exceptional — a window generous enough that a student using it is not visibly marked out. If you are unsure how a plan applies to a specific assignment, ask the case manager before the deadline rather than after.

    Should practice work and assessments have the same late policy?

    No, and running one rule over both is why many late policies feel wrong in practice. Practice work — problem sets, drafts, reading checks — exists to tell you what to reteach, so when it arrives late the instructional cost is already paid and scoring it down recovers nothing. Collect it, use it, record the lateness as conduct. An assessment is different: the deadline is the point at which you claim to know what a student can do. Keep a window there too, but a narrow one, and after it closes have the student demonstrate the standard another way rather than submitting the same artifact months later.

    Won’t accepting late work bury me in grading?

    It can, and a policy that ignores that will not survive to March. Two things keep it manageable. Cap what comes back: late work gets a score and a one-line comment rather than the full written feedback a punctual draft earns, and say so in advance so it reads as a stated rule and not as neglect. And schedule the make-up session mid-unit rather than at the window’s close, so work trickles in instead of landing as a stack the night before the unit assessment.

    Does removing late penalties actually hurt students?

    There is real evidence that it costs something, and it should not be waved away. In one high school chemistry class, homework completion fell by more than a third when late penalties were removed. Separately, students assigned to tougher-grading teachers scored higher in math — in that class and in later courses, across every subgroup studied. None of this proves zeros are correct, and none of it is a randomized trial of a specific late policy. What it does say is that “no deadline, no consequence” is a position with measured costs, which is why the policy described here keeps a real cost and moves it from points to time.

    What should a student do if they know in advance they cannot meet a deadline?

    Ask before it, not explain after it — and the teacher’s job is to make that worth doing. A student working a closing shift or caring for siblings usually knows on Tuesday that Friday will not happen. Most classrooms give them nothing to do with that information. Say out loud, more than once, that a request made before a deadline gets a different conversation than an apology made after one, then honor it: a straight answer rather than a lecture. The deadline does not move for the student who never asks. Planning ahead is the behavior the policy claims to be teaching, so it should get something.

    Sources

    1. Retraction notice: Ariely, D., & Wertenbroch, K. (2002), “Procrastination, Deadlines, and Performance: Self-Control by Precommitment,” Psychological Science. Retracted 2 September 2026. The notice cites a replication study and Data Colada analyses that “raised questions about the underlying data” and “called into question the veracity of the overall findings.” Coauthor Wertenbroch is quoted: “much or all of the data — and therefore the results — are false.” https://retractionwatch.com/?p=135929
    2. Data Colada, post 138, on Study 2 of Ariely & Wertenbroch (2002). Eighteen of twenty participants in the Last Day Deadline condition had exact duplicates across all three proofreading tasks, with ID numbers ten positions apart; expected correlations were absent (original task-performance correlations +.03 to +.27 against +.74 to +.90 in replication); self-reported times showed 11.7% rounding against 85% in the replication. Conclusion: the data “were severely tampered with or fabricated.” https://datacolada.org/138
    3. Hyndman, K., & Bisin, A. (2025). Replication of Ariely & Wertenbroch (2002). 124 participants, three randomly assigned deadline conditions (none, evenly spaced, self-imposed), three proofreading tasks over three weeks. No statistically significant effect of deadline condition on errors found (F = 0.181), days late (F = 0.353) or payment (F = 0.215), all p > 0.1. https://bpb-us-e1.wpmucdn.com/wp.nyu.edu/dist/c/16384/files/2025/09/Ariely_Replication-1.pdf
    4. Korpusik, M., Freitas, J., & Dionisio, J. D. N. (2022). “Impact of Late Policies on Submission Behavior and Grades.” ASEE Annual Conference. Loyola Marymount University, introductory programming lab, 248 submissions, 31 students consenting to analysis, mostly first-years, taught online during the pandemic. Average submission timing: no policy +2.43 days, early incentive +4.71, late penalty +0.44, combined −1.45 days (early). Grades 94.2% / 96.6% / 97.1% / 99.5% respectively. Undergraduates, not grades 6–12 — the mechanism may transfer, the numbers do not. https://people.csail.mit.edu/korpusik/asee22.pdf
    5. Guskey, T. R., Fisher, D., & Frey, N. “The Unwinnable Battle Over Minimum Grades.” Educational Leadership (ASCD). Argues minimum-grade floors treat a symptom: “The true problem is not the zero; it’s the use of the 100-point percentage grading scale and the practice of averaging scores.” Cites Starch and Elliott (1913) on percentage-scale problems and Lozano et al. (2008) and Preston and Colman (2000) on four-to-seven-level scales producing optimal discrimination, validity and reliability. https://www.ascd.org/el/articles/the-unwinnable-battle-over-minimum-grades
    6. Reeves, D. B. (2004). “The Case Against the Zero.” Phi Delta Kappan, 86(4). The arithmetic argument: on a 100-point scale the interval between passing grades is 10 points while “the interval between the D and F is not 10 points but 60 points,” making the mathematically consistent value of an F 50 rather than 0. https://www.researchgate.net/publication/285846142_The_Case_against_the_Zero
    7. Thomas B. Fordham Institute. “Think Again: Does ‘equitable’ grading benefit students?” Cited as a review, not as the primary studies. This is where the grading-strictness findings above come from: Gershenson (2020) on high school math students with tougher-grading teachers scoring higher in that class and in later courses across all subgroups; Figlio and Lucas (2004) on elementary test-score growth; and the high school chemistry case in which removing late penalties cut homework completion “by more than one-third.” The review’s own position is that there is no hard evidence that more lenient grading benefits students long term — a contested claim, and presented here as the strongest version of the case against loosening a late work policy. https://fordhaminstitute.org/national/research/think-again-does-equitable-grading-benefit-students

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Gallery Walk Activity: How to Run One So the Feedback Is Worth Reading

    Gallery Walk Activity: How to Run One So the Feedback Is Worth Reading

    By Clay Shumate

    A gallery walk activity posts student work around the room and sends classmates around to read it and leave written feedback. The walking is not what makes it work. The peer assessment underneath it is the part with research behind it, and that research says it only pays off when students are told exactly what they are looking for before they stand up.

    What follows is what the evidence actually supports, the one study at the right grade level and why it is weaker than it looks, and a six-step version that produces feedback worth reading instead of thirty sticky notes that say “good job.”

    Key Takeaways

    • Peer assessment has real evidence. The gallery walk format does not, separately. A meta-analysis of 54 control-group studies put peer assessment at g = 0.31 on achievement, and at secondary level specifically, g = 0.44.
    • Peer assessment beat teacher assessment in that analysis (0.31 versus 0.28) and tied with self-assessment. It is not a second-best substitute for your own marking.
    • Attaching a grade to peer feedback helped university students and did not help school students. Keep the walk ungraded in grades 6–12.
    • Feedback is not automatically good. Across 607 effect sizes, feedback raised performance on average but made it worse in over a third of cases — mostly when it pointed at the person instead of the work.
    • The single highest-leverage change is naming one focus before students circulate. Open-ended walks produce praise, not information.
    • If nobody revises anything afterward, you ran a walking tour. The revision block is the lesson, not the extra.

    Free Download · PDF

    The Gallery Walk Planner (2 pages)

    Page 1 plans the walk: the four decisions, the three sentence stems, the access check and the two rules that keep it fair. Page 2 runs it: the six steps in order, a student revision slip, and the four failure modes with their fixes.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    What Is a Gallery Walk Activity?

    It is a critique protocol: student work goes up on the walls or tables, students circulate in small groups, and each group leaves written feedback on what it sees. Then the authors read their feedback and change something.

    That last clause is the one teachers drop, and it is the one that turns the activity into instruction. Everything before it is logistics.

    The format is flexible in the ways that do not matter much and rigid in the one way that does. It works with drafts, lab write-ups, solved problems, design sketches, annotated sources, project prototypes. It works on paper taped to a wall or on a shared digital board. What it does not work without is a stated thing students are supposed to judge.

    Does a Gallery Walk Activity Actually Improve Learning?

    The peer assessment inside it does. The format itself has almost no rigorous evidence, and that is worth saying out loud because most articles about gallery walks imply otherwise.

    Double, McGrane and Hopfenbeck published a meta-analysis of peer assessment in Educational Psychology Review covering 54 experimental and quasi-experimental control-group studies. The overall effect on academic performance was g = 0.31 — small to medium, and statistically significant. Peer assessment outperformed no assessment and outperformed teacher assessment (g = 0.28). It showed no advantage over self-assessment.

    Two details in that analysis matter more for a grades 6–12 classroom than the headline number.

    First, the effect at secondary level was higher than the average: g = 0.44 across 13 studies. That runs against the usual pattern on this site, where secondary tends to be the weakest band in cooperative-learning research. Second, the effect held up across implementation choices — online or in person, anonymous or named, frequent or occasional, trained or untrained. None of those moderators came out significant. That is good news for a teacher deciding whether to buy clipboards: the fiddly choices are not where the result lives.

    There was one moderator that did split, and it splits against school students. Peer feedback that carried a grade significantly helped university students (g = 0.55). The same effect could not be shown for primary or secondary students. Do not grade the sticky notes.

    Three-part evidence summary comparing well-evidenced peer assessment at g equals 0.31 overall and 0.44 at secondary level, thin evidence for the gallery walk format itself from one Grade 8 study with no control group, and the finding that feedback made performance worse in over a third of cases
    Two different claims get mixed together. Only one of them is well supported.

    What About Research on Gallery Walks Specifically?

    Thin, and the one study at the right grade level found less than its abstract suggests.

    The closest match is a 2023 study of 198 Grade 8 students at a public high school in the Philippines, measuring knowledge, interest and attitude across three social studies lessons taught through gallery walk activities. It is a genuine secondary-level sample, which almost nothing else in this literature is.

    It is also a one-group pre-test/post-test design with no control group. Students gained on the researcher-made test — mean gained scores of 7.88, 8.15 and 8.01 across the three lessons — but with no comparison class, there is no way to separate the gallery walk from three weeks of teaching. More pointed: when the researchers correlated the specific features of how the walk was implemented against outcomes, none of the correlations with cognitive skills were significant, and none of the correlations with interest were significant either. Only attitude moved reliably.

    So the honest summary is: students liked it, and this study cannot tell you whether it taught them anything. Most of the remaining gallery-walk literature is small single-classroom work outside the United States, often in language instruction, and frequently without a control group either.

    The best-known practitioner guide, from PBLWorks, is also worth reading accurately. It describes a protocol and reports that giving participants a specific focus and criteria “helps generate higher-quality feedback,” while open-ended walks produced superficial comments like “good DQ!” That is a credible observation from people who run these constantly. It is not a study, and the article cites none. It should be quoted as what it is.

    None of this means skip the activity. It means the reason to run it is the peer assessment evidence, and the way to run it should be built to make that peer assessment real.

    Why Does Feedback Sometimes Make Students Worse?

    Because a lot of feedback is about the student instead of the work, and that redirects attention away from the task.

    Kluger and DeNisi’s review of feedback interventions remains the uncomfortable finding in this whole area. Across 607 effect sizes and 23,663 observations, feedback raised performance on average, d = 0.41. But in over a third of cases it lowered performance. Their explanation is that when feedback draws attention to the self rather than the task, the recipient spends their effort on defending themselves rather than improving the thing.

    Anybody who has watched a fifteen-year-old read a comment on their essay knows what that looks like. “This is confusing” is a verdict on a person. “I couldn’t find your claim in the second paragraph” is information about a page. The second one gets acted on; the first one gets argued with.

    That is also why the sentence stems matter more than they sound like they should. “I notice…” and “I wonder…” are not politeness training. They are a grammatical trick that forces the comment to describe the work. The EL Education framing that runs through a lot of critique practice — feedback should be kind, specific and helpful — is making the same move from a different direction, and the people who use it are explicit that it depends on an existing culture. As Ron Berger puts it in that work, this is not a strategy you drop into a classroom without a respectful one already in place.

    That is a real precondition, not a throat-clearing caveat. A gallery walk in a room where students are not already being treated decently by each other and by you produces insults on sticky notes, and you will spend the period on discipline instead of drafts.

    How Do You Run a Gallery Walk That Produces Useful Feedback?

    Six steps. Five of them cost nothing; the sixth costs ten minutes of class time and is the one that makes the other five worth doing.

    Six numbered steps for running a gallery walk activity: name the one thing students look for, hand out criteria before they walk, require written feedback, give descriptive sentence stems, put a clock on each station, and build in revision time
    None of these adds a planning period. The last one is the one that gets cut.

    0. Model it once, on one piece of work, before anybody walks

    Put a single piece of work up on the projector — ideally one of yours, or an anonymous sample from a previous year — and have the whole class write one note about it together. Then read three of their notes aloud and say plainly which one is useful and why. That takes eight minutes and it is the difference between a protocol students understand and a protocol they are performing. Skipping it is the most common reason a first gallery walk falls flat, and it is also the thing a coach watching your room will ask about first.

    1. Name one thing they are looking for

    Not “give feedback.” One focus: content accuracy, or evidence quality, or whether the claim is actually answerable. If you name three, you get one sentence about the easiest of the three. This single decision is what separates a walk that generates information from a walk that generates compliments.

    2. Hand over the criteria before they walk, not after

    Students should be holding the same rubric you will grade against while they circulate. Giving it to them afterward turns peer feedback into a guessing game about what you wanted. If you do not have a rubric for this task yet, three success criteria written on the board will do.

    3. Require it in writing

    Spoken feedback evaporates at the bell. Sticky notes work. So does a feedback slip per station, or a shared document with a row per project. The requirement is that the author can read it later, because step six depends on that.

    4. Give stems that describe rather than judge

    Post three and require them: I notice… / I wonder… / One thing that would strengthen this is… Teenagers will write “nice” if you let them, and they will write something sharper than you expect if you make the sentence start with a verb that forces specificity.

    5. Put a clock on each station

    Two or three minutes, then rotate on a signal. Without a timer the quick groups finish in forty seconds and the room drifts, which is a transition problem dressed up as a behavior problem. A posted rotation order prevents the traffic jam at station one. My Station Rotation Toolkit on TPT has eight reusable station frames, signs, and recording sheets if you would rather not build them.

    6. Build in the revision

    Ten minutes at the end: read your notes, pick one change, make it. Not “consider the feedback.” Pick one, make it, and be ready to say what you changed and why. This is the step that gets cut when the period runs long, and cutting it is what turns a critique into a field trip around your own classroom.

    One more thing belongs in that ten minutes, and it is the part students care about most: they do not have to take the advice. Some peer feedback is confidently wrong. A student who reads “your thesis is unclear” from someone who skimmed two sentences is entitled to decide that note is not useful. The requirement is that they can say which note they acted on and why, or which note they rejected and why. Both are evidence of judgment, which is the actual skill being taught. A protocol that forces a student to obey bad advice is teaching the opposite of what peer assessment is for.

    Whose Work Goes on the Wall?

    This is the question a parent asks first, and most guides to this activity never raise it.

    Putting a student’s work in front of thirty classmates is not a neutral act. For the student who knows their draft is the weakest in the room, a gallery walk can be the worst twenty minutes of their week, and no amount of “I wonder…” stems fixes that by itself.

    Three things keep it fair. First, everything goes up, every time — a walk where only the strong work is displayed is a showcase, and students read the selection as the verdict. Second, do not use personal writing. Narrative and reflective pieces where a student has written about their own life do not belong on a wall; use the analytical task, the lab, the problem set, the project artifact. Third, let a student put work up unnamed if they ask. Peer assessment effects in the meta-analysis held up whether feedback was anonymous or not, so you are giving away nothing that the research says you need.

    If a family asks why their child’s work was displayed, the answer should be a sentence you already have: everyone’s work goes up, nobody’s name has to, and nothing personal is ever used. Worth putting in a back-to-school message the first time you run it, rather than after a phone call.

    What Goes Wrong, and How Do You Tell It Is Your Structure?

    Four failure modes cover almost all of it, and all four are things you set up rather than things students did.

    Four numbered failure modes of a gallery walk activity: no focus so no substance, feedback aimed at the person rather than the work, no revision after the walk, and running critique before the classroom culture can support it
    Every one of these is a structure problem, not a student problem.

    The first is no focus, which produces “good job” and a smiley face. Students are not being lazy. They have not been told what to judge, so they default to the safest comment available.

    The second is feedback aimed at the person. See above. The fix is mechanical — stems, and a rule that every note names a specific part of the work.

    The third is no revision, which is the most common and the most expensive. If the notes go into a backpack, the walk taught judgment to nobody and cost you a period.

    The fourth is running it too early. Critique is a high-trust activity. In a room where that trust is not there yet, start with something lower stakes — structured peer response on a single paragraph at their desks, with you reading over shoulders — and build toward the full walk over a few weeks.

    Two practical notes, because the guides tend to skip both. Prep is roughly ten minutes — deciding the focus, writing three criteria, and finding the sticky notes. There is no packet to build, which is a large part of why this protocol survives in real classrooms. And the student who has nothing to post is an ordinary Tuesday, not a crisis. They post what exists, even if it is a title and two sentences, and they do the walk and give feedback like everybody else. Giving feedback is most of the cognitive work in this activity anyway. Exempting them from the one part they can still do teaches them that the room has stopped expecting anything from them.

    Where Does This Fit in a Unit?

    Mid-draft, not at the end. A gallery walk on finished work is a showcase, which is a fine thing but a different thing. The activity earns its period when the work is still changeable.

    In a project, the natural slot is the day the first full draft exists and before any of it is graded. The critique and revision tools in a PBL toolkit are built around exactly that moment. In a writing unit, it is after the first full draft and before the revision conference. In a problem-based math or science lesson, it is after groups have committed to an approach and before they have invested two days in it.

    One per unit is plenty. Run it weekly and it becomes the thing students perform rather than the thing they use, which is the same failure mode that eventually catches every student-centered routine that gets over-used.

    Does This Work in a Class of Thirty-Five?

    Yes, and it is one of the few protocols that gets easier as the class gets bigger, because it runs in parallel.

    Eight or nine stations, groups of four, three minutes each, and the whole room is working at once. The constraint is wall space and traffic flow, not headcount. If the room is tight, put work flat on desks and rotate groups between desk clusters instead of around the perimeter — same protocol, less collision.

    Where class size does bite is the revision block. Thirty-five students each wanting to ask you about one piece of feedback does not fit in ten minutes. The answer is that they do not ask you. They pick one change and make it. You circulate. The ones who are genuinely stuck will find you.

    Access is worth planning for rather than improvising. A student who reads well below grade level cannot absorb six peers’ drafts in three minutes, and a student newer to English may be able to judge the work and not able to write the note quickly. Both are solved the same way: let feedback be spoken to a partner who writes it, or give a sentence frame with the technical words already in it. Neither reduces the demand — the judgment is still theirs — and both stop the activity quietly sorting the room by reading speed.

    What to Do Next

    Pick a task you already have, where the work is half-finished and you were going to collect it anyway. Decide the one thing students will look for. Write three success criteria on the board. Give them sticky notes, three sentence stems and three minutes a station, and reserve the last ten minutes of the period for one change each.

    Do not grade it. The research says the grade does nothing for students this age, and the moment it is graded you are back to students writing what they think you want to read.

    Then look at the revisions rather than the sticky notes. The sticky notes tell you how well you set up the protocol. The revisions tell you whether anybody learned anything, which is the only question that matters.

    Before you go: grab the free The Gallery Walk Planner (2 pages) (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Is there actual research behind the gallery walk activity, or just blog posts?

    Both, and they are not the same strength. The peer assessment that happens inside a gallery walk is well evidenced: a meta-analysis of 54 control-group studies found g = 0.31 on academic performance, rising to g = 0.44 at secondary level. The gallery walk format is a different claim, and the evidence there is thin. The closest study at the right grade level followed 198 Grade 8 students but had no control group, and found no significant relationship between how the walk was run and students’ knowledge gains or interest. Run the activity for the peer assessment, and build it so the peer assessment is real.

    How do you keep a gallery walk from humiliating the student whose work is weakest?

    Three rules handle most of it. Everyone’s work goes up every time — a wall of selected work tells the room exactly who was selected. Never use personal or narrative writing; use the analytical task, the lab, the problem set or the project artifact. And let a student display work without their name if they ask for that. Anonymity was not a significant moderator in the peer assessment research, so you lose nothing measurable by allowing it. If a student still refuses, take it privately rather than in front of the room.

    Should peer feedback on a gallery walk be anonymous or signed?

    Either works, and you can decide it on classroom grounds rather than research grounds. In the Double, McGrane and Hopfenbeck meta-analysis, anonymity was one of several implementation choices that showed no significant difference in effect. Signed notes tend to be more careful and let the author follow up with a question; anonymous notes tend to be more honest early in the year. A reasonable default is signed, with the caveat that you will read them.

    Should you grade the feedback students give?

    No, not in grades 6–12. The one place the research splits by age is exactly here: peer feedback carrying a grade significantly improved university students’ performance, and that effect could not be shown for primary or secondary students. Grading it also changes what students write — they start producing what they think you want rather than what they noticed. Hold them accountable for completing it, not for a score on it.

    What if a student gets feedback that is simply wrong?

    They are allowed to reject it, and saying so out loud is part of the lesson. Require that every student can name one note they acted on and why, or one note they decided not to act on and why. Both answers show judgment, which is the skill peer assessment is supposed to build. A protocol that makes a student obey bad advice from a classmate is teaching compliance, not critique.

    How long does a gallery walk take, and how much preparation?

    One class period, and about ten minutes of prep. The prep is deciding the single focus, writing three success criteria, and finding sticky notes — there is no packet to build. In the period, allow eight minutes to model the protocol the first time, two to three minutes per station, and a protected ten minutes at the end for revision. If the period is short, cut the number of stations, never the revision block.

    What do you do when a student writes something unkind on a note?

    Collect it, handle it privately, and do not re-run the whole protocol as a punishment for the class. Then look at what made it possible: unkind notes are far more common when the focus was vague, because “say something about this” invites personal commentary while “find the claim and say whether the evidence supports it” does not. Critique also depends on an existing classroom culture — if the room is not there yet, scale back to partner feedback at desks and build up to the full walk.

    How do you know whether the gallery walk actually taught anything?

    Look at the revisions, not the sticky notes. The notes tell you how well you set up the protocol; the changes students made tell you whether the feedback was understood and used. A quick version: collect the one-sentence statement of what each student changed and why. If most of them name a surface change — spelling, formatting, length — the focus you set was too broad and the next walk needs a narrower one.

    Sources

    1. Double, K. S., McGrane, J. A., & Hopfenbeck, T. N. (2019). “The Impact of Peer Assessment on Academic Performance: A Meta-analysis of Control Group Studies.” Educational Psychology Review, 32. Meta-analysis of 54 experimental and quasi-experimental control-group studies; overall g = 0.31 (p < .001) on academic performance; peer assessment outperformed no assessment and teacher assessment (0.28) and was comparable to self-assessment; effects robust across delivery mode, frequency and educational level. https://ora.ox.ac.uk/objects/uuid:0a09975c-e7e6-416f-b888-977230b29ba4
    2. Clearinghouse Unterricht (TUM), Short Review 28 on Double et al. (2020). Cited as a summary, not as the primary study. This is where the secondary-level figure used above comes from — g = 0.44 across 13 studies — together with the moderator finding that peer feedback carrying a grade significantly helped university students (g = 0.55) but could not be shown to help primary or secondary students. https://www.clearinghouse.edu.tum.de/wp-content/uploads/2023/10/CHU-KR-28_ENG_Double_2020.pdf
    3. “Gallery Walk Activities in Teaching Social Studies: Inputs in Enhancing Knowledge, Interest, and Attitude of Grade 8 Students” (2023). International Journal of Research Publications. 198 Grade 8 students, Philippines. One-group pre-test/post-test and correlational design with no control group. Mean gained scores 7.88 / 8.15 / 8.01 across three lessons; correlations between implementation features and cognitive skills were not significant, and correlations with interest were “not significant at 0.05 level”; only attitude correlations reached significance. https://ijrp.sfo3.cdn.digitaloceanspaces.com/pubjournal/5094/1001281720235215.pdf
    4. PBLWorks / Buck Institute for Education. “Using Gallery Walks for Critique & Revision in PBL.” Practitioner protocol. The claim that a specific focus and criteria produce higher-quality feedback is reported by the authors as their own experience — “We’ve found” — and the article cites no studies. Quoted here as practice knowledge, not as research. https://www.pblworks.org/blog/using-gallery-walks-critique-revision-pbl
    5. Kluger, A. N., & DeNisi, A. (1996). “The Effects of Feedback Interventions on Performance: A Historical Review, a Meta-analysis, and a Preliminary Feedback Intervention Theory.” Psychological Bulletin, 119(2). 607 effect sizes across 23,663 observations; average d = 0.41; feedback interventions reduced performance in over a third of cases. The primary article was not reachable from this session; the figures above are taken from a secondary summary of it and are labelled as such. https://explore.psychsafety.com/n/kluger-denisi-1996/
    6. Varlas, L. (2017). “Peer Feedback Without the Sting.” Educational Leadership (ASCD), May 2017. Source of the “kind, specific, helpful” framing attributed to Ron Berger of EL Education, and of the point that peer critique depends on existing classroom culture rather than on the protocol alone. https://www.ascd.org/el/articles/peer-feedback-without-the-sting

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Group Work Roles for Students: What the Evidence Supports, and What It Doesn’t

    Group Work Roles for Students: What the Evidence Supports, and What It Doesn’t

    By Clay Shumate

    Group work roles for students are assigned jobs inside a team — a scheduler, a source keeper, a build lead — meant to stop one student doing everything. They are the standard fix, and the research supports what sits underneath them rather than the roles themselves. Individual accountability and a real shared goal are well evidenced. Role cards are not.

    That distinction is not an academic quibble. It is the difference between a role that changes who works and a laminated card that changes nothing.

    Key Takeaways

    • Cooperative learning works — and works least well in secondary school. A meta-analysis of 51 studies put the achievement effect at 0.54, with the secondary level significantly lower than both primary and university.
    • The evidence for assigned roles specifically is thin. A 2023 systematic review of 36 studies concluded that empirical evidence for the assumption that roles improve collaborative problem-solving “has not yet been provided.”
    • What is well supported is individual responsibility plus a clear common goal. Both are necessary conditions. Neither requires a role card.
    • Making effort matter beats making effort watched. In one experiment, students who believed their effort genuinely affected the outcome worked just as hard unobserved as observed students did.
    • A role has to produce something. If nobody would notice the role vanishing, it is decoration, and students work that out faster than adults do.

    Free Download · PDF

    Group Work That Does Not Collapse (2 pages)

    Page 1 is the five tests to run before you make a single role card, plus the one sentence that answers the parent email about group grades. Page 2 is the task planner and a plain table of which claims the evidence actually supports.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Two-column graphic separating well-supported group work findings including the 0.54 achievement effect size from unestablished claims including that assigning named roles improves outcomes
    The left column is why you run group work. The right column is why your role cards are not working.

    What Are Group Work Roles for Students?

    A group work role is a named job inside a team, assigned rather than negotiated, with responsibilities attached to it. The familiar set is project manager, researcher, recorder, timekeeper, presenter, materials manager.

    The logic is sound on its face. Left to themselves, a group of four will not divide work evenly. One student will take over because they want the grade, one will drift because taking over is exhausting, and two will do what is asked and no more. Assigning roles is supposed to pre-empt that by giving everybody a defined share before the negotiation can happen.

    The problem is that most of the roles on that standard list are not shares of the work. They are labels attached to students who are then expected to do the work anyway. “Encourager” is the clearest case: it produces nothing, it cannot be done badly in any way you could point to, and if the student simply does not do it, the project finishes on time regardless.

    Do Assigned Roles Actually Improve Group Work?

    Nobody has shown that they do, and the researchers who looked hardest say so directly.

    He, Shi, Choi and Zhai published a systematic review in Thinking Skills and Creativity in 2023, covering 36 empirical studies of student roles in collaborative learning from 2013 to 2022. Their conclusion is unusually blunt for a review: the assumption that roles drive collaborative problem-solving competency is widespread, and “empirical evidence for this assumption has not yet been provided.”

    Two further findings from that review are worth carrying into a classroom:

    • Roles are fluid. Whatever you assign, students renegotiate it once the work starts. The review distinguishes scripted roles, which the teacher sets, from emergent roles, which the group actually settles into — and the emergent ones are what you end up observing.
    • Effects are conditional. Whatever roles do depends on the students, the context and the teacher support around them. There is no version where the card does the work.

    And the sample problem again: of those 36 studies, 23 were in higher education, 7 in primary schools, and 6 in secondary. The advice circulating in secondary professional development is mostly borrowed from somewhere else.

    None of this means stop assigning roles. It means stop expecting the assignment to be the intervention. What you are actually reaching for when you hand out roles is accountability, and accountability has much better evidence behind it than roles do.

    What Does the Research Actually Support?

    Two things, and both are conditions rather than techniques.

    Kyndt and colleagues published a meta-analysis of face-to-face cooperative learning in Educational Research Review in 2013, covering 65 articles with 51 providing usable effect sizes. The headline is good news: an achievement effect of 0.54, with a confidence interval well clear of zero. Attitudes moved far less, at 0.15.

    Their summary of what makes it work is short. Effectiveness depends on the individual responsibility each learner takes for completing the task and on a clearly defined common group goal. That is the whole mechanism. Everything else — the cards, the contracts, the seating — is machinery for producing those two conditions, and the machinery is replaceable.

    Two findings that cut against the standard advice

    Group rewards did not beat individual rewards. Kyndt’s team compared the two and found no significant difference. They flag this themselves as contradicting earlier reviews, which had consistently favoured group rewards, and they suggest the more useful distinction is between result interdependence — your grade depends on the team’s outcome — and task interdependence — you literally cannot finish without the others. The second has the more consistent evidence behind it. Build the task so students need each other, and you need the grading lever less.

    And group work does measurably worse in secondary school. Kyndt’s moderator analysis found the secondary level had a significantly lower effect size than both primary and tertiary — differences of roughly 0.20 and 0.18 respectively. That finding deserves more attention than it gets on a site like this one. The age group that gets the most group work in the name of collaboration skills is the age group where the achievement payoff is smallest. It is still positive. It is just not what the slide deck claimed.

    Why Does One Student End Up Doing Everything?

    Because effort that nobody can see is effort most people reduce — and because in a lot of group tasks, one student’s effort genuinely does not change the outcome.

    The research term is social loafing, and it has been studied since the 1970s. Karau and Williams’s 1993 meta-analysis in the Journal of Personality and Social Psychology is the standard reference for it. The classic finding is that people work less hard in groups than alone, and that making individual contributions identifiable reduces the drop.

    Shepperd and Taylor ran a more interesting version in 1999. Their control condition replicated the standard effect: participants whose work would be evaluated individually produced 28.4 ideas, against 20.0 for participants who knew nobody would look. But their main finding was about something else. Participants who were not going to be evaluated but who believed their effort genuinely affected the group’s result produced 28.4 — matching the evaluated group exactly. Participants who were unevaluated and did not believe their effort mattered produced 21.1.

    Read that carefully, because it reframes the whole problem. Surveillance works. But believing your effort matters works just as well, and it does not require you to watch anybody. Most role systems are built as surveillance — here is your job, I will be checking. The better design question is whether the role is one where a student can see their own effect on whether the thing succeeds.

    Five numbered tests a group role must pass: does the task need it, does it produce something, would the group notice if it vanished, can it be done badly visibly, and does it rotate
    Run a role you already use through these five. Most standard roles fail at least two.

    How Do You Build Roles That Come From the Task?

    Start with the work, not with a list of role names. Look at what the project actually requires, find the parts that are genuinely separable, and name those. A role invented before the task exists will never fit it.

    Five tests. A role that fails two of them is decoration:

    1. Does the task actually need it? If one motivated student could do the whole project faster alone, you do not have a group task, you have an individual task with an audience. No role card repairs that. Fix the task. If you want the signs and recording sheets done, my Station Rotation Toolkit on TPT covers eight reusable stations.

    2. Does the role produce something? A document, a build, a dataset, a written objection — something with that student’s name on it that exists at the end of the period. If the output of a role is a behaviour, it is not a role.

    3. Would the group notice if it vanished? Imagine the student assigned to it did nothing at all. If the project still finishes, the role was never load-bearing.

    4. Can it be done badly in a visible way? Accountability needs something you can point at. “You weren’t encouraging enough” is not a conversation anyone can have. “Your source list has three dead links and one of them is where our main claim came from” is.

    5. Does it rotate? A permanent project manager is just the student who was already doing everything, now with institutional backing. Rotate on a fixed schedule and say so in advance.

    One more thing to watch, and it will not show up in any of the five tests: who you hand which role to. If you assign by who seems suited to it, you will reliably end up with the same students leading and the same students recording, and the pattern tends to break along lines nobody intended. Rotation fixes most of this automatically, which is a second reason to do it. If you are letting groups pick their own roles instead, understand the trade — the systematic review found roles get renegotiated anyway, so student choice is honest about what happens regardless, but the student who always takes charge will take charge again, and you have given it your blessing.

    Which Roles Do Real Work?

    The ones with a deliverable attached. These four transfer across subjects, and the names matter less than the artefact:

    • Source keeper. Owns the reference list and has to defend, out loud at a check-in, where every factual claim came from. Produces a document. Can fail visibly.
    • Build lead. Owns the actual product and the version everyone else works from. Produces the thing. Cannot be faked.
    • Sceptic. Writes down the strongest objection to the group’s own argument each working day and brings it to the check-in. This one is underused and it is the best of the four — it is the only role that makes disagreement a job rather than a personality trait.
    • Scheduler. Owns the deadline map and reports in writing what slipped and why. Produces a record. The record is also your early-warning system.

    Three that are usually decoration: encourager, which produces nothing; materials manager, which is ten seconds of real work dressed up as a responsibility; and group leader, which in practice names the student who was already carrying the group. If you want a fuller set of ready-made structures, the project-based learning toolkit has team contracts and role cards you can adapt — but run them through the five tests first rather than printing them as they come.

    List of four group roles with real deliverables — source keeper, build lead, sceptic and scheduler — contrasted with three roles that produce nothing: encourager, materials manager and group leader
    The test is not whether a role sounds responsible. It is whether it leaves evidence.

    How Do You Grade This Without Punishing Cooperation?

    Grade the individual deliverable, not the student’s share of the group’s grade. And here the evidence genuinely conflicts, so you should know that before you decide.

    Every modern treatment says individual accountability is essential, and Kyndt’s meta-analysis names individual responsibility as one of two necessary conditions. But a research synthesis on instructional grouping from 1987 draws a sharper line: group-level recognition encourages cooperation, while “evaluation of each individual student’s contribution to a group score discourages cooperation.” That synthesis is old, and it covers elementary through secondary rather than secondary alone, so weight it accordingly — but the mechanism it describes is not hard to recognise. If helping a teammate improves their slice of the grade and not yours, you have built a reason not to help.

    The way out is the distinction Kyndt’s team drew. Do not score “how much of the group grade did this student earn.” Score the artefact the student personally produced — the source list, the build, the written objections — and let the group’s shared product carry its own separate grade. Now the individual work is assessed, the shared work is assessed, and nothing in the system pays a student for withholding help. Grading individuals inside a group project is a longer problem than this section, and it is worth reading on its own.

    Keep self-assessment separate from all of it. Asking students to rate their own contribution is useful for the conversation it starts; feeding those ratings into a grade turns them into negotiation. There is a whole method for group-work self-assessment that keeps it honest by keeping it ungraded.

    This is also the answer to the complaint you will get from a parent, usually in week three: my child did all the work and everyone got the same grade. Under this structure they did not get the same grade. Their child’s own artefact was marked on its own merits, and so was everyone else’s. Being able to say that in one sentence is worth the setup.

    And be honest about the cost, because the setup is not free. Four individual artefacts per group is more to look at than one project per group. The thing that makes it survivable is that most of those artefacts are short and you are checking them rather than marking them — a source list is a two-minute read, a scheduler’s slip report is thirty seconds. If you find yourself writing comments on all of them, you have turned a checking task into a marking task and you will quit by November.

    Partner Work as the Default, Not the Reward

    My room is set up for this and it changes what roles have to do. Trapezoid and rectangular tables pushed together in pairs and fours, no desks in rows, nobody sitting alone facing the front. I am on a wheeled stool moving between groups rather than standing at the front, and students get pulled to a whiteboard table I built for small-group work.

    Students working alongside somebody is the resting state of that room, not a thing we move into on project days. That matters for roles more than it sounds like it should. When collaboration happens twice a term, every group task needs heavy scaffolding because nobody has practised. When it is the ordinary condition of the room, students have already built the habits the role cards are trying to install, and the roles can be lighter and more specific — a job for this project, not a personality assignment for the year.

    It also makes the circulating part work. The scheduler’s written report of what slipped is only useful if an adult reads it the same day and does something. That is not a documentation system. It is a reason to be standing at that table on Wednesday asking a specific question of a specific student, which is also roughly how expecting every student to answer works in a whole-class discussion.

    Where Group Work Roles Go Wrong

    • Assigning roles to a task that does not need a group. The most common failure and the least often diagnosed. Students know when four people are doing one person’s work.
    • Permanent roles. The organised student is the manager all year, learns nothing new, and resents it. Rotate.
    • Roles with no deliverable. If you cannot name the artefact, you have named a mood.
    • Treating the role card as the accountability. The card is a label. The accountability is somebody looking at the artefact and saying something about it.
    • Grading share-of-group-work. It teaches students that helping a teammate costs them, which is the opposite of the thing you are trying to build.
    • Assuming the roles survive contact. They do not — the 2023 review is explicit that roles shift once work starts. Check what students actually ended up doing rather than what you assigned.

    What to Do Next

    Take the next group task you have planned and run it through two questions before you print anything. First: could one student do this faster alone? If yes, the task needs redesigning and no role system will save it. Second: for each role you were going to assign, what does that student hand me at the end of the period?

    Keep the roles that answer the second question. Delete the rest — you will usually find you are down to three, and three real roles beat six decorative ones. Rotate them on the next project, and tell students that is the plan so the organised kid is not quietly serving a life sentence as project manager.

    Then look at what the work asks of students generally. Roles are a structure for responsibility, and structures only hold where responsibility is already something the room expects. If the task is real and the students are used to being accountable for their own share, light roles are enough. If neither is true, heavy roles will not rescue it.

    Before you go: grab the free Group Work That Does Not Collapse (2 pages) (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Do group work roles for students actually improve learning?

    There is no good evidence that the roles themselves do. A 2023 systematic review of 36 studies on student roles in collaborative learning concluded that empirical evidence for the assumption that roles improve collaborative problem-solving competency "has not yet been provided." What is well evidenced is what roles are meant to produce: individual responsibility for a share of the task, and a clearly defined common goal. Both of those were named as necessary conditions in a meta-analysis of 51 cooperative learning studies. Assign roles if they help you create those conditions. Do not expect the assignment to be the intervention.

    What are the best roles to assign for a group project?

    The ones with a deliverable attached. A source keeper who owns the reference list and has to defend where each claim came from. A build lead who owns the product and the working version. A sceptic who writes down the strongest objection to the group’s own argument each day. A scheduler who reports in writing what slipped. Each produces an artefact with a name on it. Encourager, materials manager and group leader usually do not, which is why they get ignored by the second week.

    How do I stop one student doing all the work?

    Two levers, and the second is the stronger one. Make individual contributions visible — each student hands in something of their own, not just a share of the group product. And make the task one where a student can see their own effort affecting whether it succeeds. In one experiment, participants who believed their effort genuinely mattered produced exactly as much unobserved as participants who knew they were being evaluated individually. Surveillance works, but so does genuine consequence, and consequence does not need policing.

    Does group work even work in middle and high school?

    Yes, but less well than most professional development implies. The 2013 meta-analysis that puts cooperative learning’s achievement effect at 0.54 also found the secondary level had a significantly lower effect size than both primary and university level — differences of roughly 0.20 and 0.18. It is still a positive effect and still worth running. It is just not the strongest case for collaboration, and the age group getting the most group work is the one where the measured payoff is smallest.

    Should I grade each student on their contribution to the group?

    Grade the artefact the student personally produced, not their estimated share of a group score. The evidence here conflicts and it is worth knowing: individual accountability is named as a necessary condition in the modern meta-analysis, but an older research synthesis found that evaluating each student’s contribution to a group score actually discourages cooperation. Both can be true if you separate them — score the individual deliverable on its own, score the shared product on its own, and never create a situation where helping a teammate costs a student points.

    Should group roles be fixed or rotated?

    Rotated, on a schedule you announce in advance. A permanent project manager is normally the student who was already doing everything, now with your authority behind it — which entrenches exactly the pattern roles are supposed to break. It also means the student who most needs practice at coordinating never gets it. Fixed roles are defensible only inside a single short task where there is no time to learn a new one.

    What if students ignore the roles I assigned?

    Expect it, because the research does. The 2023 review found student roles are fluid and renegotiated once work begins, and distinguishes the scripted roles a teacher assigns from the emergent roles a group settles into. The useful response is to check what students actually ended up doing rather than whether they followed your chart. If the emergent division is working and everyone is contributing, that is the outcome you wanted. If one student has absorbed three roles, that is the problem to solve — and it is a task-design problem more often than a compliance problem.

    How many students should be in a group?

    The cooperative learning research does not settle this, and anyone quoting a magic number is going beyond their evidence. What the social loafing literature does establish is that effort drops as groups get larger, because each person’s contribution becomes less visible and less consequential. That points toward three or four for most secondary tasks — small enough that every student has a load-bearing share, large enough that the work genuinely needs dividing. If you cannot write a real deliverable for a fifth student, the group is too big.

    Sources

    1. Kyndt, Eva, Elisabeth Raes, Bart Lismont, Fran Timmers, Eduardo Cascallar, and Filip Dochy. “A Meta-Analysis of the Effects of Face-to-Face Cooperative Learning. Do Recent Studies Falsify or Verify Earlier Findings?” Educational Research Review, vol. 10, 2013, pp. 133–149. 65 articles, 51 with usable effect sizes; achievement ES 0.54 (95% CI .47–.60, p < .001); attitudes 0.15; perceptions 0.18 (n.s.). Secondary level significantly lower than primary and tertiary (differences of −.20 and −.18, both p < .05). No significant difference between group-reward and individual-reward methods (Qm = 1.64, p = .20). Authors flag possible publication bias and small subgroup samples. https://motivatingus.wordpress.com/wp-content/uploads/2017/03/kyndt-et-al-2013-edurev-cooperative-learning.pdf
    2. He, Shan, Xiaoyan Shi, Tae-Hee Choi, and Junqing Zhai. “How Do Students’ Roles in Collaborative Learning Affect Collaborative Problem-Solving Competency? A Systematic Review of Research.” Thinking Skills and Creativity, vol. 50, 2023, article 101423. 36 empirical studies published 2013–2022; 23 higher education, 6 secondary, 7 primary. Roles found to be fluid and renegotiated during group work; scripted vs emergent roles distinguished; effects conditioned by student characteristics, learning context and teacher support. Authors state that empirical evidence for the assumption that roles influence collaborative problem-solving competency “has not yet been provided.” https://www.sciencedirect.com/science/article/abs/pii/S1871187123001918
    3. Shepperd, James A., and Kevin M. Taylor. “Social Loafing and Expectancy-Value Theory.” Personality and Social Psychology Bulletin, vol. 25, no. 9, 1999. Control condition replicated the standard social loafing effect: 28.4 uses generated under individual evaluation vs 20.0 with no evaluation (t = 1.94, p < .05). High-instrumentality participants with no evaluation generated 28.4, matching the evaluated condition; low-instrumentality no-evaluation participants generated 21.1. This paper is also the source used here for the citation of Karau and Williams’s 1993 social loafing meta-analysis, which was not read directly. https://people.clas.ufl.edu/shepperd/files/PSPB1999.pdf
    4. Ward, Beatrice A. Instructional Grouping in the Classroom. School Improvement Research Series, Research You Can Use, Close-Up #2, 1987. Research synthesis. Recommends heterogeneous grouping; holds individual accountability to be essential; finds that group-level rewards or recognition encourage cooperation while “evaluation of each individual student’s contribution to a group score discourages cooperation”; states the teacher must specify subtasks and assign responsibility. Covers elementary through secondary rather than secondary alone, and is now dated — cited here for the reward tension it documents. https://educationnorthwest.org/sites/default/files/InstructionalGrouping.pdf
    5. Farrell, Mark, Tobias Schönbeck, and the CHU Research Group. Cooperative Learning in the Classroom: New Findings Substantiate the Effectiveness of This Method. Clearinghouse Unterricht, Short Review 4, 2023. Practice-facing summary of the Kyndt meta-analysis; reports achievement ES 0.54 and attitudes 0.15, and a g = 0.32 advantage for science and mathematics over social sciences and languages. States that effectiveness “depends on the individual responsibility that each learner takes for completing the task and on a clearly defined common group goal.” Cited as a summary, not as the primary study. https://www.clearinghouse.edu.tum.de/wp-content/uploads/2023/12/CHU_KR4_ENG_Kyndt_2013.pdf

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Cold Calling in the Classroom: What the Research Says and How to Do It Without the Gotcha

    Cold Calling in the Classroom: What the Research Says and How to Do It Without the Gotcha

    By Clay Shumate

    Cold calling in the classroom means calling on a student who did not raise a hand. The research says it works: more students answer voluntarily in classes where it is normal, and the gap between who speaks and who stays quiet narrows. The same research says it makes students anxious. Both are true, and the second one is the part most teaching advice leaves out.

    What follows is what the studies actually measured, who they measured it on, and the version of cold calling that holds a high standard without turning a question into a punishment.

    Key Takeaways

    • Cold calling raises voluntary participation, and the effect is not small. 77% of students answered voluntarily in high cold-call sections, against 55% in low cold-call sections — and it grew over the term rather than wearing off.
    • It closes a gender gap in who talks. Women answered less often than men in low cold-call classes and more often than men in high cold-call classes.
    • It also makes students anxious, reliably. Of 52 students interviewed about active-learning practices, 31 said cold call only increased their anxiety. Not one said it decreased it.
    • Every one of those studies was run on college undergraduates. There is no comparable randomized evidence from a grades 6–12 classroom, and anyone who tells you there is has not read the papers.
    • The design decisions are what separate the two outcomes. Rehearsal time, three seconds of silence, visible randomness and a return path cost nothing and change what the practice feels like from a seat.

    Free Download · PDF

    The Cold Call Card and Participation Log (2 pages)

    Page 1 is the six moves, the sentence to say out loud in week one, and what to do about a flat refusal, a student newer to English, and a documented accommodation. Page 2 is a five-day log for finding the names that never speak.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Two-column comparison of cold calling research: participation findings including 77 percent versus 55 percent voluntary answering, beside anxiety findings including 31 of 52 students reporting increased anxiety
    The two halves of the evidence, side by side. Most articles quote one column.

    What Is Cold Calling in the Classroom?

    Cold calling is asking a question and then naming a student who did not volunteer. That is the whole definition. It is not a discipline move, it is not a pop quiz, and it is not the thing where you catch the kid who was looking out the window.

    The confusion matters because the two versions produce opposite results. A question asked of a class, followed by silence, followed by a name, is an instructional move. The same name said to a student who has stopped paying attention, with a question attached as the consequence, is a correction wearing a question’s clothes. Students can tell the difference in about half a second, and so can anyone watching.

    Hands-up questioning has an obvious problem that cold calling is meant to solve. If you only call on raised hands, you learn what your four most confident students think. Everyone else gets to opt out of thinking, because they know in advance that nothing will be asked of them. That is not a participation problem. It is an information problem — you are teaching a room you cannot see.

    Which means the answer has to be worth collecting. If a student says something that reveals half the room is stuck on the same step and you carry on to the next slide anyway, you have not run a check, you have run a performance. The point of hearing from the students who were not going to volunteer is that they are the ones most likely to tell you something you did not already know.

    Does Cold Calling Actually Increase Participation?

    Yes, and the increase is larger and more durable than I expected when I went looking for the numbers.

    The most careful study on this is Dallimore, Hertenstein and Platt (2013), published in the Journal of Management Education. They tracked 632 undergraduate sophomores across 16 sections of a required accounting course, with seven different instructors. Observers sat in classrooms twice per section and recorded every answer and whether it was volunteered or cold-called. Sections were then sorted into high cold-call and low cold-call environments.

    The results, in the order that matters:

    • 77% of students in high cold-call sections answered at least one voluntary question, against 55% in low cold-call sections.
    • Voluntary answers per student: 2.25 in high cold-call sections, 1.60 in low.
    • The effect grew. High cold-call sections went from 68% at the first observation to 86% at the second. Low cold-call sections sat at 55% both times and did not move.

    That last line is the one worth sitting with. The usual objection to cold calling is that students will participate only under threat and will shut down the moment you stop. The observed pattern is the opposite — voluntary participation climbed as the term went on. The students in those rooms were not being dragged into speaking. They were getting used to it.

    And it narrows a gap in who speaks

    The same research team published a follow-up in 2019 looking specifically at gender. In low cold-call sections, 52% of women answered voluntarily against 57% of men, and men answered more questions each (1.78 against 1.33). In high cold-call sections that flipped: 82% of women answered voluntarily against 73% of men, and the per-student difference disappeared. The rate of increase for women was statistically significant.

    If you have ever looked at a discussion and thought the same six boys are carrying this, that is the finding that speaks to it. The quiet half of a room is not quiet because it has nothing to say. It is quiet because volunteering is a social risk, and it is a bigger social risk for some students than others. Taking volunteering out of the equation changes who is heard.

    Does Cold Calling Make Students Anxious?

    Yes, and this is the part of the literature that rarely makes it into a staff meeting.

    Cooper, Downing and Brownell (2018), in the International Journal of STEM Education, interviewed 52 students in large active-learning biology courses about which practices raised or lowered their anxiety. Cold call was the single most negative practice they studied. Thirty-two students brought it up unprompted. Thirty-one of them said it only increased their anxiety. One said it had no effect. Zero students said it decreased their anxiety. The driver students named was fear of being evaluated badly in front of a large group.

    That is not an isolated finding. A 2024 study of 186 students at a Chinese university found cold calling significantly and positively correlated with anxiety, and the authors recommended using it with caution for exactly that reason.

    And here is the uncomfortable detail: in the Dallimore work, students’ self-reported comfort with class participation did not significantly improve in high cold-call sections either. Participation went up. Comfort did not follow it. Those are two different things and the studies measured both.

    So the honest summary is that cold calling makes more students speak and makes some of them feel worse while doing it. Anyone who tells you it is a free win is selling something.

    Does Any of This Apply to a Grades 6–12 Classroom?

    Not directly, and this is the biggest caveat on the page.

    Every study above was conducted on university undergraduates — accounting sophomores, biology majors, design students. Teenagers in a required course they did not choose, in a building where they will see the same classmates for four years, are not undergraduates. The social cost of being wrong in front of people is higher in a high school than in a 300-seat lecture hall, and the floor on it is lower. A fourteen-year-old who gets laughed at in October remembers it in May.

    What transfers is the mechanism, not the number. The mechanism is that when answering is optional, most students opt out of thinking, and when it is not optional, they do not. That is about how attention works, not about how old you are. The specific percentages do not transfer, and I would not quote them at a faculty meeting as if they did. If you want the broader picture of what counts as real evidence that a class understood something, that is a separate question with its own answer.

    Why Three Seconds of Silence Changes What You Get Back

    Most teachers wait about one second between asking a question and doing something else. That measurement comes from Mary Budd Rowe’s wait-time research, and it is one of the few findings in education that has survived essentially unchallenged for fifty years.

    When Rowe trained teachers to extend that pause to three to five seconds, the changes were not subtle. Average student response length went from 8 words to 27. Failures to respond dropped from 7 to 1. Unsolicited appropriate responses rose from 5 to 17. Students the teachers had described as slow started contributing.

    The caveat, stated plainly: Rowe’s classrooms were primary grades. That is elementary science, not a ninth-grade room, and this site does not pretend elementary findings are secondary findings. But the mechanism behind it — that a one-second pause only has time to surface an answer that was already sitting there, and three seconds has time to produce a new one — is not age-specific. The reason to believe it for teenagers is that it describes thinking, and the reason to be careful about it is that nobody has run Rowe’s study on a high school.

    Put together with a cold call, the sequence is: ask, wait three seconds, say a name, wait three more seconds. The first pause is for the class. The second is for the student. Filling either one is the most common way a well-intentioned cold call turns into an interrogation.

    Six numbered moves for fair cold calling: ask before naming, leave three seconds of silence, let everyone rehearse first, make it visibly random, keep a return path open, and never use it as a correction
    None of these costs lesson time. Four of them cost nothing at all.

    How Do You Cold Call Without Making It a Gotcha?

    You make it predictable, you let everyone rehearse, and you make sure no student can lose by being called on. Six moves, in the order you would use them.

    1. Ask the question before you say the name. Say the name first and twenty-nine students stop thinking, because the question is now somebody else’s problem. Question, pause, name. That order is the entire difference between a cold call that teaches the room and one that teaches one kid.

    2. Leave three seconds of silence. Covered above, and it is the move teachers find hardest. Three seconds is longer than it sounds when you are standing up in front of people.

    3. Let everyone rehearse first. Thirty seconds of writing, or thirty seconds of turning to a partner, before any name is said. Now nobody is actually answering cold — every student has already produced something and is being asked to read it out. This single change removes most of what the anxiety research is describing, because the fear is not of speaking, it is of having nothing to say.

    4. Make it visibly random, not targeted. Name cards, a list you work down, a seating chart you move through. The point is not fairness in the abstract. The point is that a student who gets called on knows it was not about them. The moment the selection looks like judgment, you have the gotcha back.

    5. Keep a return path open. “I don’t know” has to lead somewhere other than embarrassment — a hint, permission to ask a partner, or a straight “I’ll come back to you,” followed by actually coming back with a question they can answer. The return is the part people skip, and it is the part that makes the whole thing safe.

    6. Never use it as a correction. The moment a question is the consequence for inattention, it stops being a question and the class learns that your questions are weapons. Handle the off-task student the way you would handle anything else — quietly, at close range, and not as a performance for the room.

    What About the Student Who Genuinely Cannot Answer?

    If a student freezes badly enough to matter, the task was mis-scoped, and that is on the plan rather than the kid.

    I believe this about classrooms generally. A student cannot screw up that badly in a well-planned room. If they can, the question was wrong, the scaffolding was missing, or the adult was not paying attention. The honest fix after a cold call goes wrong is not to stop cold calling and it is not to apologise at length in front of everyone. It is to come back to that student inside the same period with something they can answer, and to look at why the first question had no entry point.

    This is also where the anxiety research should change your practice rather than end it. Students in those interviews were not afraid of being asked things. They were afraid of being evaluated badly in public. Those are separable. You can ask a student something hard and make it structurally impossible for the answer to count against them — nothing said out loud is graded, “I don’t know” is a legitimate answer, and the return path exists. Do that and you have kept the demand while removing most of the threat.

    Respecting students is not complicated, and it mostly comes down to doing exactly what you said you would do, every time. If you tell a class that nothing they say when called on is graded, that has to be true in week fourteen as well as week one.

    Two groups need a decision made in advance rather than in the moment. A student who is newer to English, or who stammers, or who has a documented anxiety accommodation, should not be finding out mid-lesson what your rules are. For the first two, written rehearsal time does most of the work and a quiet “would you rather read what you wrote or tell me after?” does the rest. For the third, the plan governs, not the technique — if a 504 or IEP says the student is not called on without warning, that is the end of the discussion and no research finding overrides it.

    And occasionally a student will simply refuse. Not freeze — refuse, flatly, in front of everyone. Take it at face value, move on without comment, and deal with it later and privately, because a standoff in front of thirty people is a fight you cannot win and should not be having. Nine times out of ten the refusal is about something that happened before the bell.

    Three-phase cold calling checklist covering what to say before the first cold call, the question-pause-name-wait sequence in the moment, and how to repair it the same period when it goes wrong
    Before, during, after. The third block is the one that gets skipped.

    Organization to the Minute, and Then the Pause

    The hardest part of all of this is the silence, and I have an expensive data point about silence.

    I once waited over five minutes. It was one of the most awkward moments of my life. That section finished nine minutes behind my other classes, and it is the only time I have ever assigned homework — I do not otherwise assign it at all. That is the price tag, and most advice about waiting students out does not come with one.

    What I would take from it, applied to a three-second pause rather than a five-minute one: the silence works, and it costs something, and you should know what it costs before you spend it. Three seconds is cheap. Three seconds times forty questions is two minutes of a lesson. That is a trade worth making. Standing there for five is a different decision, and you should make it on purpose.

    My room runs as a workshop — tables in pairs and fours, nobody sitting in a row facing front, me on a wheeled stool moving between groups. Cold calling in a room like that is less of an event than it is in a lecture hall, because students are already talking to each other all period. That is part of why the rehearse-first move matters so much. In a room where partner talk is the default condition rather than a treat, “say what you just told your partner” is barely a cold call at all.

    Where Cold Calling Goes Wrong

    • Using it to catch people. The fastest way to poison it. One targeted question in front of thirty people and every subsequent question in your room reads as a threat.
    • Grading it. Participation points attached to cold-call answers convert a thinking task into a performance task, and the students who most need to practise speaking are the ones who will now dread it most.
    • Skipping the rehearsal. Cold calling without think time is the version the anxiety studies describe. With think time, it is a different practice wearing the same name.
    • Filling the silence. Rephrasing your own question after two seconds, which teachers do constantly, resets the clock and tells the room you did not mean it.
    • Doing it inconsistently. If it happens twice a month, it is an event, and events are frightening. If it happens every day, it is the weather.
    • Only cold calling the strong students. It feels safe and it defeats the purpose. The information you need is from the students who were not going to volunteer.

    What to Do Next

    Pick one class. Tomorrow, before the first question, say out loud how it is going to work: you will call on people who did not raise a hand, nothing said out loud is graded, and “I don’t know” gets you a hint and a second try. Then run the sequence — thirty seconds of writing, ask, three seconds, a name off a list, three more seconds.

    Do it for two weeks before you judge it. The Dallimore data showed the effect growing between the first observation and the second, which means a first week that feels awkward is not evidence of anything. Then look at who is talking and compare it to your memory of September.

    It is also worth putting one line in a newsletter or a back-to-school email: this class expects everyone to answer questions, nothing said out loud is graded, and no student is ever stuck with no way out. A parent who hears about cold calling first from an upset teenager will assume the worst. A parent who heard it from you in August usually does not.

    If you want this inside a wider system of checks rather than as a single technique, the full list of formative assessment strategies puts cold call alongside the other moves that tell you what a room actually understands, and a structured seminar is what this turns into once students are used to being expected to speak. Both work better once the expectations you set in the first week have made being asked a question an ordinary event rather than a surprise.

    Before you go: grab the free The Cold Call Card and Participation Log (2 pages) (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Is cold calling in the classroom bad for anxious students?

    It raises anxiety for a lot of students — in one interview study, 31 of 52 said cold call only increased theirs, and none said it decreased it. But the thing students named as the cause was fear of being evaluated badly in public, not being asked a question. You can remove most of that while keeping the demand: give everyone thirty seconds to write or talk to a partner first, say out loud that nothing spoken aloud is graded, and make sure "I don’t know" leads to a hint rather than an audience. A student with a documented anxiety accommodation should be handled through that plan, not through a classroom technique.

    Does cold calling actually work, or does participation just drop again when you stop?

    In the best study available, it did not drop — it grew. Dallimore, Hertenstein and Platt tracked 632 undergraduates across 16 sections and found 77% of students in high cold-call sections answered voluntarily, against 55% in low cold-call sections, and the high cold-call figure rose from 68% at the first observation to 86% at the second while the low cold-call sections stayed flat. That pattern is the opposite of compliance wearing off.

    Is there any research on cold calling in middle or high school?

    Not of the same quality, and that is worth saying plainly. The participation studies were run on college sophomores in an accounting course, the anxiety study on undergraduate biology students, and a third on university design students. There is no comparable randomized evidence from a grades 6–12 classroom. The mechanism — that optional answering lets most students opt out of thinking — is about attention rather than age, so it is reasonable to expect it to transfer. The specific percentages are not yours to quote.

    How is cold calling different from putting a student on the spot?

    Order and intent. A cold call asks the question first, pauses so the whole class thinks, then names someone from a list or a set of cards. Putting a student on the spot names the student first, usually because they were not paying attention, and attaches a question as the consequence. The first is an instructional move that happens to land on one person. The second is a correction, and the room reads it as one immediately.

    Should I tell students in advance that I am going to cold call?

    Yes, and it costs you nothing. Surprise is not the active ingredient — the active ingredient is that answering is not optional, and students can know that in advance. Telling them the rules, including that nothing said aloud is graded and that there is always a way back from "I don’t know", converts the practice from an ambush into a routine. Routines are not frightening. Events are.

    How long should I wait after asking the question?

    Three seconds before you say a name, and three more after. Mary Budd Rowe’s wait-time research found teachers typically wait about one second; extending the pause to three to five seconds raised average response length from 8 words to 27 and cut failures to respond from 7 to 1. That work was done in primary-grade science classrooms rather than secondary ones, so treat the mechanism as transferable and the exact numbers as not yet tested on teenagers.

    Does cold calling help close participation gaps between boys and girls?

    The one study that looked directly at this found that it did. In low cold-call sections, 52% of women answered voluntarily against 57% of men, and men answered more questions each. In high cold-call sections it reversed: 82% of women against 73% of men, with the per-student gap gone. The authors are clear this was a single institution and a single discipline, so read it as a reason to try it rather than a settled result.

    What do I do when a student freezes and cannot answer at all?

    Give the return path immediately — a hint, permission to ask a partner, or "I will come back to you" followed by actually coming back with a question they can answer inside the same period. Then look at the question itself. If a student froze badly enough that it mattered, the task was mis-scoped or the scaffolding was missing, and that is a planning problem rather than a student problem. Fix the entry point, not the student.

    Sources

    1. Dallimore, Elise J., Julie H. Hertenstein, and Marjorie B. Platt. “Impact of Cold-Calling on Student Voluntary Participation.” Journal of Management Education, vol. 37, no. 3, 2013, pp. 305–341. 632 undergraduate sophomores, 16 sections, 7 instructors, two classroom observations per section; high cold-call sections 77% voluntary participation vs 55% low; 2.25 vs 1.60 voluntary answers per student; high cold-call rose 68% to 86% across observations while low cold-call held at 55%; no significant effect on self-reported comfort in the main analysis. Authors note the single-course design limits generalizability. https://journals.sagepub.com/doi/abs/10.1177/1052562912446067
    2. Dallimore, Elise J., Julie H. Hertenstein, and Marjorie B. Platt. “Leveling the Playing Field: How Cold-Calling Affects Class Discussion Gender Equity.” Journal of Education and Learning, vol. 8, no. 2, 2019, pp. 14–24. Low cold-call: women 52% / men 57% voluntary participation, 1.33 vs 1.78 questions per student (p < .01). High cold-call: women 82% / men 73%, per-student difference not significant; women’s rate of increase significant at p = 0.004. No significant effect on comfort ratings. Single institution, single discipline. https://files.eric.ed.gov/fulltext/EJ1207291.pdf
    3. Cooper, Katelyn M., Virginia R. Downing, and Sara E. Brownell. “The Influence of Active Learning Practices on Student Anxiety in Large-Enrollment College Science Classrooms.” International Journal of STEM Education, vol. 5, no. 1, 2018, article 23. Interviews with 52 undergraduates in large active-learning biology courses; 32 raised cold call, 31 reported it only increased anxiety, 1 reported no effect, 0 reported decreased anxiety; fear of negative evaluation in front of large groups named as the driver. Exploratory interview design. https://link.springer.com/article/10.1186/s40594-018-0123-6
    4. Yang, Hongze, Bo Li, Jiajing Yu, and Lichen Sun. “Impact of Active Learning Instruction in Blended Learning on Students’ Anxiety Levels and Performance.” Frontiers in Education, vol. 9, 2024, article 1332778. 186 students at a Chinese university, 80.6% undergraduate; cold calling significantly positively correlated with anxiety (SE = 0.632, p < 0.001); authors recommend cautious use. https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2024.1332778/full
    5. Rowe, Mary Budd. Wait-Time and Rewards as Instructional Variables: Their Influence on Language, Logic, and Fate Control. ERIC ED061103, presented at the National Association for Research in Science Teaching, April 1972. Over 300 initial recordings plus 84 further tapes; baseline teacher wait time averaged about one second; extending to 3–5 seconds raised mean response length from 8 to 27 words, cut failures to respond from 7 to 1, and raised unsolicited appropriate responses from 5 to 17. Primary-grade science classrooms — not a secondary sample. https://files.eric.ed.gov/fulltext/ED061103.pdf

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.