Author: Clay Shumate

  • Socratic Seminar: How to Run One That Grades 6–12 Take Seriously

    Socratic Seminar: How to Run One That Grades 6–12 Take Seriously

    By Clay Shumate

    A Socratic seminar is a structured, whole-group discussion in which students question a shared text and each other while the teacher stays mostly silent. There is no winner and no correct answer to arrive at. The point is that students do the reasoning out loud, in front of each other, using evidence from something they all read.

    That is the definition. What follows is the part most guides skip: what the research actually supports, what it does not, and how to run one in a real secondary class without the same five kids carrying it.

    Key Takeaways

    • A seminar is not a debate. Nobody is assigned a side, nobody is trying to win, and the question has to be one you cannot settle by looking something up.
    • The strongest evidence is about discussion generally, not this format specifically. A study of 64 middle and high school English classrooms found discussion-based teaching predicted higher spring literacy performance — for low-achieving students as well as high-achieving ones.
    • The most honest finding is a mixed one. A meta-analysis found discussion produced big jumps in student talk and real gains in text comprehension, but few approaches moved critical thinking and reasoning.
    • Preparation is the whole game. A seminar with an unread text is twenty minutes of opinions.
    • Silence is not failure. The most common way a teacher wrecks a seminar is by answering the question they just asked.

    Free Download · PDF

    The Socratic Seminar Prep Sheet

    A text-selection check, the opening-question test, the six moves in order, six norms worth posting, a turn-count tally for you, and the four-way autopsy for a seminar that died.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    What Is a Socratic Seminar?

    A Socratic seminar is a formal, text-based dialogue where students ask and answer open-ended questions with each other, and the teacher facilitates instead of teaching. The National Paideia Center — which has trained schools in this format for decades — defines it as “a collaborative intellectual dialogue facilitated with open-ended questions about a text.” That definition is doing more work than it looks like.

    Every word in it rules something out. Collaborative rules out debate. Open-ended rules out questions with answers. About a text rules out a free-floating conversation on a topic. Take away any one and you have something else — possibly useful, but not a seminar. Here is the distinction that matters most in a secondary room, because teenagers assume it is a debate unless you tell them otherwise.

    Comparison table of a Socratic seminar, a class discussion and a debate across goal, positions, teacher role, what the talk is anchored in, and whether changing your mind counts as success
    A seminar is not a debate. Nobody is assigned a side and nobody wins.

    Does the Research Actually Support Discussion-Based Teaching?

    Yes for discussion in general — and more cautiously than most professional development admits when it comes to thinking. Two studies point in slightly different directions, and both are worth knowing before you build a unit around this.

    The strong one is squarely in our grade band. Applebee, Langer, Nystrand and Gamoran studied 64 middle and high school English classrooms across grades 7, 8, 10, 11 and 12. Controlling for fall performance and background, discussion-based approaches were significantly related to spring literacy performance. Their own summary of why: students in discussion-heavy rooms “internalize the knowledge and skills necessary to engage in challenging literacy tasks on their own.” And the finding held “for low-achieving as well as high-achieving students” — which is the opposite of what people assume about discussion, and the single best argument for not reserving it for honors sections.

    An earlier study from the same research group found the mechanism. Across sixteen middle schools feeding nine high schools in eight Midwestern communities, more than 1,100 eighth and ninth graders were tested in fall and spring. Students whose teachers asked a higher proportion of authentic questions — questions the teacher did not already have an answer to — and who practiced uptake, building the next question out of what a student just said, scored significantly higher on literature achievement. Not more questions. Different questions.

    There is an equity finding buried in the first study that deserves more attention than it gets. Applebee and colleagues note that their interpretation is “complicated because instruction is unequally distributed across tracks.” Translated: the discussion-based teaching that helps low-achieving students the most was not what low-achieving students were mostly getting. If a school is going to act on this research at all, the action is not “add seminars” — it is checking which sections currently get them.

    Now the cautious one, and it is the reason this section exists. Murphy and colleagues ran a meta-analysis of classroom discussion approaches in the Journal of Educational Psychology. They found “strong increases in the amount of student talk and concomitant reductions in teacher talk, as well as substantial improvements in text comprehension.” Then they found this: “Few approaches to discussion were effective at increasing students’ literal or inferential comprehension and critical thinking and reasoning.”

    Read that twice. The talk goes up reliably. Comprehension of the text goes up. Whether students get better at thinking is not established by the pooled evidence, and the majority of the studies in that analysis were conducted in fourth through sixth grade, not with teenagers. If someone tells you seminars teach critical thinking, that is a hope, not a result.

    The same gap shows up in a study built specifically around Socratic dialogue. Mahoney and colleagues ran an eight-week Socratic dialogue series in Dutch vocational secondary education with 85 students and five teachers. Teachers found it deliverable but cognitively demanding, and all five asked for more facilitation training. The study’s own limitations section says there was “no measurement of actual critical thinking skill development.” A paper with “learning to think critically” in the title did not measure whether anyone did. The researchers said so themselves; it is a warning about how this format gets sold, not about them.

    One more, because the number gets quoted at teachers without its age attached. England’s Education Endowment Foundation ran a randomised controlled trial of Dialogic Teaching across 76 schools and roughly 5,000 pupils and found two additional months’ progress in English and science, with larger gains for pupils on free school meals. Those pupils were nine and ten years old. The practices probably transfer to a secondary room. The effect size does not automatically come with them.

    The defensible version: structured, text-anchored discussion with authentic questions is one of the better-evidenced things you can do with a secondary class, and the evidence for it is about comprehension and engagement rather than about producing better reasoners in eight weeks.

    One implication for anyone above the classroom level: this is not a practice you mandate in August and audit in October. Every teacher in the Dutch study asked for more training in facilitation specifically, and facilitation is the whole skill. A department that runs four seminars and talks about them afterward will get further than a district that requires one per unit.

    What Do You Need Before the First Seminar?

    Three things: a text worth arguing about, a question with no answer, and students who have actually read it. Skip any one and the seminar fails in a predictable way.

    The text. Short beats long. One or two pages that every student can hold and mark up beats a chapter half of them skimmed. It needs genuine ambiguity — a primary source with a point of view, a court opinion, a poem, a paragraph from a science paper where the authors hedge. If the text has one obvious reading, there is nothing to discuss and students will know it within four minutes.

    Two constraints on that choice, both easy to miss when you are excited about a text. Every student has to be able to read it. A twelfth-grade primary source in a mixed ninth-grade class produces a conversation among the strongest readers while everyone else decodes. If the text is hard, read it aloud together or gloss the four words that will stop people — but do not let reading level decide who participates. And check the text against your own building’s ground. Ambiguity is the point; a text likely to pull a student into disclosing something personal in front of twenty-eight peers is a different situation, and it is worth a conversation with an administrator before it is worth a seminar.

    The opening question. Write it down before the seminar and test it against one standard: could a well-prepared student answer this in one sentence and be finished? If yes, it is a comprehension check, not a seminar question. “What does the author claim?” is a warm-up. “Is the author being honest about what they do not know?” is a seminar. If writing them is the part you get stuck on, there are forty subject-general seminar questions and a ten-minute way to write your own.

    The reading. This is what quietly kills most first attempts. The Paideia model puts multiple close readings in a pre-seminar phase before any discussion and writing in a post-seminar phase afterward. That three-part shape is the format, not an optional extra. Read the text once ten minutes before the circle forms and you get opinions instead of evidence — then conclude your class cannot handle seminars. Your class is fine. The prep was missing.

    A note on when they read it. The obvious move is to send the text home the night before, and I would push back on that — not on principle, but because it decides in advance that the students who read at home get to participate and the ones who do not, do not. Read it in class. Ten minutes of everyone reading the same page in the same room is not lost time; it is the thing that makes the next twenty minutes possible.

    One practical note about the room. The tables in my room are trapezoids pushed together in groups rather than desks in rows, which is closer to a discussion shape than a lecture hall — but it still is not a circle, and the furniture has to move. Build that into the clock. A seminar that starts with four minutes of scraping chairs starts four minutes late, every time, and you will feel it at the end.

    How Do You Actually Run One?

    Six moves, in order, and the hardest is the fifth. You can run this sequence in a single period and repeat it until it stops feeling like an event.

    Numbered list of six moves for running a Socratic seminar: read the text together in class, everyone writes one question, move the furniture into a circle, restate two norms by name, ask the question once then stop, and write afterward naming whose comment changed your position
    Six moves you can run in a single period. The fifth is the hard one.

    The sequence is deliberately boring in the middle. Step five is the interesting one, and the one I get wrong most often.

    In my own room I have landed on roughly a fifteen-count before I respond to something that is off track — and I mean roughly. It is an average, not a rule, because real classes change day to day and a number you enforce mechanically stops being judgment. A seminar asks for that same discipline in a different situation. You ask the opening question, nobody talks, and everything in you wants to rephrase it. Rephrasing is how you teach a room that if they wait long enough, you will do the work. Ask it once. Then sit there.

    The related move is refusing to evaluate. The instant you say “good point,” you have told twenty-eight teenagers that the goal is to produce points you approve of, and the conversation reorients toward you. Write the comment down instead. Nod. Look at somebody else. The Applebee study’s mechanism — authentic questions and uptake — only works if students believe you do not already have the answer, and praise is the fastest way to convince them you do.

    What Are the Rules, and Who Enforces Them?

    Keep the list short enough that students can hold it, and hand enforcement to them by the third seminar. Long norm lists get read aloud once and ignored. Five or six behaviors, posted, referred to by name, is enough.

    Six Socratic seminar norms: cite the page, disagree with the claim not the person, invite before you add, silence is allowed, ask rather than announce, and you may change your mind out loud
    Short enough that students can hold them, and short enough to enforce by name.

    The norms that matter protect the conversation from its two failure modes: nobody talking, and three people talking. Everything else is manners.

    There is one behavior worth planning for: the student who is performing for the room rather than talking to it. Do not handle that inside the circle. Naming it publicly makes it a bigger performance, and the seminar stops. Let it pass, keep the conversation moving to someone else, and take it up with the student afterward at normal volume. A seminar is a bad place to win an argument with a teenager.

    Worth saying clearly, because it is where teenagers get this wrong: disagreeing with an idea is the job. A seminar where everyone agrees has not happened yet. What is off-limits is going after the person instead of the claim, and the fix is a sentence frame, not a lecture — “I read that line differently, because…” does the whole job. The broader argument for why this works in a secondary room is the same one behind holding students to an adult standard of disagreement rather than a polite one.

    What About the Fishbowl Format?

    A fishbowl Socratic seminar puts one group in an inner circle discussing while an outer circle observes, then swaps them. It exists to solve two real problems: a class of thirty is too big for one conversation, and students who never get observed never get feedback on how they discuss.

    It works, with one condition. The outer circle needs a job. An outer ring with nothing to do is a ring of spectators with phones. Give each outer student one inner-circle partner to track, and one thing to record — how many times their partner cited the text, or the one question their partner asked that moved the conversation. Keep the record behavioral: what the partner did, not how well they did it. “You went back to the text three times” is feedback. “You seemed nervous” is not, and a fifteen-year-old should not be handed that by a classmate. Hand the note to the partner rather than reading it to the room. That feedback is usually more useful than anything I would have said, because it comes from a peer who was watching only them — but it belongs to the person it is about.

    Do not make it your first seminar of the year, though. Students end up learning two unfamiliar procedures at once and the period goes to logistics.

    How Do You Handle the Students Who Won’t Talk?

    Stop treating silence as a single problem, because it is at least three problems with different fixes. Some students have nothing prepared. Some have something prepared and cannot find an opening. Some are genuinely afraid of the sound of their own voice in a quiet room, and no amount of encouragement touches that.

    A study of quiet students by Medaille and Usinger is instructive here, with a caveat stated up front: the participants were ten undergraduates, not teenagers, so treat it as a description of an experience rather than a finding about your seventh period. Nine of the ten struggled with instructor expectations for verbal participation. Six reported physical reactions — trembling, blushing, stuttering — when speaking aloud. And one of them described the thing I now believe is the actual mechanism: “When I speak [in class] I plan out what I want to say.”

    That is not shyness. That is a student who needs a draft, and the fix is not calling on them warmly — it is giving the whole room ninety seconds to write before the circle opens, so the students who need a draft have one and nobody is singled out for needing it.

    Three more moves that cost nothing:

    • Give the quiet student a job that is not an opinion. Tracking which page the group keeps returning to, or reporting the one question nobody answered, is a real contribution that does not require volunteering a view.
    • Use written entry. A sticky note with one question on it, collected at the door and read aloud anonymously, gets a reticent student’s thinking into the room without their voice attached to it. Read them yourself before you read any of them out. Anonymous is not the same as safe — a question can identify the student who wrote it, or disclose something that should not be discussed in a circle, and the two seconds it takes to scan the stack is the whole safeguard.
    • Do not grade volume. More on that next, because it is the single change that most reliably changes who talks.

    I would not promise that a seminar fixes this. Some students will speak twice all year and write the best post-seminar reflection in the class. That is a real outcome, and the format should be built to catch it.

    Should You Grade a Socratic Seminar?

    Grade the preparation and the reflection. Do not grade how many times a student spoke. The moment participation counts, you have created an incentive to talk rather than to think, and the students who most need the practice are the ones the incentive punishes.

    Count what you can defend. Did they annotate the text? Did they bring a written question? Does the post-seminar writing show they changed or sharpened a position — and can they name whose comment did it? That last one is the best assessment item this format produces, because only a student who was listening can answer it. Three comments in a circle are easy to fake. “I came in thinking X, and the point about the second paragraph is why I do not think that anymore” is not.

    This is also the answer when a parent asks why their quiet child is not being penalized, or why their talkative one is not being rewarded: the grade is on the thinking you can see on paper, and it is the same standard for everyone in the room. If your school requires a participation grade, take it from the written pre- and post-work and say out loud, to the students, that the talking is not scored. There is a four-criterion rubric that scores it this way, free and printed on the page. Then hold to it. The first seminar after they believe you is a different conversation. This is the same logic behind treating ungraded checks as information rather than points — the second you attach a score, you stop learning what students actually think.

    Four Ways a Seminar Goes Wrong

    Every one of these is a design problem, not a student problem. If a seminar dies, the autopsy is almost always in the prep.

    1. The question had an answer. Students find the answer in six minutes and then sit there. Fix: write the question the night before and try to answer it yourself in one sentence. If you can, it is not the question.
    2. The teacher kept talking. You clarified, you rephrased, you summarized, and the conversation routed through you every time. Fix: count your own turns. More than four in a thirty-minute seminar and you are running a recitation with the chairs moved.
    3. Three students ran it. Not because they are hogging it — because they are the only ones who prepared. Fix: require a written question from everyone, collected as they come in. Make it an expectation, not a bar — a student who arrives without one writes it in the first two minutes while the furniture moves. Nobody gets locked out of the conversation for being unprepared; they get two minutes and then they are in it.
    4. It happened once. A seminar in October and a seminar in March are two events. Students never get past the awkwardness of the format itself. Fix: shorter and more often. Fifteen minutes on a half-page every other week will outperform two forty-minute showpieces.

    The fourth is the one I would push hardest. Frequency turns this from a special activity into the way the class talks, and the same logic runs through any structure that hands students the thinking: the first run is about the structure, and only the fourth or fifth is about the content.

    What to Do Next

    Pick a one-page text you already teach — something with a hedge or a contradiction in it — and write one question you cannot answer in a sentence. Give students ten minutes in class to read it, mark two places they disagree, and write one question. Move the furniture. Ask your question once, and then do not talk for as long as you can stand.

    Then do it again in two weeks. The first one will be rough, and that is not information about your students. Run four before you decide whether this works in your room.

    If you want the underlying argument for why handing teenagers this much of the conversation is worth the mess, it is the same one behind giving students real decisions rather than cosmetic ones. A seminar is just the version of that argument you can run on a Tuesday.

    Before you go: grab the free The Socratic Seminar Prep Sheet (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    What is a Socratic seminar in simple terms?

    It is a structured discussion where students question a shared text and each other while the teacher mostly stays quiet. There is no assigned side, no winner, and no single correct answer to land on. The National Paideia Center calls it a collaborative intellectual dialogue facilitated with open-ended questions about a text. If the question has an answer, or if the teacher is doing most of the talking, it is a class discussion rather than a seminar.

    How long should a Socratic seminar last?

    Fifteen to thirty minutes of actual discussion is plenty for grades 6–12, and shorter is better when you are starting. A full class period sounds ambitious and usually produces twenty good minutes followed by fifteen of people repeating themselves. Frequency beats length. A fifteen-minute seminar every other week on a half-page text will build the skill faster than two long ones a year.

    What is the difference between a Socratic seminar and a debate?

    A debate assigns positions and produces a winner. A seminar assigns nothing and produces a better-understood question. That difference changes student behavior immediately: in a debate, changing your mind is losing, and in a seminar, changing your mind is the evidence that it worked. Tell students this explicitly before the first one, because teenagers default to debate unless you rule it out.

    Do Socratic seminars actually improve student learning?

    The evidence is genuinely mixed and worth knowing honestly. A study of 64 middle and high school English classrooms found discussion-based approaches significantly predicted spring literacy performance, including for low-achieving students. But a meta-analysis in the Journal of Educational Psychology found that while discussion produced strong increases in student talk and substantial gains in text comprehension, few approaches improved critical thinking and reasoning — and most of those studies were in grades 4 through 6. Discussion is well supported. The claim that it teaches thinking is not settled.

    How do you get quiet students to participate in a Socratic seminar?

    Give the whole class ninety seconds of writing time before the circle opens, so students who need to plan what they say have a draft and nobody is singled out for needing one. Beyond that, offer jobs that are not opinions — tracking which passage the group keeps returning to, or reporting the question nobody answered. And do not grade how often a student speaks. That single change does more than any amount of encouragement, because it removes the reason a struggling student is performing rather than thinking.

    Should I grade a Socratic seminar?

    Grade the preparation and the reflection, not the talking. Annotated text, a written question brought to the circle, and a post-seminar piece of writing are all defensible and all reward thinking. Counting comments rewards volume and penalizes the students who most need the practice. If a participation grade is required, pull it from the written work and tell students plainly that speaking is not scored.

    What makes a good Socratic seminar question?

    One you cannot answer in a sentence and cannot settle by looking something up. Test your question by trying to answer it yourself before the seminar; if you finish in one line, it is a comprehension check. Questions that work usually ask about a tension inside the text — whether the author is being honest about what they do not know, whether two of their claims can both be true, or what the text refuses to say.

    How many students should be in a Socratic seminar?

    Twelve to fifteen is the practical ceiling for one circle. Above that, the students at the edges stop being participants. That is what the fishbowl format is for: put half the class in an inner circle discussing and half in an outer circle observing a specific partner, then swap. Do not run a fishbowl as your first seminar though — students end up learning two unfamiliar procedures at once and the period goes to logistics.

    Sources

    1. Applebee, Arthur N., Judith A. Langer, Martin Nystrand, and Adam Gamoran. “Discussion-Based Approaches to Developing Understanding: Classroom Instruction and Student Performance in Middle and High School English.” American Educational Research Journal, vol. 40, no. 3, 2003, pp. 685–730. 64 middle and high school English classrooms, grades 7, 8, 10, 11 and 12; discussion-based approaches significantly related to spring performance controlling for fall performance, and effective for low-achieving as well as high-achieving students. https://eric.ed.gov/?id=EJ782328
    2. Murphy, P. Karen, Ian A. G. Wilkinson, Anna O. Soter, Maeghan N. Hennessey, and John F. Alexander. “Examining the Effects of Classroom Discussion on Students’ Comprehension of Text: A Meta-Analysis.” Journal of Educational Psychology, vol. 101, no. 3, Aug. 2009, pp. 740–764. Strong increases in student talk and substantial improvements in text comprehension; “few approaches to discussion were effective at increasing students’ literal or inferential comprehension and critical thinking and reasoning.” Majority of included studies were conducted in grades 4–6. https://eric.ed.gov/?id=EJ861185
    3. Nystrand, Martin, and Adam Gamoran. Student Engagement: When Recitation Becomes Conversation. National Center on Effective Secondary Schools, 1990. ERIC ED323581. Eighth- and ninth-grade English classes in sixteen middle schools feeding nine high schools across eight Midwestern communities; over 1,100 students tested fall and spring, 1987–88 and 1988–89. Higher proportions of authentic questions and uptake predicted significantly higher literature achievement. https://files.eric.ed.gov/fulltext/ED323581.pdf
    4. Jay, Tim, et al. Dialogic Teaching: Evaluation Report and Executive Summary. Education Endowment Foundation, July 2017. ERIC ED581114. Three-level clustered randomised controlled trial, 38 intervention schools (2,492 pupils) and 38 control schools (2,466 pupils); +2 months English, +2 months science, +1 month maths; three-padlock security rating. Pupils were in Year 5, aged nine and ten — this is not a secondary-school result. Evaluators note 21% of pupils excluded from analysis. https://files.eric.ed.gov/fulltext/ED581114.pdf
    5. Mahoney, Bethany, Ron Oostdam, Hessel Nieuwelink, and Jaap Schuitema. “Learning to Think Critically Through Socratic Dialogue: Evaluating a Series of Lessons Designed for Secondary Vocational Education.” Thinking Skills and Creativity, vol. 50, 2023, article 101422. 85 students and 5 teachers, Netherlands, eight weekly lessons; teachers found the format deliverable but demanding and all five requested more facilitation training. The study states as a limitation that there was no measurement of actual critical thinking skill development. https://www.sciencedirect.com/science/article/pii/S1871187123001906
    6. Medaille, Ann, and Janet Usinger. “Quiet Students’ Experiences with the Physical, Pedagogical, and Psychosocial Aspects of the Classroom Environment.” Educational Research: Theory and Practice, vol. 31, no. 2, 2020, pp. 41–55. Qualitative study of ten upper-division undergraduates — not secondary students; nine of ten struggled with verbal participation expectations and six reported physical symptoms when speaking aloud. https://files.eric.ed.gov/fulltext/EJ1274336.pdf
    7. National Paideia Center. “Paideia Seminar.” Definition of the seminar as a collaborative intellectual dialogue facilitated with open-ended questions about a text, and the pre-seminar / seminar / post-seminar cycle. https://paideia.org/blogs/npc/paideia-socratic-seminar

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Civics Project Ideas: The Mock Election That Took Over a School

    Civics Project Ideas: The Mock Election That Took Over a School

    By Clay Shumate

    One of the most fun group projects I have ever run was a mock presidential election with seventh graders. It started in my classroom during a presidential race, and by the time it was over, the entire school was lined up to vote.

    If you are looking for civics project ideas that get middle schoolers arguing about issues instead of insulting candidates, this is the one I would hand you first. Here is where it came from, how it ran, and what it honestly takes.

    Key Takeaways

    • Assign students a candidate’s positions, not a candidate to cheer for. Once they have to defend a platform, they argue the issues instead of the personalities.
    • Change the format every week. Poster board one week, PowerPoint the next, group defenses after that, for three to four weeks.
    • Keep it inside your standards. Civics carried this entire project.
    • Ask for help early. An instructional coach shaped the idea, and a community partner supplied the “I Voted” stickers.
    • A real booth, a real ballot, and running results turn a class project into something the whole building wants in on.

    Free Download · PDF

    25 Project-Based Learning Examples and Planning Template

    Twenty-five driving questions with real public products and the standards they carry, plus a one-page planning template to fill in before you commit a single class period.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Where the Idea Came From

    It was a radically difficult election, for the candidates and for the American people. I was teaching seventh grade at the time, and all of it was seeping down from home into my classroom.

    Mom and dad may very well be right about what they believe. That is not the point. Our job is to put as much information in front of students as we possibly can.

    This is where reaching out to others always helps. My instructional coach had taught the same course, so I went and asked a simple question: what is a good way to really embrace this election and bring middle schoolers into it? We landed on the one goal nobody could call partisan. Get them excited about voting when they turn 18.

    So I suggested we hold a seventh grade presidential election, with both major-party tickets on the ballot.

    How the Project Ran

    Everybody was split into groups. Each group was given topics and a side: one major-party candidate’s platform or the other’s.

    By side, I mean students took that candidate’s views. They read them, applied them, and had to use them in a group setting. Then the groups defended their arguments. Without realizing it, they were practicing a lot of the same techniques you see in Advanced Placement classes.

    The format changed every week for about three or four weeks. We did poster board. We did PowerPoint. You name it. By the time anyone voted, students had been through three to five group projects, with note taking built into every one of them.

    There was direct instruction in there too, just not much of it. You know that salt guy? That is how I use direct instruction in a project. A little sprinkle here and there.

    A little sprinkle of direct instructionIllustration of a hand pinching salt and sprinkling a few grains labeled direct instruction onto a large bowl labeled the project.Direct instructionTHE PROJECT
    A little sprinkle, not the whole shaker. The project is the meal; direct instruction is the seasoning.

    And it covered what I was required to teach. Civics was a seventh grade standard, taught all in one block, so the project was the standard rather than a break from it. You always have to be conscious of your standards, and this one was built around them from day one.

    Numbered four-week plan for a mock election civics project: week one poster board, week two slides, week three group defense of the assigned platform, and week four the vote with a booth and ballot
    Four weeks, four formats, one argument. Note taking is built into every round.

    From Name-Calling to Actual Arguments

    Before this project, the political conversation in my room sounded like “Trump sucks” and “Obama blows.” Those are two direct quotes from students, and they were the polite ones. I heard others I will not print here.

    Then something shifted. Students were arguing points. Instead of fighting about the candidates, they were applying what they had learned and arguing about the issues themselves.

    Along the way they figured out the thing I most wanted them to leave with: you vote for the candidate who, in your opinion, is going to meet the most of your needs. That lesson outlasts any single election.

    Then the Whole School Wanted In

    Word got around fast. Before long we were running the election for the entire building.

    I reached out to a community partner and asked for a few hundred “I Voted” stickers. They worked some kind of miracle and got me several hundred.

    Voting looked like the real thing. Students walked into a booth, pulled the curtain closed, and filled in a Scantron for one candidate or the other.

    We also had our own version of Fox and CNN. I think we called it FXN. I honestly do not remember, and it does not matter. At the end of every class, I ran the Scantrons or sent a student to run them. We reworked the percentages, and our news team announced the updated results before the next group voted.

    I want to be honest about one thing. A lot of the sway in those votes probably still came from home and how parents talked about the race. But by that point, my students had been fully vetted on both sides. They had argued each candidate’s positions themselves. Whatever they chose in that booth, they could tell you why.

    It was incredible. They loved it. They became super involved in the rest of the unit and the rest of the year, and they got genuinely interested in what civics was and how they could learn to vote.

    What Does the Research Say About Mock Elections?

    The evidence is real, it is modest, and the most interesting part of it says the classroom is only half the machine.

    The best-studied program of this kind is Kids Voting USA, which ran mock elections in schools alongside real ones. A CIRCLE panel study followed 497 student–parent pairs across Arizona, Colorado and Florida, then checked actual voting records two years later. Students who went through it showed lasting gains in political knowledge a full year afterward, widened the set of people they talked politics with, and — the finding I care about most — got measurably better at disagreeing with someone without it turning into a fight.

    Then the part that reframed my own project for me. The strongest path to students actually voting when they came of age ran through conversations at home. The curriculum got students paying attention to the news, the news got them talking at the dinner table, and that talk predicted turnout. The researchers also found parents shifted — they became more likely to vote themselves, without ever seeing a lesson, purely because their kid brought it up.

    I said further up that a lot of the sway in my students’ votes probably still came from home. I had that filed as the weakness of the project. The research says it is the mechanism. You are not competing with the kitchen table. You are trying to give a twelve-year-old something worth bringing to it.

    Table of findings from the Kids Voting USA research: lasting gains in political knowledge, increased willingness to disagree civilly, wider discussion networks, parents more likely to vote, no direct effect on stated intention to vote, and a sample of high school juniors and seniors rather than middle schoolers
    The numbers, including the two rows that cut against the pitch.

    The honest limits

    Four things about that evidence that a professional development slide would leave off, and you should know all of them before you quote a number to an administrator.

    • The sample was high school juniors and seniors. My project was seventh grade. The practices transfer; the findings do not automatically come with them, and anyone telling you a mock election is proven for middle schoolers is going past the data.
    • It was quasi-experimental, not a randomized trial, and attrition skewed the surviving sample toward higher-income families. Low-income and Spanish-speaking households were underrepresented.
    • There were null results. No direct effect on students’ stated intention to vote. The authors’ own summary is that effects were “not remarkable across the board.”
    • The overall strategy rates “some evidence” in the County Health Rankings review — meaning likely to work, tested more than once, trending positive, and not yet confirmed. Which is roughly where most of the good stuff in teaching sits.

    None of that is a reason to skip it. It is a reason to run it because it is good teaching and students learn the content, not because you were promised a turnout statistic.

    Four safeguards for running a classroom mock election: tell your administrator first, send a note home, assign the sides rather than letting students choose, and build a non-speaking role such as fact checker or ballot counter
    Four things that keep this safe, and none of them is difficult.

    Running Your Own Mock Election

    If you want to try this, here is the shape of it:

    1. Talk to your instructional coach or a colleague who has taught the course before you plan anything.
    2. Tell your administrator the design before you build it — the standard it covers and the assign-the-sides rule.
    3. Send a short note home explaining that students are assigned a position rather than choosing one.
    4. Assign sides. Do not let students pick the candidate they already like.
    5. Give each group specific topics, and require them to read, apply, and defend that candidate’s positions.
    6. Rotate the format weekly for three to four weeks, with note taking in every round and a sprinkle of direct instruction when they need it.
    7. Ask a community partner for “I Voted” stickers well ahead of election day.
    8. Build a booth with a curtain and use Scantrons as ballots.
    9. Start a student news team and announce running percentages before each new class votes.

    There is no presidential race this November, but a midterm works the same way. Pick one race or ballot question with two clear sides and build the groups around it.

    If you want more projects built on the same idea, there are ten more in history project ideas, including a Constitutional Convention that pairs well with this one, and a longer list in these examples of project based learning. Build the rubric before you assign anything.

    Where This Goes Wrong, and How to Not Let It

    A mock presidential election is the highest-risk project I have ever run, and every one of the risks is manageable if you handle it before week one.

    Be honest with yourself about the exposure. You are putting live partisan politics in a room full of twelve-year-olds whose families have strong feelings, in a building with a principal who will get the phone call. Four things keep it safe, and none of them is difficult.

    • Tell your administrator before you plan it, not after a parent calls. Walk them through the standard it covers and the assign-the-sides rule. An administrator who heard it from you first will defend it. One who hears it from a parent first has to investigate it.
    • Send something home. A short note explaining that students are assigned a position rather than choosing one, that the goal is understanding both platforms, and that nobody is being told who to support. Most objections never happen if the family already knows the design.
    • Assigning sides is the neutrality mechanism, and it is also the thing families worry about. The answer to “my child had to argue for someone we cannot stand” is that arguing a position you do not hold is the oldest exercise in civic education, it is what debate has always been, and it is the single most reliable way to stop students from treating an election as a team sport. Say that plainly and say it early.
    • Give a student a way out of the spotlight without leaving the project. Some students have a genuine reason not to want to stand up and argue a position — a family situation, an immigration-adjacent worry in the room, a kid who simply will not perform. Research lead, fact checker, news team, and ballot-counting are all real jobs that require the content and require no public advocacy. Build them in from the start instead of after someone has to ask.

    What It Actually Takes

    Project-based learning can be extremely powerful. It also takes a lot of work, and it takes asking others for help.

    If you are not super outgoing, this will not be easy for you at first. You will not know the answers to everything, and you are going to have to ask. That is fine — it is how I got half of what I use. If you want to talk one through, the contact page is open.

    Because the gratification from a correctly executed PBL project is nearly as good as that feeling right after you cut the lawn, when you look back on it with a cold glass of iced tea.

    Before you go: grab the free 25 Project-Based Learning Examples and Planning Template (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    What is a good civics project for middle school?

    A mock election works well because students cannot take part without learning the content. Assign each group a candidate’s positions rather than letting them pick a side, require them to read, apply and defend those positions, and finish with a real vote using a booth and a ballot. The project carries the standard instead of interrupting it.

    How long does a mock election project take?

    Mine ran about three to four weeks with a different format every week. Students completed three to five group projects, each with note taking built in, before anybody voted. You could run a smaller version in one week with a single ballot question and two groups, and it would still work.

    How do you keep a mock election from getting too political?

    Assign sides instead of letting students choose, and keep every argument on the issues rather than the candidates. That one rule does most of the work: a student defending a platform they did not pick cannot treat the election as a team sport. Tell your administrator the design before you start, and send a short note home explaining it.

    What do I say to a parent whose child was assigned a candidate the family opposes?

    Say it early rather than defending it later. Arguing a position you do not personally hold is the oldest exercise in civic education and the whole point of debate: it is how a student finds out what the other half of the country actually believes rather than what they have been told it believes. Nobody is being asked to change their mind, and no student’s own view is ever graded.

    What if a student does not want to argue a position in front of the class?

    Give them a real job that needs the content and needs no public advocacy. Research lead, fact checker, news team, ballot counting. Build those roles in from day one so a student can take one without it reading as an exemption, and so nobody has to ask in front of anybody.

    Does research actually show mock elections work?

    Partly, and it is worth knowing the shape of it. A CIRCLE panel study of Kids Voting USA found lasting gains in political knowledge, wider political discussion networks, and students who got better at disagreeing civilly — but that sample was high school juniors and seniors, not middle schoolers, the design was quasi-experimental rather than randomized, and there was no direct effect on stated intention to vote. The County Health Rankings review rates youth civics education as “some evidence.” Run it because it is good teaching, not because you were promised a statistic.

    There is no presidential race this year. Does this still work?

    Yes, and a smaller race is easier to run. A midterm, a statewide office, a local ballot question, or a single policy proposal with two defensible sides all produce the same argument structure. The presidential version is louder, not better — and a ballot question has the advantage that neither side comes with a personality attached.

    What does this cost?

    Almost nothing. Scantrons or paper ballots you already have, a refrigerator box or a rolling whiteboard for a booth, and stickers if a community partner will donate them. The expensive part is your planning time during the three to four weeks, not materials.

    Sources

    1. McDevitt, Michael, and Spiro Kiousis. Education for Deliberative Democracy: The Long-Term Influence of Kids Voting USA. CIRCLE Working Paper 22, 2004. Panel study, 497 student–parent dyads in Arizona, Colorado and Florida; high school juniors and seniors; lasting knowledge gains, wider discussion networks, greater willingness to disagree civilly; no direct effect on intention to vote; quasi-experimental, sample skewed upper-SES. https://circle.tufts.edu/sites/default/files/2019-12/WP22_Long-termInfluenceofKidsVotingUSA_2004.pdf
    2. McDevitt, Michael. Experiments in Political Socialization: Kids Voting USA as a Model for Civic Education Reform. CIRCLE Working Paper 49, 2006. Three-year longitudinal design with verified voting records; political communication in the home increased the probability of voting once students reached voting age. https://eric.ed.gov/?id=ED494074
    3. County Health Rankings & Roadmaps, University of Wisconsin Population Health Institute. Youth Civics Education. Rated “Some Evidence” — likely to work, tested more than once, results trending positive, further research needed; effects vary by curriculum and by student population. https://www.countyhealthrankings.org/strategies-and-solutions/what-works-for-health/strategies/youth-civics-education
    4. Simon, Jesse, and Bruce Merrill. “Trickle Up Political Socialization: The Impact of Kids Voting USA on Voter Turnout in Kansas.” State Politics & Policy Quarterly. Cambridge Core

    About Clay Shumate
    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Formative Assessment Strategies for Reading: Five Checks That Fit a Real Week

    Formative Assessment Strategies for Reading: Five Checks That Fit a Real Week

    By Clay Shumate

    Formative assessment strategies for reading are short, ungraded checks that show you what students understood from a text while the lesson is still running. A two-sentence gist, a flagged confusion, a prediction. They take two minutes, they are not graded, and they are only worth doing if the answer changes what you teach next.

    What follows is the honest version: five checks that survive a real secondary schedule, a routine that does not require reading 120 responses, and the numbers — including the ones that are smaller than your last in-service claimed.

    Key Takeaways

    • The check is not the intervention. What you do with the answer is. Studies where reading checks fed differentiated instruction showed nearly five times the effect of studies where they did not.
    • The honest effect is modest. Around +0.19 to +0.20, not the 0.40 to 0.70 that gets quoted. English language arts does better than math or science at 0.32.
    • Sort, do not mark. Twenty-eight gist statements is a five-minute sort into three piles. Marking them is what kills the practice by week three.
    • Nothing here requires reading aloud in front of the class. That is a design decision, not an oversight.
    • The high school evidence gap is real. An IES review found twelve adolescent literacy programs with positive effects and not one of them studied in a high school.

    Free Download · PDF

    Formative Assessment Quick-Use Pack

    A strategy decision matrix, an evidence tracker, four exit-ticket formats, and a next-day response planner — four pages built around changing the next instructional move.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Numbered list of five formative reading checks: the two-sentence gist, the confusion flag, predict then rule, the evidence pull, and the one-word question, each with a one-line description
    Five checks. None takes more than three minutes and none requires reading aloud.

    What Counts as a Formative Assessment Strategy for Reading?

    Any short check that tells you what a student took from a text in time for you to act on it. The format is negotiable. The timing is not.

    That definition rules more out than it rules in. A comprehension worksheet collected at the bell and returned Thursday is not formative — not because worksheets are bad, but because the information arrived after the decision it was supposed to inform. A five-question quiz you grade and record is summative wearing different clothes. The line is whether the answer reaches you while you can still change the next ten minutes.

    It also rules out the thing most reading checks actually measure, which is whether a student did the reading. That is a useful fact and it is not comprehension. “Did you read it” and “what did you understand” require different questions, and conflating them is how a teacher ends up with a stack of evidence that everyone read a chapter nobody understood. The broader version of this distinction is covered in the site’s list of formative assessment strategies; this page is the reading-specific case.

    What Does the Research Actually Show?

    A modest, consistent benefit — smaller than the number most professional development quotes.

    The most direct evidence is a 2022 meta-analysis by Xuan, Cheung and Sun covering 48 qualifying studies and 116,051 K–12 students. The pooled effect of formative assessment on reading achievement was +0.19 (95% CI 0.15–0.23). By grade band: kindergarten +0.28, elementary +0.16, and middle and high school +0.27 — though the meta-regression found grade-level differences were not significant once other variables were controlled. Secondary students are not the weak case here.

    The older and more quoted number is worse than people think. Kingston and Nash reviewed more than 300 studies, found only 13 with enough data to analyze, and reported a weighted mean effect of 0.20. Their conclusion is unusually blunt for a meta-analysis: the 0.40 to 0.70 figure “often claimed for the efficacy of formative assessment… is not supported by the existing research base.” The one genuinely encouraging row for a reading teacher is the subject breakdown — English language arts came in at 0.32, against 0.17 for mathematics and 0.09 for science.

    So: real, replicated, worth two minutes a day. Not transformational, and anyone selling it as transformational is selling something.

    Table of effect sizes for formative assessment on reading: plus 0.19 across K-12, plus 0.27 for middle and high school, plus 0.24 with differentiated instruction versus plus 0.05 without, plus 0.001 for technology involvement, and 0.32 for English language arts
    The numbers, including the one that contradicts the slide deck.

    What Makes the Difference Between a Check That Works and One That Does Not?

    Two moderators did most of the work, and neither is the check itself.

    First, who uses the information. In the Xuan meta-analysis, teacher-directed approaches alone produced significantly smaller effects than approaches that integrated teacher and student use of the results (coefficient −0.12, p < 0.001). Purely student-directed assessment showed no significant advantage either. The winning configuration is both: you see the pattern, and the student sees their own answer against what the text actually said.

    Second, whether anything got differentiated. Studies in which the information fed differentiated instruction showed an effect of +0.24. Studies where it did not: +0.05. That is close to the whole finding. A reading check that produces a number you record and nothing you change is worth almost nothing, and the meta-analysis says so with a p-value.

    And one thing that made no difference at all: technology. The moderator coefficient for technology involvement was +0.001 (p = 0.978). A Google Form and a sticky note perform identically. Pick the one your students will finish in ninety seconds.

    Which Five Reading Checks Fit a Real Secondary Schedule?

    These five take two to three minutes, need no materials beyond paper, and none requires a student to read aloud in front of peers.

    • The two-sentence gist. After a passage: “In two sentences, what is this saying?” Not a summary, not a main idea in a sentence frame. Two sentences in their own words. The students who can’t do it in their own words are the ones you need to find.
    • The confusion flag. One sticky note, one sentence: the part of this text I could not follow. Nothing else on it. Collected at the door. This is the highest-yield check on the list because it is the only one where the student, not you, locates the problem.
    • Pre-reading prediction, post-reading verdict. Before: one sentence on what you expect this will argue. After: right, wrong, or more complicated, and why. The post-reading half is the assessment; the pre-reading half is what makes them read looking for something.
    • The evidence pull. “Find the sentence in the text that best supports this claim.” Line number is enough. It is fast to scan and it catches the student who agrees with an argument without being able to find it on the page.
    • The one-word question. Give a single word from the text and ask what it means here. Vocabulary in context is the check with the strongest evidence behind it — explicit vocabulary instruction is one of the two practices the What Works Clearinghouse panel rates strong for adolescent literacy.

    The WWC practice guide is worth reading alongside these. Its five recommendations carry explicit evidence ratings: explicit vocabulary instruction (strong), direct and explicit comprehension strategy instruction (strong, based on five randomized experiments), extended discussion of text meaning (moderate), increasing motivation and engagement (moderate), and intensive individualized intervention for struggling readers (strong). The panel is honest about its own limits: most of the discussion studies used narrative texts, which is a meaningful caveat if you teach history or science.

    How Do You Run This Without Reading 120 Responses?

    You sort them. You do not mark them.

    This is the single practical thing that decides whether a teacher is still doing reading checks in November. Twenty-eight two-sentence gist statements is not a marking job. It is a five-minute sort into three piles — got it, partial, missed the point — and the only thing you write down is how many are in each pile. Three numbers. That is your instructional decision.

    Here is a week that fits a normal load:

    • Monday. Confusion flags on the week’s first text. Sort into piles. Whatever the biggest pile names becomes Tuesday’s opening five minutes.
    • Tuesday. Nothing collected. Teach into Monday’s biggest pile. This is the differentiation the research is actually measuring.
    • Wednesday. Two-sentence gist. Sort. Read four aloud — anonymized, and only with the writer’s permission, because every student recognizes their own sentences.
    • Thursday. Evidence pull on the same text. Ninety seconds. Scan for line numbers that are obviously wrong.
    • Friday. Nothing collected. Whatever Thursday showed, fix it in the room.

    Two collection days, two teaching-into days, one flexible. That is the whole system, and it is deliberately smaller than what a district rollout would design, because a smaller system that runs every week beats a comprehensive one that runs until October.

    Five-day routine showing Monday collect confusion flags, Tuesday teach into the biggest pile, Wednesday collect two-sentence gists, Thursday collect a fast evidence pull, and Friday fix what Thursday showed
    Two collection days. Two days spent on what they showed. That is the ratio that matters.

    What Does This Look Like in a Room?

    My room runs as a workshop, and the piece of furniture that does the most work is a whiteboard table I built, set in a corner under a lower lamp, with a rug and better chairs than the rest of the room has. Students ask to go there. That matters more than it sounds like it should.

    When a confusion flag says something I cannot fix from the front of the room, that table is where the conversation happens — not as a remediation station, because it is the seat everyone wants. A student reading a paragraph out loud at that table, with one other person, is doing the same thing that would humiliate them standing at their desk. The check tells you who needs the conversation. The room decides whether the conversation costs them anything.

    That is the part no meta-analysis measures and the part a teacher controls completely.

    Two ready-to-run versions of that table, one for grades 6–8 and one for grades 9–12, are laid out in the site’s guide to small group instruction in middle and high school.

    Where Does the Evidence Run Out?

    At high school, more precisely than most people admit.

    An IES review summarizing twenty years of adolescent literacy research screened 111 studies, found 33 meeting What Works Clearinghouse evidence standards, and identified 12 programs and practices with positive or potentially positive effects on reading comprehension, vocabulary or general literacy across grades 6–12. And then the line that belongs in every conversation about this: none of the 12 was conducted in a high school setting. The evidence base for adolescent literacy is, in practice, a middle school evidence base.

    Three more honest limits from the meta-analytic work. Small studies inflate the result badly — studies with 250 or fewer participants produced an effect of +0.45 against +0.13 for large ones, which is the usual sign that tightly-supervised implementations outperform real ones. The Xuan authors found no significant difference between published and unpublished studies, so publication bias is not the explanation, but the sample-size gap is a warning about what happens when a practice scales. And only 8 of the 48 studies came from Confucian-heritage contexts against 40 Anglophone, so the authors caution against moving interventions across cultures unadapted.

    None of that is a reason not to run a two-minute gist check. It is a reason not to build a district initiative on a number somebody rounded up.

    What Should You Do Next?

    Pick one check — the confusion flag if you want the highest yield for the least work — and run it twice a week for three weeks on a text you were teaching anyway. Sort, do not mark. Spend one whole class period in those three weeks teaching into whatever the biggest pile said, and notice whether the next set of flags moves.

    If it does, add a second check. If it does not, the problem is probably that the flags are not changing your lesson yet, which is the failure mode the research is loudest about. Two related pages are worth having open: the site’s guide to checking for understanding for the general-purpose version of these moves, and formative assessment strategies for writing for the same problem on the production side. If your reading happens on screens, the medium changes some of this — digital reading has its own comprehension research and its own accessibility upside. And if you want students doing more of the judging themselves, which is what the strongest moderator in the meta-analysis points at, start with student self-assessment.

    Before you go: grab the free Formative Assessment Quick-Use Pack (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    What is a formative assessment strategy for reading, exactly?

    It is any short, ungraded check that shows you what a student understood from a text while you can still do something about it. A two-sentence gist statement, a confusion flag, a one-question prediction, a retell in the student’s own words. The test is not the format. The test is whether the information changes what you do in the next ten minutes. If you collect it and file it, you ran a quiz.

    Should reading checks be graded?

    No, and the reason is practical rather than philosophical. The moment a gist statement counts for points, students write what they think you want instead of what they actually understood, and the check stops telling you anything. Keep them in a separate column or in no column at all. If your school requires a reading grade, take it from something summative and say out loud that the daily checks do not feed it.

    How do I do this without reading 120 responses every night?

    Sort, do not mark. A stack of 28 two-sentence gist statements is a five-minute sort into three piles — got it, partial, missed the point — and the only thing you write is the number in each pile. That number is the instructional decision. Reading every response carefully is what makes teachers abandon this in week three, and it is not what produces the benefit.

    Does the research actually support this for high school students?

    Partly, and the gap is worth knowing before anyone quotes a number at you. The meta-analysis evidence for formative assessment and reading achievement includes middle and high school studies and finds them performing at least as well as elementary. But an IES review of twenty years of adolescent literacy research found that of twelve programs with positive or potentially positive effects on reading outcomes, none was conducted in a high school setting. The practices are defensible. The claim that they are proven in a high school is not.

    Is the effect size big enough to be worth the class time?

    It is modest and it is real. The most recent meta-analysis puts formative assessment’s effect on reading achievement at +0.19 across 48 studies and 116,051 students. An earlier meta-analysis found a weighted mean of 0.20 overall and 0.32 for English language arts specifically — and explicitly rejected the 0.40 to 0.70 range that gets quoted in professional development. Two or three minutes a day for a modest, reliable gain is a good trade. A forty-minute assessment system for the same gain is not.

    What about the student who cannot read the text at all?

    A comprehension check tells you that, which is most of its value. What it must not do is tell the whole room. None of the checks here require reading aloud in front of peers, and that is deliberate. A written gist, a confusion flag on a sticky note, or a quiet conversation at a side table all give you the same information without making a struggling reader perform. The WWC panel rates intensive individualized support for struggling adolescent readers as a strong recommendation, and a daily check is how you find out who needs it — not a substitute for it.

    Does technology make these checks better?

    The evidence says not much. The same meta-analysis that found a +0.19 effect for formative assessment in reading tested whether technology involvement moderated it and found essentially nothing — a coefficient of +0.001. A form on a laptop and a sticky note produce the same result. Use whichever one your students will actually complete in ninety seconds.

    What actually predicted bigger effects, if not technology?

    Two things, and both are about what happens after the check. Approaches that integrated teacher and student use of the information beat teacher-directed approaches alone by a significant margin. And studies where the information fed differentiated instruction showed an effect of +0.24 against +0.05 for those where it did not. Collecting the data is not the intervention. Changing the next lesson is.

    Sources

    1. Xuan, Qingzhi, Alan C. K. Cheung, and Danping Sun. “The Effectiveness of Formative Assessment for Enhancing Reading Achievement in K-12 Classrooms: A Meta-Analysis.” Frontiers in Psychology, vol. 13, 2022. 48 studies, 116,051 students; pooled ES +0.19; middle/high school +0.27; differentiated instruction +0.24 vs +0.05; technology coefficient +0.001 (n.s.); small-sample studies +0.45 vs +0.13 for large. https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2022.990196/full
    2. Kingston, Neal, and Brooke Nash. “Formative Assessment: A Meta-Analysis and a Call for Research.” Educational Measurement: Issues and Practice, vol. 30, no. 4, 2011, pp. 28–37. 13 studies, 42 effect sizes; weighted mean 0.20, median 0.25; English language arts 0.32, mathematics 0.17, science 0.09. https://eric.ed.gov/?id=EJ951173
    3. Kamil, Michael L., et al. Improving Adolescent Literacy: Effective Classroom and Intervention Practices. IES Practice Guide, NCEE 2008-4027, What Works Clearinghouse, 2008. Five recommendations with evidence levels; explicit vocabulary instruction and explicit comprehension strategy instruction both rated strong. https://ies.ed.gov/ncee/wwc/docs/practiceguide/adlit_pg_082608.pdf
    4. Institute of Education Sciences, Regional Educational Laboratory Southeast. Summary of 20 Years of Research on the Effectiveness of Adolescent Literacy Programs and Practices. 111 studies screened, 33 met WWC evidence standards, 12 programs with positive or potentially positive effects — none conducted in a high school setting. https://ies.ed.gov/use-work/resource-library/report/systematic-literature-review/summary-20-years-research-effectiveness-adolescent-literacy-programs-and-practices
    5. Agarwal, Pooja K., Ludmila D. Nunes, and Janell R. Blunt. “Retrieval Practice Consistently Benefits Student Learning: A Systematic Review of Applied Research in Schools and Classrooms.” Educational Psychology Review, vol. 33, no. 4, 2021, pp. 1409–1453. 50 experiments, 5,374 participants; 57% of effect sizes medium or large. https://link.springer.com/article/10.1007/s10648-021-09595-9

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Classroom Review Games That Actually Review: Two Formats for Grades 6-12

    Classroom Review Games That Actually Review: Two Formats for Grades 6-12

    By Clay Shumate

    Classroom review games work because they force retrieval, not because they make content fun. The ones that hold up give every student a real job instead of four, keep questions short and frequent, and end with kids asking when you are playing again. The ones that fail almost always fail the same way: twenty-four students watching four.

    What follows is the mechanism the research actually supports, two formats broken down far enough to run on Monday, and the design failures that turn a review game into a spectator sport.

    Key Takeaways

    • Retrieval is the mechanism. Pulling an answer out of memory builds retention; re-reading it does not. The game is the packaging.
    • A format without a sideline job is not a whole-class review game. It is four students reviewing and the rest of the room waiting.
    • Two formats are broken down in full below: a four-square court built around eight assigned roles, and a five-minute trash-can format that needs nothing but a ball.
    • The evidence is good and it is not unlimited. Games reliably move content knowledge. The strongest school-based meta-analysis found no significant effect on metacognitive outcomes at all.
    • Short and repeated beats long and once. Fifteen minutes on three days does more than a full period the day before the test.

    Free Download · PDF

    Formative Assessment Quick-Use Pack

    A strategy decision matrix, an evidence tracker, four exit-ticket formats, and a next-day response planner — for turning what a review game surfaces into the next day’s lesson.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    Three-row table comparing recall after repeated studying and repeated testing at five minutes, two days and one week: 81 versus 75 percent, 54 versus 68 percent, and 42 versus 56 percent
    Five minutes after studying, re-reading looks better. A week later it is not close.

    Why Do Classroom Review Games Work?

    Because retrieving an answer from memory strengthens it more than reading it again does. That effect has a name — the testing effect, or retrieval practice — and it is one of the better-replicated findings in learning science.

    The landmark demonstration is worth knowing precisely, including the part that gets left out. Roediger and Karpicke had readers either study a passage repeatedly or study it once and then take repeated recall tests. Five minutes later, the repeated studiers were ahead — 81% recall against 75%. A week later the result had flipped hard: 56% for the testers, 42% for the studiers. In their second experiment the gap was wider still, 61% against 40%.

    Two honest caveats belong with that study. The participants were university undergraduates, ages 18 to 24, not teenagers. And the short-delay reversal is the most useful part of the finding for a teacher: immediately after a review game, students will often feel less prepared than they would after twenty minutes of re-reading their notes. They will be wrong, and they will tell you so. Knowing that in advance is what keeps you from abandoning a format that is working.

    The classroom evidence runs in the same direction. A five-year applied research project across more than 1,400 middle school students found that retrieval practice built into normal classroom quizzing improved long-term learning, that delayed quizzes were more potent than immediate ones, and that quizzes with feedback beat quizzes without. A 2021 systematic review screened nearly 2,000 abstracts and coded 50 classroom experiments covering 5,374 students: 57% of effect sizes showed medium or large benefits. The authors flag their own limit plainly — only 6% of the experiments were run outside WEIRD countries.

    A review game is a low-stakes retrieval quiz wearing a costume. The costume is not decoration; it is what gets a room of teenagers to willingly retrieve the same content twelve times without it feeling like a test. It belongs in the same toolbox as any other strategy built around real thinking rather than visible compliance.

    What Does the Research Not Say?

    It does not say that making something a game teaches students how to learn. This is where most review-game advice quietly oversells.

    The strongest recent synthesis of game-based learning in schools — Barz and colleagues, in Review of Educational Research — pooled studies of digital game-based learning and found a medium overall effect (g = .54), a solid cognitive effect (g = .67), a small affective-motivational effect (g = .32), and no significant effect on metacognitive outcomes. The authors found no evidence of publication bias, which makes the null finding harder to wave away.

    Two things follow. First, that meta-analysis covers digital games, and the formats below are a taped floor and a trash can — the content-knowledge finding transfers by mechanism, the effect size does not transfer automatically. Second, and more practically: a review game will help students remember what you taught. It will not teach them to study, to plan, or to notice what they do not know. If you want that, the reflective work has to be built in on purpose — which is what a separate student self-assessment routine is for.

    Why Do Most Classroom Review Games Fail?

    Because the format only has room for four players and no plan for everyone else.

    Almost every teacher has lived this period. Four students play a game-show format at the front. Twenty-four watch. The four are genuinely engaged and learning. The twenty-four are half-watching, half on their phones, and the review reached a sixth of the room. That is not a review game. It is a spectator sport with a curriculum connection.

    The failure is not the game and it is not the students. It is a design gap, and it is fixable before the first round: if a format has room for a handful of active participants, it needs a real job for everybody else. “Pay attention” is not a job.

    Numbered checklist of six design checks for a classroom review game: every student has a role, questions stay short, elimination is a rotation, round one uses warnings, a scaled version already exists, and there is a non-speaking way to play
    Six checks. Miss the first one and the rest do not matter.

    How Do You Design a Review Game That Actually Works?

    Run it against six checks before you build it or borrow it.

    • Every student has a role. The students not currently answering should be scoring, checking answers, timing or officiating — not spectating.
    • Questions stay short. Five to ten seconds per answer keeps the pace up. Save the multi-sentence explanations for bonus points, not the base question.
    • Elimination does not mean disappearing. A student who misses rotates into another job. A student who is “out” with nothing to do will find something to do, and it will not be what you wanted.
    • The first round uses warnings, not eliminations. Students are learning the format and the content at the same time; do not punish them for the first. This is the same logic behind teaching a routine explicitly before holding anyone to it.
    • A scaled version exists before you need it. Smaller ball, smaller court, no running, seated variant. Planned, not improvised.
    • There is a non-speaking way to play. Scorekeeper, fact checker and timekeeper all require knowing the content and none requires answering out loud. Build that in from the start rather than after a student has to ask for it.

    Those six are the difference between a format that runs itself after the first week and one you referee forever. Most of them are the same principle underneath any student-centered activity: the structure works when students genuinely own a piece of it, not when they have permission to join in.

    What Does a Full Format Look Like? The 4 Square Review Game

    A four-square court, four players, and eight assigned jobs for everyone else. It is built specifically to solve the spectator problem, and the roles do not change by subject — only the question cards do.

    Setup. Tape a 16′ × 16′ court split into four 8′ × 8′ squares, labeled A through D. Square A is the top position; D is the entry point. Split the class into four teams, one player from each team on the court, the next four in a marked Sub Zone.

    The round. Before a player serves or returns, the Question Master asks a content question. The player has about five seconds. A correct answer makes the ball live and standard four-square rules apply — one bounce before the return. A wrong answer eliminates the player unless a challenge changes the ruling, and they rotate to the end of their team’s line. The next player enters at D and everyone on the court steps up one square.

    The eight jobs. This is the part that makes it a whole-class format:

    • Question Master reads the questions and varies the difficulty.
    • Referee calls rule violations and eliminations.
    • Fact Checker confirms answers against the key and accepts reasonable equivalents.
    • Timekeeper runs a two-to-three-minute round clock and calls time.
    • Scorekeeper tracks points for court play, sideline jobs and challenges.
    • Challenger can dispute a ruling — but only with a reason or a piece of evidence, which turns disagreement into an academic move instead of an argument with you.
    • Hype Squad keeps energy up with brief, respectful encouragement and gives away no answers.
    • Sub Rotation manages the line so the next players are always ready.

    Roles rotate every round, so the student who spent one round scoring is playing the next. Scoring stays simple: a point for a correct answer, a point for a successful serve or return, a point for reaching Square A, a point for a clean sideline call, two points for a successful challenge. Small bonuses for a team lifeline, strong supporting evidence, or a team that covers every assigned role keep the sideline jobs from reading as busywork.

    Adapting 4 Square to Any Subject

    The court and the roles never change. The question bank does. A social studies class can run cause-and-effect recall one round and source-analysis the next. A math class swaps in problem-solving questions with a slightly longer answer window. An English class runs vocabulary in context. Build more questions than you think you need, sorted by difficulty, so the Question Master can adjust live when a round runs hot or cold. Running out of questions kills the pace faster than anything else on this page.

    List of eight assigned student roles for a review game: Question Master, Referee, Fact Checker, Timekeeper, Scorekeeper, Challenger, Hype Squad and Sub Rotation, each with a one-line description of the job
    Four players. Eight jobs. Nobody spectates.

    What If You Only Have Five Minutes to Set Up? Trashketball

    A question, a correct answer, then a shot at the trash can from a distance the student picks. Trashketball is a long-standing, widely used classroom format, and it earns its place here for one reason: it delivers the same retrieval mechanism with no court to tape and no roles to assign.

    I did not invent it and I want to be clear about that, because how I got it is the more useful part of the story. In my first year I did not have many review formats that were actually enjoyable, and another teacher on my hall noticed. I asked. They walked me through Trashketball, and I have been running some version of it ever since.

    Keeping your ears open and your mouth shut pays off some of the time. Asking a direct question pays off faster. If you are in your first years and your review days are flat, the fix is probably already being run two doors down by somebody who would be glad you asked. That is not a soft lesson about collegiality — it is the cheapest professional development available and it costs one conversation.

    The basic loop is a question to a team, a short answer window, and a shot worth more points from further out. The self-chosen distance is the good part — it gives a student who is sure of the content a way to press the advantage and a student who is not a way to stay in without being exposed.

    Two upgrades are worth building in. First, give the non-shooting teams a job: one team fact-checks the answer, one keeps score, and both rotate. Second, have each question lead into a one-or-two-sentence note rather than a full explanation, so movement and note-taking land in the same fifteen minutes instead of competing for them.

    The lesson from both formats is the same. A physically engaged class and a content-reviewing class are not in competition. They can be the same fifteen minutes.

    What Do You Watch For While It Runs?

    The answers you overhear, not the score.

    My room runs as a workshop — tables grouped rather than in rows, me on a wheeled stool moving between them instead of standing at the front. That habit is worth more during a review game than during almost anything else I do. Circulating while a round plays, you hear a group land on a confident wrong answer before it reaches anybody’s notes, and you can fix it in the next question instead of on the test. The scoreboard tells you who is winning. The sideline conversation tells you what you still have to teach.

    That is also the honest answer to whether a review game is “real” instruction. It is, if you are listening. It is not, if you are refereeing.

    What Are the Most Common Mistakes?

    • No job for the sideline. The single most common failure, and the one that turns a review game into a spectator sport.
    • Elimination with nowhere to go. Rotate eliminated students into a job instead of into a chair.
    • No warning period. Hard eliminations from question one punish students for not knowing a format they have never played.
    • Speed rewarded over accuracy. A five-second window is fine. A game that rewards the fastest guess over the most defensible answer teaches the wrong thing.
    • The same students in the low-status seat every time. If the fast recallers always play and the same kids always keep score, you have built a ranking system, not a review game. Rotate on a schedule, not on volunteering.
    • No safety plan. Any physical format needs its scaled version ready before somebody asks for it.
    • Treating it as a time-filler. If the game does not require retrieving real content, it is recess with a curriculum label.

    What Should You Do Next?

    Pick one format — 4 Square if your class needs structure and assigned roles, Trashketball if you need something running in five minutes — and build a question bank of about thirty questions sorted into three difficulty tiers. Run it for fifteen minutes on three separate days before your next unit test rather than for a full period the day before it. The spacing is doing as much work as the game.

    Then watch two things. Whether the sideline stays engaged, which tells you if the roles are real. And whether students ask when you are playing again, which tells you whether you have a repeatable Friday or a one-time novelty. A review game also fits an awkward period well — the kind covered in teaching after a long weekend, when a straight lecture is fighting uphill no matter how good it is.

    Before you go: grab the free Formative Assessment Quick-Use Pack (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Do classroom review games actually work, or do they just feel productive?

    They work when they force retrieval — students pulling an answer out of memory rather than recognizing it on a page in front of them. That mechanism has been tested directly. Roediger and Karpicke found repeated testing beat repeated studying by 14 points of recall a week later. The game format is the delivery system, not the mechanism. A game where students look up every answer is a worksheet with a ball in it.

    How do I keep 28 students engaged when only four can play at a time?

    Assign every non-playing student a real job and rotate the jobs every round — scorekeeper, fact checker, timekeeper, referee, challenger. “Pay attention” is not a job, it is a hope. If a format has no work for the twenty-four students who are not currently answering, it is not a whole-class review game and you should expect it to behave like one.

    Won’t the competition discourage my students who are already behind?

    It can, and the design has to answer for it. Three things keep it from landing on the same students every time: rotate the jobs so nobody sits in the low-status seat all period, score sideline work as heavily as court play, and score teams rather than individuals so a wrong answer is absorbed. If a specific student is visibly shrinking during a game, give them the Fact Checker role — it is high-status, it requires knowing the content, and it never requires answering in front of the room.

    What if my class is not ready for something physical?

    Scale it before ruling it out. A softer or larger ball, a smaller court, a no-running rule, or a fully seated version of every role keeps the structure intact. The roles and scoring do not depend on anyone moving. Plan the scaled version before the first round rather than improvising it after someone gets hurt or someone gets left out.

    How long should a review game run?

    Rounds of two to three minutes with a role rotation between them, for a total of fifteen to twenty minutes. That is long enough to cycle the whole class through several retrievals and short enough that it does not eat the period. A game that runs forty minutes has stopped being retrieval practice and started being an activity.

    Do these formats work outside social studies?

    Yes, because nothing in the court, the roles or the scoring references a subject. Only the question bank changes. A math class can lengthen the answer window for a multi-step problem; an English class can run vocabulary-in-context questions; a science class can run prediction questions. The one thing that does not travel is a question bank written the morning of.

    A student is arguing with my ruling. How do I handle it without losing the room?

    Build the argument into the format. A Challenger role that requires a brief reason or a piece of evidence turns a dispute into an academic move and takes you out of the middle of it. Award points for a successful challenge. Students who would never volunteer an answer will absolutely argue a call, and a challenge is a retrieval too.

    Is a review game a good use of a class period before a test?

    It is a good use of fifteen to twenty minutes of one, run more than once. The retrieval research is consistent that spaced, repeated retrieval beats one long session, so a short game on three different days will do more than a single period-long game the day before the exam. It is also worth being honest about what games do not do: the best meta-analysis of game-based learning in schools found solid gains in content knowledge and no significant effect on metacognitive outcomes. A game will help students remember. It will not, on its own, teach them how to study.

    Sources

    1. Roediger, Henry L., III, and Jeffrey D. Karpicke. “Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention.” Psychological Science, vol. 17, no. 3, 2006, pp. 249–255. Participants were university undergraduates aged 18–24. https://journals.sagepub.com/doi/10.1111/j.1467-9280.2006.01693.x
    2. Agarwal, Pooja K., Ludmila D. Nunes, and Janell R. Blunt. “Retrieval Practice Consistently Benefits Student Learning: A Systematic Review of Applied Research in Schools and Classrooms.” Educational Psychology Review, vol. 33, no. 4, 2021, pp. 1409–1453. 50 coded experiments, 5,374 participants; 57% of effect sizes medium or large; only 6% of experiments conducted outside WEIRD countries. https://link.springer.com/article/10.1007/s10648-021-09595-9
    3. Agarwal, Pooja K., Patrice M. Bain, and Roger W. Chamberlain. “The Value of Applied Research: Retrieval Practice Improves Classroom Learning and Recommendations from a Teacher, a Principal, and a Scientist.” Educational Psychology Review, vol. 24, no. 3, 2012, pp. 437–448. Five-year project with more than 1,400 middle school students. https://eric.ed.gov/?id=EJ977134
    4. Barz, Nathalie, Manuela Benick, Laura Dörrenbächer-Ulrich, and Franziska Perels. “The Effect of Digital Game-Based Learning Interventions on Cognitive, Metacognitive, and Affective-Motivational Learning Outcomes in School: A Meta-Analysis.” Review of Educational Research, 2024. Overall g = .54; cognitive g = .67; affective-motivational g = .32; no significant effect on metacognitive outcomes; no evidence of publication bias. https://journals.sagepub.com/doi/10.3102/00346543231167795

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

  • Formative Assessment Strategies for Writing: Checks That Fit a Real Week

    Formative Assessment Strategies for Writing: Checks That Fit a Real Week

    By Clay Shumate

    Formative assessment strategies for writing are the checks a teacher runs while a piece is still being written, so the information can still change the draft. Exit tickets and comprehension checks do not transfer here. Writing has its own set, because the thing being assessed takes days, gets revised, and is too long to read thirty times a week.

    That last constraint is the real problem. Almost every writing-feedback system that gets abandoned was abandoned because it required reading every draft. What follows is what the research supports, the number that should worry you, and a set of checks that fit an ordinary secondary schedule.

    Key Takeaways

    • Feedback works, and the size depends on who gives it. A meta-analysis found adult feedback at an effect size of 0.87, self-evaluation at 0.62 and peer feedback at 0.58 — but that study covered grades 1–8.
    • Two popular practices showed no meaningful effect in the same analysis. Teacher progress monitoring and the 6+1 Trait Writing model did not improve writing quality.
    • The federal practice guide rates the assessment recommendation as its weakest. Of three recommendations for teaching secondary writing, “use assessments to inform instruction and feedback” carries minimal evidence. That is worth knowing before anyone sells you a system.
    • Assess one thing at a time. A draft marked for everything gets revised for nothing. Pick the criterion the lesson taught and check only that.
    • Never put a grade on a draft you want revised. Self-assessment that counts toward a mark stops being honest, and the same logic applies to a draft.

    Free Download · PDF

    Formative Assessment Quick-Use Pack

    A strategy decision matrix, an evidence tracker, four exit-ticket formats, and a next-day response planner — four pages built around changing the next instructional move.

    Download the free PDF

    Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

    What Makes a Writing Check Formative?

    Timing and use. A check is formative if it happens while the piece can still change and if somebody acts on it. Both conditions. A rubric applied to a final draft, however detailed, is a grade with extra steps.

    That rules out more of the standard toolkit than teachers expect. Marking a finished essay is summative no matter how much you write in the margin. A reading quiz is not a writing check. A participation grade for peer editing measures compliance with a procedure, not the quality of anything. The general versions of these checks are covered in the list of formative assessment strategies; this page is about the ones built for a text that takes a week.

    What is distinctive about writing is that the product has stages. A thesis exists before the paragraph does, an outline before the draft, a draft before the revision. Each stage is a place to check something cheaply, while it is still cheap to fix. The whole art is checking the earliest stage at which the problem is visible — because a thesis fixed on Monday saves five paragraphs that never had to be written.

    What Does the Research Actually Show?

    That feedback on writing works, that two popular practices do not, and that the evidence for the whole assessment-driven approach is weaker than its popularity suggests. All three are worth having straight.

    The central study is Graham, Hebert and Harris’s 2015 meta-analysis in the Elementary School Journal, which pooled experimental studies of formative writing assessment and reported effect sizes by who does the assessing:

    PracticeEffect on writing quality
    Adult feedback0.87
    Students evaluating their own writing0.62
    Peer feedback0.58
    Computer feedback0.38
    Teachers monitoring student progressno meaningful improvement
    The 6+1 Trait Writing modelno meaningful improvement

    Three things to notice. First, the grade band is 1 through 8. Only the top three grades of that range are secondary, and none of it is high school. The practices are reasonable to carry upward; the numbers do not travel with them.

    Second, the two null results are the most useful rows in the table. Progress monitoring — tracking writing scores over time to inform instruction — showed no meaningful improvement, and neither did 6+1 Trait, which is a widely adopted framework. That does not make either worthless, but it does mean a department should not treat adopting them as having addressed writing feedback.

    Third, student self-evaluation at 0.62 nearly matched adult feedback and beat peer feedback. For a teacher with 120 students, that is the most consequential number in the table, because it is the only one that does not scale with your reading time.

    Now the part most articles omit. The What Works Clearinghouse practice guide on teaching secondary students to write effectively, written by a panel including Steve Graham, Jill Fitzgerald, Linda Friedrich, Katie Greene, James Kim and Carol Booth Olson, makes three recommendations and rates the evidence behind each. Explicitly teaching writing strategies through a model–practice–reflect cycle is rated strong. Integrating writing and reading using exemplar texts is rated moderate. Using assessments to inform instruction and feedback is rated minimal.

    Minimal is the lowest rating in that system, and it is attached to the recommendation this article is about. Read it as a statement about proportion rather than a reason to stop. If you have a fixed amount of energy for improving writing in your classroom, the guide says to spend it first on explicitly teaching strategies and modeling them — and the assessment practices below are what you run inside that, not instead of it.

    Table of effect sizes on writing quality from a formative assessment meta-analysis: adult feedback 0.87, student self-evaluation 0.62, peer feedback 0.58, computer feedback 0.38, and no meaningful effect for teacher progress monitoring or the 6+1 Trait Writing model
    Graham, Hebert and Harris (2015), grades 1-8. The two null rows are the most useful in the table.

    Which Checks Fit a Secondary Schedule?

    The ones that read a sentence rather than a draft. Every strategy below is designed around the fact that you cannot read 120 full drafts and still have a weekend.

    • The thesis-only check. Collect one sentence. Read all of them in ten minutes, sort into three piles — arguable, too broad, not a claim — and hand them back with the pile name. Fixing this on day one prevents most of what you would otherwise write in margins on day six.
    • The one-criterion read. Announce that today’s read is only about evidence, or only about topic sentences. Mark only that. A draft marked for everything gets revised for nothing, because a student facing thirty marks does not know where to start and usually starts with the commas.
    • The highlight-and-justify. Students highlight the sentence in their own draft that meets a specific criterion, and write one line explaining why. If they cannot find it, that is the check — and they have found it themselves, which is the part that matters.
    • The first-paragraph conference. Two minutes per student, on the opening paragraph only, while the rest write. Twelve students a period, everybody covered across a week.
    • The anonymous exemplar. Put two short pieces of writing on the board with names removed — one that does the thing, one that nearly does it — and have the class say which and why. This is the model–practice–reflect cycle the WWC rates as strong evidence, run as a five-minute check. Ask before you use a student’s work, even anonymized, because the writer always recognizes their own sentences and so do the people sitting next to them. Asking takes ten seconds, almost nobody says no, and the ones who do have a reason. Writing your own two versions, or using last year’s with permission, works just as well.
    • The revision log. One line per revision: what changed, and why. It takes a student thirty seconds and tells you whether feedback was used, which is the only question a progress tracker was ever trying to answer.

    Notice what is missing: reading every draft, writing extended comments, and any system requiring a spreadsheet. Those are the things that get abandoned in October, and abandonment is the real failure mode — not choosing the second-best check. If you would rather start from something printed, the free formative assessment quick-use pack has a general version to adapt.

    Graphic listing six formative writing checks: the thesis-only check, the one-criterion read, the highlight-and-justify, the first-paragraph conference, the anonymous exemplar, and the revision log
    Every one of these reads a sentence rather than a draft. That is what makes them survive past October.

    How Do You Make Self-Evaluation Work?

    Give the criteria first, keep it out of the gradebook, and ask for evidence rather than a rating. At 0.62 this is the best return per minute of teacher time in the whole table, and all three conditions matter.

    Heidi Andrade’s 2019 critical review in Frontiers in Education supplies the design rules. Criterion-referenced self-assessment showed main effects on every criterion assessed, and concrete task-specific criteria outperformed vague competence-based ones — a result she attributes to Fastré and colleagues. In writing terms, “every claim is followed by a quotation and an explanation of it” is checkable. “Uses evidence effectively” is not.

    The decisive rule is about grading. Andrade cites Tejeiro and colleagues, where self-assessment counted toward the final grade: overestimation rose dramatically and no correlation remained between the instructor’s assessment and the student’s. Run formatively, agreement improved substantially, and all twenty studies in her review that used self-assessment formatively showed a positive association with learning.

    Andrade makes one more point worth carrying. There is little evidence that inaccurate self-assessment produces worse learning, and students act on their predictions regardless of accuracy. You are not trying to make a student’s judgment match yours. You are trying to make them read their own draft as a reader. The broader version of the practice is in the student self assessment guide; writing only changes what sits in the criteria.

    Is Peer Feedback Worth the Class Time?

    Yes at 0.58, and only if you narrow what you ask for. Unstructured peer review produces “I liked it, maybe add more detail,” which is the outcome most teachers have seen and correctly concluded is a waste of twenty minutes.

    There is a useful finding from an adjacent literature. Falchikov and Goldfinch’s meta-analysis of 48 higher-education studies comparing peer marks with teacher marks found that agreement was closest when students made a global judgment against well-understood criteria, and worse when asked to break the judgment into many separate dimensions and score each. Those were undergraduates marking work, not teenagers giving revision advice, so treat it as a design hint rather than a transferred result — but the hint points somewhere useful: the twelve-box peer editing checklist is probably the worst available format.

    What works better is narrower and more concrete:

    • Ask for a location, not a judgment. “Underline the sentence where the argument actually starts.” A reader can do that honestly; “rate the organization” they cannot.
    • Ask what the reader could not follow. This is the one thing a peer knows that you do not, because they read it without already knowing what the writer meant.
    • Ask for one question, not one suggestion. Suggestions are advice from a novice. A genuine question — “is this the same person as in paragraph two?” — is information the writer can act on without deferring to anyone.

    Two cautions. Peer feedback has a social cost that teacher feedback does not. A fifteen-year-old asked to critique a classmate’s writing is managing a relationship as well as a text, and most will resolve that tension by being vague. Asking for locations and questions rather than evaluations removes the tension rather than asking students to override it. And decide who reads what before you start. Writing is personal in a way a math worksheet is not; a student writing about something that matters to them should know in advance whether a classmate will see it, and should have a way out.

    A One-Week Routine That Does Not Require Reading Every Draft

    Four checks across a week, none of which takes you more than fifteen minutes.

    1. Day one — collect the thesis only. One sentence per student. Sort into three piles, hand back with the pile name, and give the “not a claim” pile five minutes to try again.
    2. Day two — the anonymous exemplar. Two openings on the board, names off, one working and one nearly working. The class names the difference. This is the modeling step, and it is the one with the strongest evidence behind it.
    3. Day three — highlight and justify. Students find the sentence in their own draft that meets today’s single criterion and write one line saying why. Walk the room and read over shoulders; collect nothing.
    4. Day four — peer question round. Swap drafts. Each reader underlines where the argument starts and writes one genuine question. Ten minutes, no checklist.

    Then the revision log on day five: one line per change, what and why. That log is the whole assessment record, and unlike a score tracker it answers the question that actually matters, which is whether any of this changed the draft.

    Two adjustments. First, differentiate the container rather than the criterion — a student who cannot produce the written justification quickly can say it to you in the last minute of class, and a student working in a second language can highlight and point. The judgment against criteria is what has to survive. Second, if a student’s draft has a problem that is not today’s criterion, note it for yourself and leave it. You will get to organization on the week you are checking organization, and a student who receives one correctable thing at a time actually corrects it. For the general version of this discipline, see checking for understanding.

    Four Ways Writing Feedback Fails

    All four are common, and all four are about the system rather than the student.

    Everything gets marked. A draft returned with thirty corrections communicates that the piece is bad, not what to do next. Students respond by fixing the easiest marks, which are almost always the mechanical ones. One criterion per read is not a compromise — it is what makes revision possible.

    A grade goes on the draft. Once a number is attached, the piece is finished in the student’s mind, and the comments underneath it become an explanation of the number rather than instructions for a revision. If you need the draft in the gradebook, grade completion rather than quality, and say which you are doing.

    The one-criterion rule is also the thing families most often ask about, usually in the form of why the teacher did not correct all the errors. It deserves a straight answer rather than a defensive one: every error was noticed, and marking all of them is what produces a draft a student fixes the commas in and hands back otherwise unchanged. The errors get their turn, one at a time, on the assignment where that is the thing being taught. Said in advance, in a sentence on the assignment sheet, that lands as a deliberate method. Said after a parent email, it sounds like an excuse.

    There is no time to act on it. Feedback returned the day the final is due is a post-mortem. If the schedule does not contain a revision block after the feedback, the feedback is decorative, and it is more honest to admit that than to keep writing comments into a void.

    Graphic listing four ways writing feedback fails: everything gets marked, a grade goes on the draft, there is no time to act on it, and a framework is adopted and called done
    All four are about the system rather than the student.

    The department adopts a framework and calls it done. This is where the two null results earn their keep. Adopting 6+1 Trait or installing a progress-monitoring spreadsheet is a visible action that showed no meaningful effect on writing quality in the meta-analysis. The things that did work — someone reads a piece of writing and responds to it, or a student reads their own against criteria — are less visible on a plan, and are the ones worth protecting time for.

    What to Try on the Next Assignment

    Take the next piece of writing you have already assigned. Collect the thesis on its own, before anything else exists, and sort the sentences into three piles. That single move costs you ten minutes and changes more drafts than any set of margin comments you will write later.

    Then pick one criterion for the whole assignment and check only that — in the exemplar, in the self-evaluation, in the peer round, in your own read. Put a revision block on the calendar before you give any feedback, and keep a number off the draft.

    None of this is a system, and that is deliberate. The evidence for assessment-driven writing instruction is rated minimal by the people best placed to judge it, while explicit strategy instruction and modeling are rated strong. The checks here are worth running because they are cheap and they surface problems early. They are not a substitute for teaching students how to write, and any resource presenting them as one has the proportions backwards.

    Before you go: grab the free Formative Assessment Quick-Use Pack (PDF) — ElevateTheNorm.com branded, printable, no email required.

    Frequently Asked Questions

    Do these effect sizes apply to high school students?

    Not directly, and it matters. The meta-analysis behind the headline numbers — adult feedback 0.87, self-evaluation 0.62, peer feedback 0.58 — covered grades 1 through 8, so only its top three grades are secondary at all and none is high school. The practices are reasonable to carry upward because the mechanisms are not age-specific, but the numbers should not be quoted as a high school result. The federal practice guide that does cover secondary writing rates the assessment recommendation as having minimal evidence.

    Should I put a grade on a draft?

    Not if you want it revised. Once a number is attached, most students treat the piece as finished and read the comments as justification for the mark rather than instructions for a revision. The self-assessment research points the same way: when self-evaluation counted toward a grade, overestimation rose sharply and agreement with the instructor disappeared. If a draft has to appear in the gradebook, grade completion rather than quality and tell students that is what you are doing.

    Is 6+1 Trait Writing a waste of time?

    That is stronger than the evidence supports. What the meta-analysis found is that 6+1 Trait showed no meaningful improvement in writing quality, and the same was true of teachers monitoring student progress over time. That is a real finding and a department should not treat adopting the framework as having addressed writing feedback. It is not a finding that a shared vocabulary for talking about writing is harmful — it is a finding that the vocabulary alone does not move the writing.

    How do I give feedback to 120 students without losing every weekend?

    Stop reading whole drafts. Collect one sentence rather than one essay; mark one criterion rather than everything; run two-minute conferences on opening paragraphs while the rest write; and lean on self-evaluation, which came in at 0.62 and is the only practice in the table that does not scale with your reading time. A teacher who reads every draft thoroughly in September and nothing at all by November has given less useful feedback than one who reads one paragraph from everybody every week.

    Does peer feedback actually help, or is it busywork?

    It helped at an effect size of 0.58, which is real — but what most classrooms run is not what was studied. Unstructured peer review produces “I liked it, add more detail.” Narrow the ask instead: have readers underline where the argument starts, say what they could not follow, and write one genuine question rather than one suggestion. Those ask for information a peer actually has. Also settle who reads what before you begin, because writing is more personal than a worksheet and a student should know in advance.

    What if a student’s draft has problems that are not this week’s criterion?

    Note them for yourself and leave them alone. A student who receives one correctable thing at a time usually corrects it; a student who receives thirty fixes the commas. Keep your own running list and let it decide what the criterion is for the next assignment — that list is more useful as a planning document than as margin notes, and it means the pattern across the class shapes what you teach next rather than disappearing into thirty separate drafts.

    How is this different from just marking essays carefully?

    Timing and use. Marking a finished essay is summative however detailed it is, because nothing about that piece can change afterwards. A formative check happens while the writing is still in progress and is followed by time to act on it. The practical test is simple: if there is no revision block on the calendar after the feedback goes back, what you did was grading, and calling it formative assessment does not make it function like one.

    Where should I spend my energy if I can only change one thing?

    Not here, according to the people who reviewed the evidence. The What Works Clearinghouse panel rates explicitly teaching writing strategies through a model–practice–reflect cycle as strong evidence, integrating reading and writing with exemplar texts as moderate, and using assessments to inform instruction and feedback as minimal. If you have one change in you this year, make it the modeling. The checks on this page are cheap enough to run inside that work, and they are not a replacement for it.

    Sources

    1. Graham, Steve, Michael Hebert, and Karen R. Harris. “Formative Assessment and Writing: A Meta-Analysis.” Elementary School Journal, vol. 115, no. 4, 2015, pp. 523–547. https://eric.ed.gov/?id=EJ1068976
    2. Graham, Steve, Jill Fitzgerald, Linda D. Friedrich, Katie Greene, James S. Kim, and Carol Booth Olson. “Teaching Secondary Students to Write Effectively” (practice guide summary). What Works Clearinghouse, Institute of Education Sciences, U.S. Department of Education. https://ies.ed.gov/ncee/wwc/Docs/PracticeGuide/wwc_secwrit_summary_053117.pdf
    3. Andrade, Heidi L. “A Critical Review of Research on Student Self-Assessment.” Frontiers in Education, vol. 4, art. 87, 2019. https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2019.00087/full (The Tejeiro et al. 2012 and Fastré et al. 2010 findings are reported in this review; the primary papers were not read directly.)
    4. Falchikov, Nancy, and Judy Goldfinch. “Student Peer Assessment in Higher Education: A Meta-Analysis Comparing Peer and Teacher Marks.” Review of Educational Research, vol. 70, no. 3, 2000, pp. 287–322. https://eric.ed.gov/?id=EJ630369

    About Clay Shumate

    Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.