By Clay Shumate
Formative assessment strategies for reading are short, ungraded checks that show you what students understood from a text while the lesson is still running. A two-sentence gist, a flagged confusion, a prediction. They take two minutes, they are not graded, and they are only worth doing if the answer changes what you teach next.
What follows is the honest version: five checks that survive a real secondary schedule, a routine that does not require reading 120 responses, and the numbers — including the ones that are smaller than your last in-service claimed.
Key Takeaways
- The check is not the intervention. What you do with the answer is. Studies where reading checks fed differentiated instruction showed nearly five times the effect of studies where they did not.
- The honest effect is modest. Around +0.19 to +0.20, not the 0.40 to 0.70 that gets quoted. English language arts does better than math or science at 0.32.
- Sort, do not mark. Twenty-eight gist statements is a five-minute sort into three piles. Marking them is what kills the practice by week three.
- Nothing here requires reading aloud in front of the class. That is a design decision, not an oversight.
- The high school evidence gap is real. An IES review found twelve adolescent literacy programs with positive effects and not one of them studied in a high school.
Free Download · PDF
Formative Assessment Quick-Use Pack
A strategy decision matrix, an evidence tracker, four exit-ticket formats, and a next-day response planner — four pages built around changing the next instructional move.
Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.

What Counts as a Formative Assessment Strategy for Reading?
Any short check that tells you what a student took from a text in time for you to act on it. The format is negotiable. The timing is not.
That definition rules more out than it rules in. A comprehension worksheet collected at the bell and returned Thursday is not formative — not because worksheets are bad, but because the information arrived after the decision it was supposed to inform. A five-question quiz you grade and record is summative wearing different clothes. The line is whether the answer reaches you while you can still change the next ten minutes.
It also rules out the thing most reading checks actually measure, which is whether a student did the reading. That is a useful fact and it is not comprehension. “Did you read it” and “what did you understand” require different questions, and conflating them is how a teacher ends up with a stack of evidence that everyone read a chapter nobody understood. The broader version of this distinction is covered in the site’s list of formative assessment strategies; this page is the reading-specific case.
What Does the Research Actually Show?
A modest, consistent benefit — smaller than the number most professional development quotes.
The most direct evidence is a 2022 meta-analysis by Xuan, Cheung and Sun covering 48 qualifying studies and 116,051 K–12 students. The pooled effect of formative assessment on reading achievement was +0.19 (95% CI 0.15–0.23). By grade band: kindergarten +0.28, elementary +0.16, and middle and high school +0.27 — though the meta-regression found grade-level differences were not significant once other variables were controlled. Secondary students are not the weak case here.
The older and more quoted number is worse than people think. Kingston and Nash reviewed more than 300 studies, found only 13 with enough data to analyze, and reported a weighted mean effect of 0.20. Their conclusion is unusually blunt for a meta-analysis: the 0.40 to 0.70 figure “often claimed for the efficacy of formative assessment… is not supported by the existing research base.” The one genuinely encouraging row for a reading teacher is the subject breakdown — English language arts came in at 0.32, against 0.17 for mathematics and 0.09 for science.
So: real, replicated, worth two minutes a day. Not transformational, and anyone selling it as transformational is selling something.

What Makes the Difference Between a Check That Works and One That Does Not?
Two moderators did most of the work, and neither is the check itself.
First, who uses the information. In the Xuan meta-analysis, teacher-directed approaches alone produced significantly smaller effects than approaches that integrated teacher and student use of the results (coefficient −0.12, p < 0.001). Purely student-directed assessment showed no significant advantage either. The winning configuration is both: you see the pattern, and the student sees their own answer against what the text actually said.
Second, whether anything got differentiated. Studies in which the information fed differentiated instruction showed an effect of +0.24. Studies where it did not: +0.05. That is close to the whole finding. A reading check that produces a number you record and nothing you change is worth almost nothing, and the meta-analysis says so with a p-value.
And one thing that made no difference at all: technology. The moderator coefficient for technology involvement was +0.001 (p = 0.978). A Google Form and a sticky note perform identically. Pick the one your students will finish in ninety seconds.
Which Five Reading Checks Fit a Real Secondary Schedule?
These five take two to three minutes, need no materials beyond paper, and none requires a student to read aloud in front of peers.
- The two-sentence gist. After a passage: “In two sentences, what is this saying?” Not a summary, not a main idea in a sentence frame. Two sentences in their own words. The students who can’t do it in their own words are the ones you need to find.
- The confusion flag. One sticky note, one sentence: the part of this text I could not follow. Nothing else on it. Collected at the door. This is the highest-yield check on the list because it is the only one where the student, not you, locates the problem.
- Pre-reading prediction, post-reading verdict. Before: one sentence on what you expect this will argue. After: right, wrong, or more complicated, and why. The post-reading half is the assessment; the pre-reading half is what makes them read looking for something.
- The evidence pull. “Find the sentence in the text that best supports this claim.” Line number is enough. It is fast to scan and it catches the student who agrees with an argument without being able to find it on the page.
- The one-word question. Give a single word from the text and ask what it means here. Vocabulary in context is the check with the strongest evidence behind it — explicit vocabulary instruction is one of the two practices the What Works Clearinghouse panel rates strong for adolescent literacy.
The WWC practice guide is worth reading alongside these. Its five recommendations carry explicit evidence ratings: explicit vocabulary instruction (strong), direct and explicit comprehension strategy instruction (strong, based on five randomized experiments), extended discussion of text meaning (moderate), increasing motivation and engagement (moderate), and intensive individualized intervention for struggling readers (strong). The panel is honest about its own limits: most of the discussion studies used narrative texts, which is a meaningful caveat if you teach history or science.
How Do You Run This Without Reading 120 Responses?
You sort them. You do not mark them.
This is the single practical thing that decides whether a teacher is still doing reading checks in November. Twenty-eight two-sentence gist statements is not a marking job. It is a five-minute sort into three piles — got it, partial, missed the point — and the only thing you write down is how many are in each pile. Three numbers. That is your instructional decision.
Here is a week that fits a normal load:
- Monday. Confusion flags on the week’s first text. Sort into piles. Whatever the biggest pile names becomes Tuesday’s opening five minutes.
- Tuesday. Nothing collected. Teach into Monday’s biggest pile. This is the differentiation the research is actually measuring.
- Wednesday. Two-sentence gist. Sort. Read four aloud — anonymized, and only with the writer’s permission, because every student recognizes their own sentences.
- Thursday. Evidence pull on the same text. Ninety seconds. Scan for line numbers that are obviously wrong.
- Friday. Nothing collected. Whatever Thursday showed, fix it in the room.
Two collection days, two teaching-into days, one flexible. That is the whole system, and it is deliberately smaller than what a district rollout would design, because a smaller system that runs every week beats a comprehensive one that runs until October.

What Does This Look Like in a Room?
My room runs as a workshop, and the piece of furniture that does the most work is a whiteboard table I built, set in a corner under a lower lamp, with a rug and better chairs than the rest of the room has. Students ask to go there. That matters more than it sounds like it should.
When a confusion flag says something I cannot fix from the front of the room, that table is where the conversation happens — not as a remediation station, because it is the seat everyone wants. A student reading a paragraph out loud at that table, with one other person, is doing the same thing that would humiliate them standing at their desk. The check tells you who needs the conversation. The room decides whether the conversation costs them anything.
That is the part no meta-analysis measures and the part a teacher controls completely.
Where Does the Evidence Run Out?
At high school, more precisely than most people admit.
An IES review summarizing twenty years of adolescent literacy research screened 111 studies, found 33 meeting What Works Clearinghouse evidence standards, and identified 12 programs and practices with positive or potentially positive effects on reading comprehension, vocabulary or general literacy across grades 6–12. And then the line that belongs in every conversation about this: none of the 12 was conducted in a high school setting. The evidence base for adolescent literacy is, in practice, a middle school evidence base.
Three more honest limits from the meta-analytic work. Small studies inflate the result badly — studies with 250 or fewer participants produced an effect of +0.45 against +0.13 for large ones, which is the usual sign that tightly-supervised implementations outperform real ones. The Xuan authors found no significant difference between published and unpublished studies, so publication bias is not the explanation, but the sample-size gap is a warning about what happens when a practice scales. And only 8 of the 48 studies came from Confucian-heritage contexts against 40 Anglophone, so the authors caution against moving interventions across cultures unadapted.
None of that is a reason not to run a two-minute gist check. It is a reason not to build a district initiative on a number somebody rounded up.
What Should You Do Next?
Pick one check — the confusion flag if you want the highest yield for the least work — and run it twice a week for three weeks on a text you were teaching anyway. Sort, do not mark. Spend one whole class period in those three weeks teaching into whatever the biggest pile said, and notice whether the next set of flags moves.
If it does, add a second check. If it does not, the problem is probably that the flags are not changing your lesson yet, which is the failure mode the research is loudest about. Two related pages are worth having open: the site’s guide to checking for understanding for the general-purpose version of these moves, and formative assessment strategies for writing for the same problem on the production side. If your reading happens on screens, the medium changes some of this — digital reading has its own comprehension research and its own accessibility upside. And if you want students doing more of the judging themselves, which is what the strongest moderator in the meta-analysis points at, start with student self-assessment.
Before you go: grab the free Formative Assessment Quick-Use Pack (PDF) — ElevateTheNorm.com branded, printable, no email required.
Frequently Asked Questions
What is a formative assessment strategy for reading, exactly?
It is any short, ungraded check that shows you what a student understood from a text while you can still do something about it. A two-sentence gist statement, a confusion flag, a one-question prediction, a retell in the student’s own words. The test is not the format. The test is whether the information changes what you do in the next ten minutes. If you collect it and file it, you ran a quiz.
Should reading checks be graded?
No, and the reason is practical rather than philosophical. The moment a gist statement counts for points, students write what they think you want instead of what they actually understood, and the check stops telling you anything. Keep them in a separate column or in no column at all. If your school requires a reading grade, take it from something summative and say out loud that the daily checks do not feed it.
How do I do this without reading 120 responses every night?
Sort, do not mark. A stack of 28 two-sentence gist statements is a five-minute sort into three piles — got it, partial, missed the point — and the only thing you write is the number in each pile. That number is the instructional decision. Reading every response carefully is what makes teachers abandon this in week three, and it is not what produces the benefit.
Does the research actually support this for high school students?
Partly, and the gap is worth knowing before anyone quotes a number at you. The meta-analysis evidence for formative assessment and reading achievement includes middle and high school studies and finds them performing at least as well as elementary. But an IES review of twenty years of adolescent literacy research found that of twelve programs with positive or potentially positive effects on reading outcomes, none was conducted in a high school setting. The practices are defensible. The claim that they are proven in a high school is not.
Is the effect size big enough to be worth the class time?
It is modest and it is real. The most recent meta-analysis puts formative assessment’s effect on reading achievement at +0.19 across 48 studies and 116,051 students. An earlier meta-analysis found a weighted mean of 0.20 overall and 0.32 for English language arts specifically — and explicitly rejected the 0.40 to 0.70 range that gets quoted in professional development. Two or three minutes a day for a modest, reliable gain is a good trade. A forty-minute assessment system for the same gain is not.
What about the student who cannot read the text at all?
A comprehension check tells you that, which is most of its value. What it must not do is tell the whole room. None of the checks here require reading aloud in front of peers, and that is deliberate. A written gist, a confusion flag on a sticky note, or a quiet conversation at a side table all give you the same information without making a struggling reader perform. The WWC panel rates intensive individualized support for struggling adolescent readers as a strong recommendation, and a daily check is how you find out who needs it — not a substitute for it.
Does technology make these checks better?
The evidence says not much. The same meta-analysis that found a +0.19 effect for formative assessment in reading tested whether technology involvement moderated it and found essentially nothing — a coefficient of +0.001. A form on a laptop and a sticky note produce the same result. Use whichever one your students will actually complete in ninety seconds.
What actually predicted bigger effects, if not technology?
Two things, and both are about what happens after the check. Approaches that integrated teacher and student use of the information beat teacher-directed approaches alone by a significant margin. And studies where the information fed differentiated instruction showed an effect of +0.24 against +0.05 for those where it did not. Collecting the data is not the intervention. Changing the next lesson is.
Sources
- Xuan, Qingzhi, Alan C. K. Cheung, and Danping Sun. “The Effectiveness of Formative Assessment for Enhancing Reading Achievement in K-12 Classrooms: A Meta-Analysis.” Frontiers in Psychology, vol. 13, 2022. 48 studies, 116,051 students; pooled ES +0.19; middle/high school +0.27; differentiated instruction +0.24 vs +0.05; technology coefficient +0.001 (n.s.); small-sample studies +0.45 vs +0.13 for large. https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2022.990196/full
- Kingston, Neal, and Brooke Nash. “Formative Assessment: A Meta-Analysis and a Call for Research.” Educational Measurement: Issues and Practice, vol. 30, no. 4, 2011, pp. 28–37. 13 studies, 42 effect sizes; weighted mean 0.20, median 0.25; English language arts 0.32, mathematics 0.17, science 0.09. https://eric.ed.gov/?id=EJ951173
- Kamil, Michael L., et al. Improving Adolescent Literacy: Effective Classroom and Intervention Practices. IES Practice Guide, NCEE 2008-4027, What Works Clearinghouse, 2008. Five recommendations with evidence levels; explicit vocabulary instruction and explicit comprehension strategy instruction both rated strong. https://ies.ed.gov/ncee/wwc/docs/practiceguide/adlit_pg_082608.pdf
- Institute of Education Sciences, Regional Educational Laboratory Southeast. Summary of 20 Years of Research on the Effectiveness of Adolescent Literacy Programs and Practices. 111 studies screened, 33 met WWC evidence standards, 12 programs with positive or potentially positive effects — none conducted in a high school setting. https://ies.ed.gov/use-work/resource-library/report/systematic-literature-review/summary-20-years-research-effectiveness-adolescent-literacy-programs-and-practices
- Agarwal, Pooja K., Ludmila D. Nunes, and Janell R. Blunt. “Retrieval Practice Consistently Benefits Student Learning: A Systematic Review of Applied Research in Schools and Classrooms.” Educational Psychology Review, vol. 33, no. 4, 2021, pp. 1409–1453. 50 experiments, 5,374 participants; 57% of effect sizes medium or large. https://link.springer.com/article/10.1007/s10648-021-09595-9


