
“Does everybody understand?”
Thirty heads nod. You move on. Two weeks later the unit test comes back and roughly a third of the room clearly did not understand, and hasn’t for a while.
The question wasn’t the problem. The problem is that “does everybody understand?” is not a check — it’s a request for permission to keep going, and students give it because saying no in front of thirty peers costs more than staying quiet.
This is the most fixable gap in a typical classroom, and the research behind fixing it is unusually strong.
What does “checking for understanding” actually mean?
It means gathering evidence of what students currently know, during the lesson, in a form that lets you change what you do next. Two conditions, both required: it happens while there’s still time to act, and it produces something you can actually read.
That distinguishes it from a quiz, which measures after the teaching is finished, and from a show of hands, which measures social confidence.
The formal name is formative assessment, and it has one of the better evidence bases in education.
What does the research say?
Paul Black and Dylan Wiliam reviewed more than 250 studies for their 1998 review Inside the Black Box. Their headline finding: studies of formative assessment showed effect sizes between 0.4 and 0.7, larger than most known educational interventions.
Two details from that work matter more than the headline number.
It helps struggling students most. Formative assessment had its strongest positive effect on low-achieving students, which narrows the achievement gap. That’s the opposite of most interventions, which tend to help the students already doing well.
The mechanism is specific. Black and Wiliam describe feedback as requiring three elements: recognition of the desired goal, evidence about present position, and some understanding of a way to close the gap between the two — and all three have to be understood by the learner before they can act. A checkmark supplies none of those. A grade supplies one, badly.
Where the claim gets oversold
Here’s the part most articles on this topic leave out, and it’s worth your skepticism.
The 0.4–0.7 range is from a 1998 review, and later work has been less generous. Kingston and Nash’s 2011 meta-analysis found a much more modest weighted mean effect size of 0.20. Hattie’s review of 12 meta-analyses on feedback landed at an average effect size of 0.73, but only under the right conditions.
So the honest summary is: formative assessment works, the size of the effect depends heavily on how it’s implemented, and anyone quoting 0.7 at you as a flat fact is quoting the top of a range from the most optimistic review.
My take, not a finding: the spread in those numbers is exactly what you’d expect from a practice that got adopted as a label. Once “formative assessment” became something districts could put on a walkthrough form, a lot of things got renamed formative that were really just quizzes with a nicer word attached. I’d bet the low effect sizes are measuring that, not measuring the practice Black and Wiliam described.
The three seconds that change the room
If you do one thing from this article, do this one.
Mary Budd Rowe recorded classrooms and measured how long teachers actually wait after asking a question. Her finding: teachers allow an average of about one second for a response, and react to a student’s answer within about nine-tenths of a second.
One second. That’s not enough time to parse the question, retrieve the information, and construct a sentence — for anyone.
When wait time is extended to three to five seconds, the length of student responses increases, unsolicited appropriate responses increase, student confidence increases, speculative responses increase, students compare data with each other more, student questions increase, and responses from students rated “relatively slow” increase — while the number of teacher questions that get no response at all decreases.
The teacher changes too. Questioning becomes more flexible and varied, and teacher expectations for students rated “slow” may shift.
That last one deserves a second read. Waiting three seconds changed what teachers believed their students were capable of. The students hadn’t changed. The measurement had.
There are two places to wait, and the second matters more:
- Wait time 1 — after you ask, before anyone answers.
- Wait time 2 — after a student finishes speaking, before you respond. Rowe found this one even more important than the first.
Three seconds feels absurd from the front of the room. It is not absurd from a desk.
What actually counts as evidence
The test for any check is simple: does it produce something from every student, that you can read, before the lesson ends?
Every student. Calling on the raised hand tells you about the raised hand. Cold call, whiteboards, or written responses tell you about the room.
That you can read. Nodding, “yeah,” and eye contact are not data. A sentence is data. A worked step is data. A ranking is data.
Before it ends. Information that arrives after the bell is a grade, not a check.
Practices that clear that bar without adding paperwork:
- Mini whiteboards, all hands up at once. Thirty answers in ten seconds, and you can see the distribution instantly.
- Cold call with wait time. Ask, wait three seconds, then name someone. Everybody has to think, because nobody knows who’s up.
- Two-sentence exit ticket. Not “did you get it” — a question only someone who got it can answer.
- “Explain it to someone who missed today.” Explaining exposes gaps that recognition hides.
- Show me where you got stuck. Students who can locate their own confusion are close to resolving it.
My take, not a finding: the mini whiteboard is the highest-return purchase I’ve made for a classroom. It removes the social cost of being wrong, because thirty boards go up at once and nobody’s answer is the event. Students who never volunteer will write things down. That’s not in the research I’ve cited — it’s what I’ve seen.
What doesn’t count
“Does everybody understand?” Covered above. It’s a request for permission.
“Any questions?” Answered honestly only by students who already know enough to know what they’re missing. The students furthest behind can’t formulate the question yet.
Thumbs up / thumbs down. Students look at each other’s thumbs. You’re measuring conformity.
The raised hand. Same three students, every period, in every school in the country.
For administrators: what this looks like in a walkthrough
Checking for understanding is genuinely visible from a doorway, unlike most instructional practice. What to look for:
- After the teacher asks a question, count. If the gap is under two seconds, wait time isn’t happening — and that’s a specific, coachable, non-threatening piece of feedback.
- Ask what they’d have done differently. A teacher doing this well can tell you what the last check revealed and how the lesson changed because of it. If the check produced no change, it was decoration.
- Look for who’s answering. If it’s the same four students, the room is being sampled, not measured.
- Don’t count the quiz. A quiz at the end measures. It doesn’t inform anything still in progress.
My take, not a finding: if I could hand every new teacher one number, it would be three seconds — not because it’s the most important practice in teaching, but because it’s the cheapest one with real evidence behind it. It costs nothing, requires no materials, needs no approval, and can be started in the middle of a lesson that’s already going badly.
Frequently asked questions
How often should I check for understanding in one lesson?
There’s no research-backed number, and anyone giving you one is guessing. The useful standard is before each transition to something that depends on the previous thing being understood. If the next step assumes the last step landed, check first.
Isn’t this just more assessment on top of everything else?
It replaces rather than adds, if you do it right. A check that takes ten seconds and changes your next five minutes saves you the reteach later. The version that becomes extra work is the version that gets recorded, graded, and filed — which is no longer formative.
Does this work in every subject?
Black and Wiliam reported the effect across kindergarten through college, in different subject areas and different countries. The format changes by subject; the principle doesn’t.
What do I do when the check shows most of them didn’t get it?
Reteach it differently, right then — not louder, not slower, differently. That moment is the entire return on the practice. A check you don’t act on is just a slower way of finding out later. And when students know you check because you want to help them learn, it builds the kind of classroom responsibility that makes every other practice work better.
References
Black, P., & Wiliam, D. (1998). Inside the black box: Raising standards through classroom assessment. Phi Delta Kappan, 80(2), 139–148.
Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education, 5(1), 7–74.
Hattie, J. (2009). Visible learning: A synthesis of over 800 meta-analyses relating to achievement. Routledge.
Kingston, N., & Nash, B. (2011). Formative assessment: A meta-analysis and a call for research. Educational Measurement: Issues and Practice, 30(4), 28–37.
Rowe, M. B. (1972). Wait-time and rewards as instructional variables: Their influence on language, logic, and fate control. ERIC ED061103. https://eric.ed.gov/?id=ED061103
Rowe, M. B. (1986). Wait time: Slowing down may be a way of speeding up. Journal of Teacher Education, 37(1), 43–50.