Formative Assessment Strategies in Science That Catch a Misconception Before the Unit Test

Title card reading Formative Assessment Strategies in Science with the subtitle nine checks that catch a wrong model before the lab runs on it, and a grades 6 to 12 panel naming probe, predict, and claim evidence reasoning

By Clay Shumate

Formative assessment strategies in science are the short, ungraded checks you run mid-lesson to find out what model a student is actually using — not just whether the answer is right. In science that distinction matters more than in most subjects, because a student can reach the correct result from a wrong theory and nobody finds out until a transfer question in March.

That is the case for doing this. The case against overselling it comes next, and it is stronger than most articles on this topic will tell you.

Key Takeaways

  • The famous effect size is not real. The often-quoted 0.40–0.70 for formative assessment was not supported when somebody went and checked. The best meta-analysis found a weighted mean of 0.20 — and 0.09 in science, the lowest of the three subjects.
  • That 0.09 is contested too. A published commentary by five assessment researchers challenged the study selection and the effect-size calculations. Do not use the small number as a reason to skip formative assessment any more than you should have used the big one as proof.
  • Chase the model, not the answer. The checks that pay in science are the ones that make a student’s underlying theory visible — a written prediction, a forced choice plus an explanation, a piece of reasoning.
  • Reasoning is the part they cannot do. In the claim-evidence-reasoning framework, the claim is the easiest piece for students. A correct claim is close to no information on its own.
  • If it takes an evening to read, it will not survive October. Every strategy below is designed to be readable in the time between classes.

What Are Formative Assessment Strategies in Science?

They are checks run during instruction, for the purpose of changing instruction, and they are not graded. That last clause is the one people drop. A quiz you record is a small summative assessment, and the moment students know it counts they start writing what they think you want instead of what they believe — which destroys the only thing you were collecting it for.

The Developing Assessments for the Next Generation Science Standards report from the National Research Council puts classroom assessment at the centre of instruction and separates its two jobs cleanly: formative assessment guides instructional decisions and lesson planning, summative assessment assigns grades. The same report sets the bar that makes science assessment hard. Three-dimensional science learning asks students to use science practices and crosscutting concepts in the context of disciplinary core ideas, and assessment tasks therefore need multiple components reflecting that connected use. The committee is blunt that tasks like these are challenging to design, implement and interpret, and that teachers will need real professional development to do it.

This is the general version of the argument we make in the main list of formative assessment strategies. What follows is the part that is specific to a science room.

Does Formative Assessment Actually Work Better in Science?

On the measured evidence, science is the weakest of the three subjects studied — and the measurement itself is disputed. Anyone who tells you formative assessment is a 0.7-effect intervention in your biology class is quoting a number that was checked and did not hold up.

Four-row graphic on formative assessment evidence: the 0.40 to 0.70 claim is unsupported, Kingston and Nash found a weighted mean of 0.20 from 13 usable studies, subject effect sizes of 0.32 for English language arts, 0.17 for mathematics and 0.09 for science, and a published critique of that estimate
Science scores lowest of the three subjects, on a very thin evidence base.

Here is what happened. Neal Kingston and Brooke Nash reviewed more than 300 studies that appeared to address formative assessment in grades K–12. Their abstract states plainly that the commonly claimed effect of about 0.70, or 0.40 to 0.70, “is not supported by the existing research base.” Many studies had severely flawed designs yielding uninterpretable results. Only 13 provided enough information to calculate an effect size at all, producing 42 independent estimates. The median observed effect was 0.25; the weighted mean under a random-effects model was 0.20. Their moderator analysis estimated 0.32 for English language arts, 0.17 for mathematics and 0.09 for science.

Then the argument got better rather than worse. Derek Briggs, Maria Araceli Ruiz-Primo, Erin Furtak, Lorrie Shepard and Yue Yin published a commentary in the same journal challenging the analysis on four grounds: that the keyword search was not validated and could not be fully replicated, that the inclusion criteria were applied inconsistently — they name two studies that met the criteria and were excluded and one that was wrongly included — that effect sizes computed without pretest data diverge substantially from adjusted ones, and that the analysis never examined whether outcome measures close to the taught curriculum inflated results differently across subjects.

So the responsible reading is not “formative assessment barely works in science.” It is that 13 usable studies out of 300 is not an evidence base, and the subject-level split sits on a fraction of those 13. We treat contested evidence as contested on this site, and this is as contested as school research gets.

Which leaves you with a practical standard instead of a statistical one: use a check because it makes student thinking visible to you in time to act on it. That is a claim you can verify in your own room by Friday, and it does not depend on anyone’s meta-analysis.

Why Science Is Different From Reading and Writing

Because the wrong ideas students bring in are coherent, useful to them, and survive being taught the right answer. A student who thinks heavier objects fall faster is not missing information; they have a working theory built from a decade of watching things fall, and it predicts most of what they have seen. Telling them the correct rule does not remove the theory. It adds a second one they use on tests.

This is why the science version of “check for understanding” has to go after the model. A thumbs-up, a right answer on a plug-in-the-formula problem, or a confident nod all leave the old theory completely intact. Our broader piece on checking for understanding makes the case against compliance signals generally; in science the gap between the signal and the understanding is just wider.

The second difference is the three-dimensional structure described above. A check that only asks students to recall a core idea is assessing one third of what the standards ask for and none of the practice. The strategies below are chosen because most of them put a practice — arguing from evidence, analysing data, constructing an explanation — into a five-minute container.

Nine Formative Assessment Strategies That Fit Inside One Science Period

Every one of these is readable in a planning period and none of them needs to be graded. Pick two and run them until they are automatic rather than trying all nine in a week.

Numbered graphic listing nine formative assessment strategies for science: misconception probe, prediction before data, claim evidence reasoning exit slip, card sort, graph read, one-criterion notebook check, confidence rating, error analysis, and student-written question
Nine checks, each short enough to run inside one period.

1. The misconception probe. A short forced choice among several plausible student ideas, followed by “explain your thinking.” NSTA has published these for years as the Uncovering Student Ideas series by Page Keeley, and the teacher notes that come with them summarise the research behind each misconception. You can write your own in ten minutes once you know the common wrong answers in your unit. The choices are the hook; the explanations are the data.

2. Prediction before data. Before the lab or the demo runs, every student writes down what will happen and why. Thirty seconds. The “why” is the whole point — a wrong prediction with sound reasoning is a different student than a right prediction with no reasoning, and only one of them needs reteaching. It also stops the quiet rewrite where students adjust their memory of what they expected once they see the result.

3. The claim-evidence-reasoning exit slip. One claim, one piece of evidence from today’s data, one sentence of reasoning. Ten minutes at the end of class gives you a whole-class picture of who can connect data to a conclusion. More on what to look for below.

4. Card sort. Twelve items to sort into examples and non-examples of a concept — physical and chemical change, or inherited and acquired traits. Sorting takes four minutes. The assessment is the argument over the three hard cards, and you should be circulating with a notebook while it happens.

5. The graph read. Hand them an unfamiliar graph from the same system and ask one question: what does the slope mean here? Students who have learned a procedure answer the shape; students who understand the science answer the quantity. This is the fastest way to separate the two.

6. One-criterion notebook check. Walk the room and check every notebook against exactly one thing — units, or the controlled variable, or whether the axis is labelled. A full notebook review does not scale and so it does not happen. A single-criterion pass takes six minutes and actually happens.

7. Confidence rating. Answer the question, then rate how sure you are. The quadrant that matters is confident-and-wrong, because those students will not ask for help and will not revise. Unsure-and-right is a different problem and needs reassurance rather than reteaching.

8. Error analysis. Give them a flawed explanation and ask them to find the fault. Catching a bad inference is cognitively harder than producing a tidy one, and it tells you whether a student can evaluate reasoning or only generate it. Write the flawed version yourself from the errors you saw last year.

9. The student-written question. Ask each student to write a test question on today’s objective. What they think is worth testing tells you exactly what they think the lesson was about, which is frequently not what you thought it was about.

What If You Teach 150 Students?

Then you read a sample, not a set. Nobody reads 150 exit slips in a planning period, and pretending otherwise is how a good strategy dies in week three.

Take one class’s worth — thirty slips — and sort them into three piles as you go: has it, partly has it, does not have it. Do not write on them. Count the piles, read four or five from the middle pile properly, and that is your instructional decision made in eight minutes. If the sampled class is wildly different from what you saw while circulating in the others, sample a second class before you change anything.

The sample is for the decision. The individual feedback is a separate job and does not have to happen on every check — which is the trade most teachers get backwards, giving thin comments to everyone instead of a real decision to themselves.

Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.

How Do You Check Reasoning and Not Just the Answer?

Use claim, evidence and reasoning as three separate things to look at, and spend your attention on the third one. Katherine McNeill and Joseph Krajcik, writing for NSTA Press on assessing middle school students’ written scientific explanations, are direct about the asymmetry: the claim is the easiest component for students to construct.

Three-card graphic breaking down the claim, evidence and reasoning framework and what a teacher should check in each part, with a fourth row on giving specific feedback
Reasoning is the part that tells you whether they understand the science.

Their framework splits the work. The claim is a statement that answers the question. The evidence is scientific data supporting it, and they give you two criteria for judging it: appropriateness, meaning the data is actually relevant to the problem, and sufficiency, meaning students use more than one piece rather than resting on a single data point. The reasoning is the justification for why that evidence supports that claim, and it usually requires applying a scientific principle — which is also what tells students what counts as evidence in the first place.

They report that students often struggle to justify their claims appropriately, which is the finding with the most practical consequence here. If you mark a CER response by checking whether the claim is right, you will grade a lot of students as proficient who cannot explain why anything follows from anything.

One caution the framework does not raise and a science teacher has to. A written explanation is a literacy task sitting inside a science check, and a student who understands the chemistry but cannot yet produce a paragraph in English will score as not understanding the chemistry. That is a measurement error, not a finding. Sentence stems fix most of it — “My claim is… The evidence that supports this is… This matters because…” — and for students on language or writing accommodations, take the reasoning verbally or as a labelled diagram. You are assessing whether they can justify a claim, not whether they can type it.

On feedback, their guidance is specific rather than encouraging: identify the particular strength, identify the particular weakness, suggest an improvement, and ask a question that pushes on the thinking. “Good evidence — explain more” is not feedback. “Your evidence is relevant but it is one trial; what would a second trial have to show for your claim to hold?” is.

Where Science Formative Assessment Goes Wrong

Four failure modes, and three of them are about what happens after the data arrives.

You collect and do not act. This is the big one. A probe you read on Saturday and never mention again is a worksheet. If you are not willing to change Monday, do not run it on Friday.

You reteach by repeating. If a dozen explanations contain the same wrong idea, you have found a shared model, and volume does not displace a model. Find the case their theory cannot explain and build ten minutes around that instead.

You grade it. Covered above and worth repeating, because it is the most common way a good probe gets ruined. Record completion if you must record something.

Say that part out loud to the room, and be ready to say it to a parent. Teenagers are not stupid about ungraded work; left unexplained, “this does not count” reads as “this does not matter,” and the effort drops accordingly. The version that works is the honest one: this is not graded because I need to know what you actually think, and if it counted you would write what you think I want. A parent asking why work in the gradebook is marked complete rather than scored gets the same sentence. It is a better answer than most grading policies can give.

The other half of the dignity question is what you do with a wrong answer once you have it. Patterns go on the board; names do not. The whole arrangement depends on students believing that telling you what they really think is safe, and one wrong explanation read aloud with an author attached ends that for the year.

You run too many. Nine strategies in one article is a menu, not a schedule. Two strategies you run every week beat nine you run once, both for your sanity and because students get fluent in a format and stop spending their effort decoding the task. That is the same argument as the one for stable classroom routines, applied to assessment.

What to Do Next

Pick two. One that runs before instruction — the misconception probe or the written prediction — and one that runs after, which for most science rooms should be the CER exit slip. Run them for three weeks before you judge them, because the first week measures how well students understand the format rather than the content.

Then write down what you changed. If after three weeks you cannot name a lesson you taught differently because of what a check told you, the problem is not the strategy. It is that the data is arriving somewhere that does not feed a decision, and that is fixable in a way a 0.09 effect size is not.

One note for anyone reading this with a department rather than a classroom in mind. Eight of these nine checks run on paper and none of them needs a device, a subscription or a consumable, which matters in a building where the lab budget and the technology cart are not evenly distributed. What they do need is the thing the National Research Council flagged: the skill to design and interpret them. That is a department-level job. A shared bank of probes for the three units everyone teaches, built once and reused, is worth considerably more than nine teachers each inventing their own in September.

The sibling guides for other subjects are built the same way: formative assessment in reading and in writing, and the printable strategy guide collects the cross-subject versions in one free download.

Frequently Asked Questions

What are formative assessment strategies in science?

They are the short, low-stakes checks a science teacher runs during a lesson or unit to find out what students actually think before the summative assessment does it for them. In science specifically they have to surface the student’s model of how something works — not just whether they got the answer — because a student can produce a correct result from a wrong mental model and nobody notices until the transfer question.

Is formative assessment proven to work in science?

Less than you have been told. Kingston and Nash’s meta-analysis found a weighted mean effect size of 0.20 across subjects and an estimated 0.09 in science — the lowest of the three subject areas they examined — from only 13 studies that were methodologically usable out of more than 300 reviewed. A published commentary then challenged that estimate on several grounds. The honest summary is that the research base is too thin to give you a number in either direction.

How is science formative assessment different from reading or writing?

Two things. First, science misconceptions are durable — students arrive with working theories about force, heat, inheritance and seasons that instruction often fails to displace, so a check has to go after the model and not the vocabulary. Second, the Next Generation Science Standards ask for three dimensions at once: a disciplinary core idea, a science practice and a crosscutting concept. A check that only tests recall is not assessing what you are teaching.

How long should a formative check take?

Five to ten minutes, and the reading of it should take less than that. If a strategy costs you twenty minutes of class and an evening of marking, it will not survive October, which means it is not a strategy — it is a one-off. The one-criterion notebook check exists precisely because a full notebook review does not scale and a single-criterion pass does.

Do I have to grade formative assessments?

No, and grading them usually ruins them. The moment a probe counts, students answer what they think you want rather than what they believe, and the data you wanted disappears. Record completion if your gradebook needs something. Keep the content ungraded.

What is a misconception probe?

A short question — often a forced choice between several plausible student ideas — followed by “explain your thinking.” NSTA publishes a long-running series of them under the title Uncovering Student Ideas, written by Page Keeley, with teacher notes that summarise the research on each misconception. The choice is only the hook. The explanation is the assessment.

What do I do when half the class gets it wrong?

Reteach, but not by repeating. If the same wrong idea shows up in a dozen explanations you have found a shared model, and saying the correct version louder does not displace a model — confronting it with a case it cannot explain does. Pick the piece of evidence their idea gets wrong and build the next ten minutes around that.

Can I use these strategies without lab equipment?

Most of them, yes. The misconception probe, card sort, graph read, error analysis and student-written question all run on paper. Prediction before data needs a demonstration rather than a full lab, and a single front-of-room demo works. The claim-evidence-reasoning exit slip needs data, which can be a data set you hand out rather than one the class collected.

Sources

  1. Kingston, Neal, and Brooke Nash. “Formative Assessment: A Meta-Analysis and a Call for Research.” Educational Measurement: Issues and Practice, vol. 30, no. 4, 2011, pp. 28–37. More than 300 K–12 studies reviewed; only 13 allowed effect-size calculation, yielding 42 independent effect sizes. Median 0.25; random-effects weighted mean 0.20. Moderator estimates: ELA 0.32, mathematics 0.17, science 0.09. Cited from the published abstract via the ERIC record; the full article sits behind a publisher paywall and was not read in full. https://eric.ed.gov/?id=EJ951173
  2. Briggs, Derek C., Maria Araceli Ruiz-Primo, Erin Furtak, Lorrie Shepard, and Yue Yin. “Meta-Analytic Methodology and Inferences About the Efficacy of Formative Assessment.” Educational Measurement: Issues and Practice, 2012. Commentary challenging Kingston and Nash on four grounds: unvalidated and non-replicable keyword search; inconsistently applied inclusion criteria, naming Andrade et al. (2008) and Bonner (2009) as wrongly excluded and one study as wrongly included; divergence between effect sizes computed with and without pretest data; and unexamined effects of curriculum-proximal outcome measures. A commentary, not an independent meta-analysis — it argues the estimate is unreliable, not that a different number is correct. https://www.colorado.edu/education/sites/default/files/attached-files/EMIP_Commentary_FINAL_082012.pdf
  3. National Research Council. Developing Assessments for the Next Generation Science Standards. National Academies Press, 2014. Places classroom assessment at the centre of instruction; distinguishes formative assessment (guiding instructional decisions) from summative (assigning grades); states that three-dimensional assessment tasks require multiple components reflecting the connected use of science practices, crosscutting concepts and disciplinary core ideas, and that such tasks are challenging to design, implement and interpret. https://nap.nationalacademies.org/read/18409/chapter/2
  4. McNeill, Katherine L., and Joseph S. Krajcik. “Assessing middle school students’ content knowledge and reasoning through written scientific explanations.” National Science Teachers Association Press. Defines the claim-evidence-reasoning framework; identifies the claim as the easiest component for students to construct, sets appropriateness and sufficiency as the two criteria for evidence, and describes reasoning as the justification that applies a scientific principle. Reports that students often have difficulty appropriately justifying their claims, and recommends explicit, specific feedback over general comments. A book chapter written for practitioners, not a peer-reviewed empirical report; the version read is the authors’ accepted manuscript and carries no publication year on its face. https://websites.umich.edu/~krajcik/McNeill&Krajcik_NSTA.pdf
  5. “Uncovering Student Ideas Probes.” National Science Teaching Association. Describes the probe series by Page Keeley as questions designed to reveal what all students are thinking and to uncover initial ideas and misconceptions about core concepts, with teacher notes containing research summaries and instructional suggestions. A publisher’s description of a commercial series, cited for what a probe is rather than as evidence of effect. https://www.nsta.org/uncovering-student-ideas-probes

About Clay Shumate

Clay Shumate is a certified secondary Social Studies teacher in the public schools of West Alabama, with seven years of classroom experience, a B.A. in History, and an M.Ed. in Secondary Education. He writes about project-based learning, student responsibility, respect, and practical ways to hold young people to a higher standard while giving them room to learn from mistakes. He is a member of the Society of Professional Journalists and writes to its Code of Ethics; this site’s editorial standards and corrections policy are published in full. More about Clay.

Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.Teacher Emergency Toolkit — practical resources, real classroom support. Shop on TPT.

BOO! DON’T BE SCARED!

There aren’t any tricks here, only treats!
Subscribe to claim your exclusive Halloween ebook, free and only for subscribers.

We don’t spam! Read our privacy policy for more info.