By Clay Shumate
Formative assessment strategies for writing are the checks a teacher runs while a piece is still being written, so the information can still change the draft. Exit tickets and comprehension checks do not transfer here. Writing has its own set, because the thing being assessed takes days, gets revised, and is too long to read thirty times a week.
That last constraint is the real problem. Almost every writing-feedback system that gets abandoned was abandoned because it required reading every draft. What follows is what the research supports, the number that should worry you, and a set of checks that fit an ordinary secondary schedule.
Key Takeaways
- Feedback works, and the size depends on who gives it. A meta-analysis found adult feedback at an effect size of 0.87, self-evaluation at 0.62 and peer feedback at 0.58 — but that study covered grades 1–8.
- Two popular practices showed no meaningful effect in the same analysis. Teacher progress monitoring and the 6+1 Trait Writing model did not improve writing quality.
- The federal practice guide rates the assessment recommendation as its weakest. Of three recommendations for teaching secondary writing, “use assessments to inform instruction and feedback” carries minimal evidence. That is worth knowing before anyone sells you a system.
- Assess one thing at a time. A draft marked for everything gets revised for nothing. Pick the criterion the lesson taught and check only that.
- Never put a grade on a draft you want revised. Self-assessment that counts toward a mark stops being honest, and the same logic applies to a draft.
Free Download · PDF
Formative Assessment Quick-Use Pack
A strategy decision matrix, an evidence tracker, four exit-ticket formats, and a next-day response planner — four pages built around changing the next instructional move.
Free. No email address required. Designed for grades 6–12. Browse every printable in Your Free Library.
What Makes a Writing Check Formative?
Timing and use. A check is formative if it happens while the piece can still change and if somebody acts on it. Both conditions. A rubric applied to a final draft, however detailed, is a grade with extra steps.
That rules out more of the standard toolkit than teachers expect. Marking a finished essay is summative no matter how much you write in the margin. A reading quiz is not a writing check. A participation grade for peer editing measures compliance with a procedure, not the quality of anything. The general versions of these checks are covered in the list of formative assessment strategies; this page is about the ones built for a text that takes a week.
What is distinctive about writing is that the product has stages. A thesis exists before the paragraph does, an outline before the draft, a draft before the revision. Each stage is a place to check something cheaply, while it is still cheap to fix. The whole art is checking the earliest stage at which the problem is visible — because a thesis fixed on Monday saves five paragraphs that never had to be written.
What Does the Research Actually Show?
That feedback on writing works, that two popular practices do not, and that the evidence for the whole assessment-driven approach is weaker than its popularity suggests. All three are worth having straight.
The central study is Graham, Hebert and Harris’s 2015 meta-analysis in the Elementary School Journal, which pooled experimental studies of formative writing assessment and reported effect sizes by who does the assessing:
| Practice | Effect on writing quality |
|---|---|
| Adult feedback | 0.87 |
| Students evaluating their own writing | 0.62 |
| Peer feedback | 0.58 |
| Computer feedback | 0.38 |
| Teachers monitoring student progress | no meaningful improvement |
| The 6+1 Trait Writing model | no meaningful improvement |
Three things to notice. First, the grade band is 1 through 8. Only the top three grades of that range are secondary, and none of it is high school. The practices are reasonable to carry upward; the numbers do not travel with them.
Second, the two null results are the most useful rows in the table. Progress monitoring — tracking writing scores over time to inform instruction — showed no meaningful improvement, and neither did 6+1 Trait, which is a widely adopted framework. That does not make either worthless, but it does mean a department should not treat adopting them as having addressed writing feedback.
Third, student self-evaluation at 0.62 nearly matched adult feedback and beat peer feedback. For a teacher with 120 students, that is the most consequential number in the table, because it is the only one that does not scale with your reading time.
Now the part most articles omit. The What Works Clearinghouse practice guide on teaching secondary students to write effectively, written by a panel including Steve Graham, Jill Fitzgerald, Linda Friedrich, Katie Greene, James Kim and Carol Booth Olson, makes three recommendations and rates the evidence behind each. Explicitly teaching writing strategies through a model–practice–reflect cycle is rated strong. Integrating writing and reading using exemplar texts is rated moderate. Using assessments to inform instruction and feedback is rated minimal.
Minimal is the lowest rating in that system, and it is attached to the recommendation this article is about. Read it as a statement about proportion rather than a reason to stop. If you have a fixed amount of energy for improving writing in your classroom, the guide says to spend it first on explicitly teaching strategies and modeling them — and the assessment practices below are what you run inside that, not instead of it.

Which Checks Fit a Secondary Schedule?
The ones that read a sentence rather than a draft. Every strategy below is designed around the fact that you cannot read 120 full drafts and still have a weekend.
- The thesis-only check. Collect one sentence. Read all of them in ten minutes, sort into three piles — arguable, too broad, not a claim — and hand them back with the pile name. Fixing this on day one prevents most of what you would otherwise write in margins on day six.
- The one-criterion read. Announce that today’s read is only about evidence, or only about topic sentences. Mark only that. A draft marked for everything gets revised for nothing, because a student facing thirty marks does not know where to start and usually starts with the commas.
- The highlight-and-justify. Students highlight the sentence in their own draft that meets a specific criterion, and write one line explaining why. If they cannot find it, that is the check — and they have found it themselves, which is the part that matters.
- The first-paragraph conference. Two minutes per student, on the opening paragraph only, while the rest write. Twelve students a period, everybody covered across a week.
- The anonymous exemplar. Put two short pieces of writing on the board with names removed — one that does the thing, one that nearly does it — and have the class say which and why. This is the model–practice–reflect cycle the WWC rates as strong evidence, run as a five-minute check. Ask before you use a student’s work, even anonymized, because the writer always recognizes their own sentences and so do the people sitting next to them. Asking takes ten seconds, almost nobody says no, and the ones who do have a reason. Writing your own two versions, or using last year’s with permission, works just as well.
- The revision log. One line per revision: what changed, and why. It takes a student thirty seconds and tells you whether feedback was used, which is the only question a progress tracker was ever trying to answer.
Notice what is missing: reading every draft, writing extended comments, and any system requiring a spreadsheet. Those are the things that get abandoned in October, and abandonment is the real failure mode — not choosing the second-best check. If you would rather start from something printed, the free formative assessment quick-use pack has a general version to adapt.

How Do You Make Self-Evaluation Work?
Give the criteria first, keep it out of the gradebook, and ask for evidence rather than a rating. At 0.62 this is the best return per minute of teacher time in the whole table, and all three conditions matter.
Heidi Andrade’s 2019 critical review in Frontiers in Education supplies the design rules. Criterion-referenced self-assessment showed main effects on every criterion assessed, and concrete task-specific criteria outperformed vague competence-based ones — a result she attributes to Fastré and colleagues. In writing terms, “every claim is followed by a quotation and an explanation of it” is checkable. “Uses evidence effectively” is not.
The decisive rule is about grading. Andrade cites Tejeiro and colleagues, where self-assessment counted toward the final grade: overestimation rose dramatically and no correlation remained between the instructor’s assessment and the student’s. Run formatively, agreement improved substantially, and all twenty studies in her review that used self-assessment formatively showed a positive association with learning.
Andrade makes one more point worth carrying. There is little evidence that inaccurate self-assessment produces worse learning, and students act on their predictions regardless of accuracy. You are not trying to make a student’s judgment match yours. You are trying to make them read their own draft as a reader. The broader version of the practice is in the student self assessment guide; writing only changes what sits in the criteria.
Is Peer Feedback Worth the Class Time?
Yes at 0.58, and only if you narrow what you ask for. Unstructured peer review produces “I liked it, maybe add more detail,” which is the outcome most teachers have seen and correctly concluded is a waste of twenty minutes.
There is a useful finding from an adjacent literature. Falchikov and Goldfinch’s meta-analysis of 48 higher-education studies comparing peer marks with teacher marks found that agreement was closest when students made a global judgment against well-understood criteria, and worse when asked to break the judgment into many separate dimensions and score each. Those were undergraduates marking work, not teenagers giving revision advice, so treat it as a design hint rather than a transferred result — but the hint points somewhere useful: the twelve-box peer editing checklist is probably the worst available format.
What works better is narrower and more concrete:
- Ask for a location, not a judgment. “Underline the sentence where the argument actually starts.” A reader can do that honestly; “rate the organization” they cannot.
- Ask what the reader could not follow. This is the one thing a peer knows that you do not, because they read it without already knowing what the writer meant.
- Ask for one question, not one suggestion. Suggestions are advice from a novice. A genuine question — “is this the same person as in paragraph two?” — is information the writer can act on without deferring to anyone.
Two cautions. Peer feedback has a social cost that teacher feedback does not. A fifteen-year-old asked to critique a classmate’s writing is managing a relationship as well as a text, and most will resolve that tension by being vague. Asking for locations and questions rather than evaluations removes the tension rather than asking students to override it. And decide who reads what before you start. Writing is personal in a way a math worksheet is not; a student writing about something that matters to them should know in advance whether a classmate will see it, and should have a way out.
A One-Week Routine That Does Not Require Reading Every Draft
Four checks across a week, none of which takes you more than fifteen minutes.
- Day one — collect the thesis only. One sentence per student. Sort into three piles, hand back with the pile name, and give the “not a claim” pile five minutes to try again.
- Day two — the anonymous exemplar. Two openings on the board, names off, one working and one nearly working. The class names the difference. This is the modeling step, and it is the one with the strongest evidence behind it.
- Day three — highlight and justify. Students find the sentence in their own draft that meets today’s single criterion and write one line saying why. Walk the room and read over shoulders; collect nothing.
- Day four — peer question round. Swap drafts. Each reader underlines where the argument starts and writes one genuine question. Ten minutes, no checklist.
Then the revision log on day five: one line per change, what and why. That log is the whole assessment record, and unlike a score tracker it answers the question that actually matters, which is whether any of this changed the draft.
Two adjustments. First, differentiate the container rather than the criterion — a student who cannot produce the written justification quickly can say it to you in the last minute of class, and a student working in a second language can highlight and point. The judgment against criteria is what has to survive. Second, if a student’s draft has a problem that is not today’s criterion, note it for yourself and leave it. You will get to organization on the week you are checking organization, and a student who receives one correctable thing at a time actually corrects it. For the general version of this discipline, see checking for understanding.
Four Ways Writing Feedback Fails
All four are common, and all four are about the system rather than the student.
Everything gets marked. A draft returned with thirty corrections communicates that the piece is bad, not what to do next. Students respond by fixing the easiest marks, which are almost always the mechanical ones. One criterion per read is not a compromise — it is what makes revision possible.
A grade goes on the draft. Once a number is attached, the piece is finished in the student’s mind, and the comments underneath it become an explanation of the number rather than instructions for a revision. If you need the draft in the gradebook, grade completion rather than quality, and say which you are doing.
The one-criterion rule is also the thing families most often ask about, usually in the form of why the teacher did not correct all the errors. It deserves a straight answer rather than a defensive one: every error was noticed, and marking all of them is what produces a draft a student fixes the commas in and hands back otherwise unchanged. The errors get their turn, one at a time, on the assignment where that is the thing being taught. Said in advance, in a sentence on the assignment sheet, that lands as a deliberate method. Said after a parent email, it sounds like an excuse.
There is no time to act on it. Feedback returned the day the final is due is a post-mortem. If the schedule does not contain a revision block after the feedback, the feedback is decorative, and it is more honest to admit that than to keep writing comments into a void.

The department adopts a framework and calls it done. This is where the two null results earn their keep. Adopting 6+1 Trait or installing a progress-monitoring spreadsheet is a visible action that showed no meaningful effect on writing quality in the meta-analysis. The things that did work — someone reads a piece of writing and responds to it, or a student reads their own against criteria — are less visible on a plan, and are the ones worth protecting time for.
What to Try on the Next Assignment
Take the next piece of writing you have already assigned. Collect the thesis on its own, before anything else exists, and sort the sentences into three piles. That single move costs you ten minutes and changes more drafts than any set of margin comments you will write later.
Then pick one criterion for the whole assignment and check only that — in the exemplar, in the self-evaluation, in the peer round, in your own read. Put a revision block on the calendar before you give any feedback, and keep a number off the draft.
None of this is a system, and that is deliberate. The evidence for assessment-driven writing instruction is rated minimal by the people best placed to judge it, while explicit strategy instruction and modeling are rated strong. The checks here are worth running because they are cheap and they surface problems early. They are not a substitute for teaching students how to write, and any resource presenting them as one has the proportions backwards.
Before you go: grab the free Formative Assessment Quick-Use Pack (PDF) — ElevateTheNorm.com branded, printable, no email required.
Frequently Asked Questions
Do these effect sizes apply to high school students?
Not directly, and it matters. The meta-analysis behind the headline numbers — adult feedback 0.87, self-evaluation 0.62, peer feedback 0.58 — covered grades 1 through 8, so only its top three grades are secondary at all and none is high school. The practices are reasonable to carry upward because the mechanisms are not age-specific, but the numbers should not be quoted as a high school result. The federal practice guide that does cover secondary writing rates the assessment recommendation as having minimal evidence.
Should I put a grade on a draft?
Not if you want it revised. Once a number is attached, most students treat the piece as finished and read the comments as justification for the mark rather than instructions for a revision. The self-assessment research points the same way: when self-evaluation counted toward a grade, overestimation rose sharply and agreement with the instructor disappeared. If a draft has to appear in the gradebook, grade completion rather than quality and tell students that is what you are doing.
Is 6+1 Trait Writing a waste of time?
That is stronger than the evidence supports. What the meta-analysis found is that 6+1 Trait showed no meaningful improvement in writing quality, and the same was true of teachers monitoring student progress over time. That is a real finding and a department should not treat adopting the framework as having addressed writing feedback. It is not a finding that a shared vocabulary for talking about writing is harmful — it is a finding that the vocabulary alone does not move the writing.
How do I give feedback to 120 students without losing every weekend?
Stop reading whole drafts. Collect one sentence rather than one essay; mark one criterion rather than everything; run two-minute conferences on opening paragraphs while the rest write; and lean on self-evaluation, which came in at 0.62 and is the only practice in the table that does not scale with your reading time. A teacher who reads every draft thoroughly in September and nothing at all by November has given less useful feedback than one who reads one paragraph from everybody every week.
Does peer feedback actually help, or is it busywork?
It helped at an effect size of 0.58, which is real — but what most classrooms run is not what was studied. Unstructured peer review produces “I liked it, add more detail.” Narrow the ask instead: have readers underline where the argument starts, say what they could not follow, and write one genuine question rather than one suggestion. Those ask for information a peer actually has. Also settle who reads what before you begin, because writing is more personal than a worksheet and a student should know in advance.
What if a student’s draft has problems that are not this week’s criterion?
Note them for yourself and leave them alone. A student who receives one correctable thing at a time usually corrects it; a student who receives thirty fixes the commas. Keep your own running list and let it decide what the criterion is for the next assignment — that list is more useful as a planning document than as margin notes, and it means the pattern across the class shapes what you teach next rather than disappearing into thirty separate drafts.
How is this different from just marking essays carefully?
Timing and use. Marking a finished essay is summative however detailed it is, because nothing about that piece can change afterwards. A formative check happens while the writing is still in progress and is followed by time to act on it. The practical test is simple: if there is no revision block on the calendar after the feedback goes back, what you did was grading, and calling it formative assessment does not make it function like one.
Where should I spend my energy if I can only change one thing?
Not here, according to the people who reviewed the evidence. The What Works Clearinghouse panel rates explicitly teaching writing strategies through a model–practice–reflect cycle as strong evidence, integrating reading and writing with exemplar texts as moderate, and using assessments to inform instruction and feedback as minimal. If you have one change in you this year, make it the modeling. The checks on this page are cheap enough to run inside that work, and they are not a replacement for it.
Sources
- Graham, Steve, Michael Hebert, and Karen R. Harris. “Formative Assessment and Writing: A Meta-Analysis.” Elementary School Journal, vol. 115, no. 4, 2015, pp. 523–547. https://eric.ed.gov/?id=EJ1068976
- Graham, Steve, Jill Fitzgerald, Linda D. Friedrich, Katie Greene, James S. Kim, and Carol Booth Olson. “Teaching Secondary Students to Write Effectively” (practice guide summary). What Works Clearinghouse, Institute of Education Sciences, U.S. Department of Education. https://ies.ed.gov/ncee/wwc/Docs/PracticeGuide/wwc_secwrit_summary_053117.pdf
- Andrade, Heidi L. “A Critical Review of Research on Student Self-Assessment.” Frontiers in Education, vol. 4, art. 87, 2019. https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2019.00087/full (The Tejeiro et al. 2012 and Fastré et al. 2010 findings are reported in this review; the primary papers were not read directly.)
- Falchikov, Nancy, and Judy Goldfinch. “Student Peer Assessment in Higher Education: A Meta-Analysis Comparing Peer and Teacher Marks.” Review of Educational Research, vol. 70, no. 3, 2000, pp. 287–322. https://eric.ed.gov/?id=EJ630369


