By Clay Shumate
Teacher professional development is any structured learning a teacher does to get better at the job — workshops, coaching, course work, collaborative planning, book studies, graduate hours. The research is clearer about what fails than what works. Sustained, content-focused, coached programs beat one-off sessions, but even the best-studied versions produce modest, uneven effects on students.
That is an uncomfortable thing to write on an education site. It would be easier to hand you seven features of effective PD and let you rearrange the faculty meeting calendar. But the honest version is more useful, because once you know where the evidence is thin you can stop blaming yourself for a session that changed nothing. It was probably never going to. What follows is what the studies found, where they disagree, and how to tell whether an hour is worth giving away.
Key Takeaways
- Short PD does not show measurable student effects. In the federal review that applied What Works Clearinghouse standards, the three studies with the least contact time — 5 to 14 hours total — produced no statistically significant effect on student achievement.
- Coaching is the best-evidenced delivery model, and it shrinks when scaled. A meta-analysis of 60 causal studies found large effects on instruction and smaller effects on achievement, with large programs producing only a fraction of the effects seen in small ones.
- The famous “features of effective PD” lists are correlational. The reviews that produced them generally studied programs that already worked, so they cannot tell you which ingredient did the work.
- PD can raise teacher knowledge and still not move students. A randomized federal study delivered 93 hours of math content PD, measurably improved teacher knowledge and explanations, and found no positive impact on student achievement.
- Spending is not the constraint. One large district study estimated roughly $18,000 per teacher per year on development, with no common thread distinguishing the teachers who improved — though that study has real methodological critics.
- You are allowed to evaluate the offer. Professional judgment includes judging professional learning. A teacher who can say why a session will not transfer is doing the job, not avoiding it.
What counts as teacher professional development?
Anything structured and intentional that is meant to improve a teacher’s practice counts, which is exactly why the category is so hard to study. A ninety-minute compliance briefing on a new attendance module and a two-year content-and-coaching partnership both get logged as professional development for educators. They share a line on a spreadsheet and almost nothing else.
The usual buckets are workshops and conferences, in-district training days, instructional coaching, professional learning communities and collaborative planning, lesson study, mentoring for new teachers, graduate course work, and self-directed reading or online modules. Staff development for teachers in most American districts leans heavily on the first two, because they are the cheapest way to reach every adult in a building on the same afternoon.
Keep that mess in mind whenever somebody says “PD does not work.” That claim is about as precise as “meetings do not work.” The question is which, under what conditions, and for how long.
Why does most professional development fail to change teaching?
Because the typical session ends at the moment the hard part begins. A presenter explains a practice, the room nods, everyone goes back to a classroom where the old routine is already running, and nobody ever returns to see whether the new thing was attempted. Explaining a practice and installing a practice are different jobs, and most PD is only funded for the first one.
There is a structural reason for this. Follow-up is expensive and invisible. A district can report that 400 teachers received training on formative assessment; it cannot easily report that 90 of them changed a lesson. So the metric collected is attendance, and the thing funded is whatever produces attendance.
TNTP’s 2015 report The Mirage put numbers on the scale of the investment. Surveying more than 10,000 teachers and 500 school leaders across three large districts and a mid-size charter network, it estimated roughly $18,000 per teacher per year on development and about 19 school days’ worth of time, and reported that evaluation ratings for close to 70 percent of teachers stayed flat or declined over two to three years, with no identifiable pattern separating the improvers (TNTP, 2015).
That report deserves its caveat, and it is a serious one. Heather Hill’s review for the National Education Policy Center argued that the study leaned on teacher evaluation ratings that are unreliable year to year, measured PD by format rather than by topic, fit or quality, and never established whether the development teachers received was even aimed at the things they scored poorly on. She called the sweeping conclusions “overly broad and not supported by the study’s methods” (Hill, NEPC, 2015). Both things can be true: the spending is enormous, and we do not have clean evidence about what it bought.

What does the research actually say about effective teacher professional development?
The most-cited synthesis identifies seven features, and its own authors say they cannot tell you which feature mattered. The Learning Policy Institute’s 2017 review of 35 methodologically rigorous studies found that effective programs tended to be content focused, to use active learning, to support collaboration, to model the practice being taught, to include coaching and expert support, to build in feedback and reflection, and to be sustained over time (Darling-Hammond, Hyler & Gardner, 2017).
Those seven features are the reason your district’s PD plan says “sustained, collaborative, job-embedded.” They are a reasonable starting point. But read the report’s own limitations section. Because the studies examined whole programs with many elements bundled together, the authors state plainly that they cannot draw conclusions about the efficacy of individual components, and that they were unable to comment on studies of PD that did not produce positive results.
That is not a technicality. It means the list describes programs that worked. It does not demonstrate that adding collaboration or duration to a weak program will make it work. Sims and Fletcher-Wood made this argument directly in a critical review in School Effectiveness and School Improvement, concluding that the influential feature lists rest on methodological weaknesses — chiefly inappropriate inclusion criteria — and that the field would do better to build on basic learning science than on those syntheses (Sims & Fletcher-Wood, 2021).
Mary Kennedy reached a related conclusion by a different route. Sorting 28 rigorous studies by the theory of action behind each program rather than by its design features, she found that many popular design features were not associated with program effectiveness, and that programs with similar content differed sharply in impact depending on how they supported teachers in enacting it (Kennedy, 2016).
How strong is the evidence for instructional coaching?
Coaching has the strongest causal evidence of any professional development model, and it still has a scale problem. Kraft, Blazar and Hogan reviewed 60 studies using causal designs and found pooled effects of 0.49 standard deviations on instructional practice and 0.18 standard deviations on student achievement (Kraft, Blazar & Hogan, 2018).
An effect of 0.49 on instruction is large for anything in education research. It is the one model where watching a teacher teach, talking about that lesson, and coming back next week reliably changes the lesson.
The caution is in the same paper. Average effects from effectiveness trials of larger programs were only a fraction of the effects found in efficacy trials of smaller programs. Small, carefully run, well-staffed coaching works. The district-wide rollout with one coach for 60 teachers is a different intervention wearing the same name. If your school is adding coaching, the ratio and the protected time are the program — not the job title.
Does more time automatically mean more learning?
No, but below a floor of about 14 contact hours the studies stop finding effects at all. The federal regional lab review by Yoon and colleagues screened more than 1,300 studies and found only nine that met What Works Clearinghouse evidence standards. Across those nine the average effect size was 0.54, and the three studies with the least professional development — 5 to 14 hours total — showed no statistically significant effects on student achievement (Yoon et al., REL 2007–033).
Two honest notes. Nine studies out of 1,300 is a thin base for a field this large, and the review is from 2007. And all nine were in elementary grades, so applying the 14-hour floor to a secondary department is an inference, not a finding. I still think it is the most useful number in the literature for a classroom teacher, because it reframes the question from “was the workshop good?” to “is there enough of this, over enough weeks, for anything to stick?”
And duration alone clearly is not sufficient, which brings us to the most sobering study in this whole area.

Can professional development raise teacher knowledge and still not move students?
Yes, and a randomized federal study showed exactly that. Garet and colleagues delivered an 80-hour summer math content workshop plus 13 hours of collaborative meetings and observation-based coaching to over 200 fourth-grade teachers across six districts in five states, randomly assigned. The PD improved teachers’ mathematical knowledge and the use and quality of their mathematical explanations in class. It did not have a positive impact on student achievement (Garet et al., NCEE 2016–4010).
That study was content focused, sustained, collaborative, modeled and coached — six of the seven features. The teachers learned. The students did not show it on the test within the study window.
Note the population: fourth-grade mathematics. This site is about grades 6–12, and I am not going to pretend a fourth-grade math result transfers cleanly to a tenth-grade history classroom. What does transfer is the logic of the chain. PD has to change teacher knowledge, then change instruction, then change what students do with their minds, then show up on a measure somebody chose. Every link can break, and the last two are the ones programs almost never plan for.
Why do the “features of effective PD” lists deserve skepticism?
Because a list of features is a description of winners, not a recipe. If you study only programs that raised achievement and then catalog what they had in common, you will find they were long, collaborative and content-focused — and you will have no idea how many long, collaborative, content-focused programs did nothing, because those were never in your sample.
It is why a district can faithfully implement every feature on the list — stretch the training across the year, put teachers in teams, anchor it in content — and get nothing. The features are present; the mechanism is absent. The Education Endowment Foundation’s 2021 systematic review of randomized trials, by Sims, Fletcher-Wood, O’Mara-Eves, Cottingham, Stansfield, Van Herwegen and Anders, reframed the field around mechanisms rather than surface features for exactly this reason (EEF, 2021). That review and its companion guidance are UK-based; the funding and staffing context is not ours, but the argument about mechanisms is.
The practical translation: stop asking whether a professional learning plan has the right shape and start asking what it is supposed to do to a teacher’s head and hands, and how anyone would know.
How should a teacher judge whether a professional development offer is worth their time?
Ask five questions before the first slide, and judge the answers against what the research found. You usually cannot decline required training, and I am not suggesting you try. But you can decide how much of your limited attention and your optional time to invest, and you can ask these out loud without being difficult about it.
| Question to ask | Answer that predicts transfer | Warning sign |
|---|---|---|
| How many hours total, across how many weeks? | Multiple sessions spread over a term, totaling well past the 14-hour floor | “It’s a one-off, but it’s a really good one” |
| Who is coming back to see me try it? | A named person, a scheduled visit, a protected debrief | Nobody, or “your evaluator will look for it” |
| Will I see the practice done with students like mine? | Modeled or filmed in a comparable grade band and subject | Modeled with compliant volunteers or with elementary students |
| What am I expected to stop doing to make room for this? | Something named and removed | Nothing — it is added on top |
| What would count as evidence this worked, and who checks? | A student-work artifact or a specific observable change | An exit survey about satisfaction |
If four of five answers land in the warning column, you are at an awareness session. Awareness sessions are not worthless — they can introduce a vocabulary and tell you something exists. Just do not expect your practice to move, and do not let anyone tell you later that you were trained.

What should schools and districts stop doing?
Stop buying reach when you can only afford depth, and stop counting attendance as implementation. A principal with a fixed budget faces a real trade-off: one day of training for ninety teachers, or a year of coaching for twelve. The coaching evidence says the second is more likely to change instruction. The optics say the first is more equitable. That tension is genuine and I do not think it resolves cleanly.
What does resolve: a few defensible moves that cost little.
- Protect the follow-up before you book the session. If the calendar cannot hold a return visit, the session is awareness-only — label it that way honestly.
- Narrow the number of initiatives per year. Three well-supported changes beat nine announced ones, and teachers can tell the difference immediately.
- Give choice where choice is possible. Required compliance training is required. Instructional learning does not have to be identical for a first-year teacher and a twenty-year veteran.
- Use the people already in the building. Peer observation and lesson study cost release time rather than consultant fees, and build the collective capacity a one-day visitor cannot.
- Report honestly. “Forty teachers trained” is an input. Say what changed in classrooms, or say you do not know yet.
For schools trying to build that kind of shared capacity, the research on what happens when a staff believes it can move students together is worth reading alongside this. It is also correlational, and I would hold it loosely for the same reasons.
What does professional learning look like when it is actually working?
It is specific, it is about your students, and somebody sees your actual teaching more than once. The common thread in the programs with real effects is not a format. It is proximity to the work.
In practice that looks like a small cycle: pick one narrow problem of practice, learn one technique that addresses it, try it in a named class on a named day, bring back a piece of student work or a short recording, and have one honest conversation about what happened. Then do it again. That cycle costs almost nothing in materials and a great deal in scheduling, which is why it is rare.
It also pairs with the kind of reflective habit a teacher can run alone. Structured reflection on your own practice is not a substitute for coaching — the evidence base is much weaker — but it is the part that does not require anyone else’s budget.
Where does PD fit with treating students as young adults?
How adults are trained tends to show up in how students are treated. This is the part where I am giving you my professional opinion rather than a research finding, so I will mark it as such.
Development that hands teachers a script and checks for compliance teaches a lesson about what learning is: do the steps, get the credit. Development that gives teachers a real problem, some room to solve it, and someone to think with teaches a different lesson. Teachers who are treated as capable adults with judgment are, in my experience and in my argument, more likely to extend the same thing to a sixteen-year-old. There is at least suggestive evidence in this direction — a study of 254 teachers of grades 1 through 12 found that teachers who perceived more pressure from above were less self-determined about teaching and in turn more controlling with students (Pelletier, Séguin-Lévesque & Legault, 2002). That is one correlational model from 2002, not proof, and I will not oversell it.
If you want the fuller version of that argument, I wrote it up separately in what it means to treat teaching as a profession. And if you want the narrower question of when a single session is worth running at all, that is the workshop format on its own terms.
What to do next with teacher professional development
Pick one thing, get enough hours behind it, and arrange for someone to come back. That is the whole recommendation, and it is supported about as well as anything in this literature is supported.
If you are a classroom teacher: choose one narrow problem this semester — openings, questioning, feedback turnaround, whatever is actually costing you — and spend your discretionary learning time there instead of spreading it across five interests. Ask a colleague to watch twice. If you are leading a department, protect the second visit before you book the first session. If you are a principal or a district leader, be willing to serve fewer teachers more deeply, and be honest in your reporting about what you did and did not measure.
And hold the evidence the way it deserves to be held. The strongest finding in teacher professional development is that short, unsupported sessions do not show up in student outcomes. The second strongest is that coaching helps and gets weaker as it scales. Everything past that is more contested than the slide decks suggest, and you are allowed to say so.
Sources
- Kraft, M. A., Blazar, D., & Hogan, D., “The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence,” Review of Educational Research, 2018. Source for the pooled coaching effects of 0.49 SD on instruction and 0.18 SD on achievement across 60 causal studies, and for the finding that large-scale programs show only a fraction of small-program effects. Caveat: a meta-analysis pools programs of widely differing design, and the included studies span grade levels and subjects. https://annenberg.brown.edu/publications/effect-teacher-coaching-instruction-and-achievement-meta-analysis-causal-evidence
- Darling-Hammond, L., Hyler, M. E., & Gardner, M., Effective Teacher Professional Development, Learning Policy Institute, June 2017. Source for the seven features of effective PD and for the review’s 35-study base. Caveat: the authors state they cannot draw conclusions about individual program components and could not comment on PD studies that did not yield positive results. https://learningpolicyinstitute.org/sites/default/files/product-files/Effective_Teacher_Professional_Development_REPORT.pdf
- Sims, S., & Fletcher-Wood, H., “Identifying the characteristics of effective teacher professional development: a critical review,” School Effectiveness and School Improvement, 32(1), 47–63, 2021. Source for the argument that influential feature lists rest on inappropriate inclusion criteria and flawed inference. Caveat: this is a critique of method, not evidence that sustained or collaborative PD fails. https://discovery.ucl.ac.uk/id/eprint/10101575/
- Kennedy, M. M., “How Does Professional Development Improve Teaching?” Review of Educational Research, 86(4), 945–980, 2016. Source for the finding that many popular design features are not associated with program effectiveness across 28 rigorous studies. Caveat: the classification by theory of action is the author’s own framework rather than a standard taxonomy. https://journals.sagepub.com/doi/abs/10.3102/0034654315626800
- Yoon, K. S., Duncan, T., Lee, S. W.-Y., Scarloss, B., & Shapley, K. L., Reviewing the evidence on how teacher professional development affects student achievement, REL Southwest, REL 2007–No. 033, October 2007. Source for the 14-hour threshold finding, the nine qualifying studies out of more than 1,300 screened, and the average effect size of 0.54. Caveat: all nine qualifying studies were in elementary grades, so the hour threshold is an inference for grades 6–12, not a finding. https://files.eric.ed.gov/fulltext/ED498548.pdf
- Garet, M. S., et al., Focusing on Mathematical Knowledge: The Impact of Content-Intensive Teacher Professional Development, NCEE 2016–4010, September 2016. Source for the randomized study of 80 hours of summer content PD plus 13 hours of meetings and coaching that improved teacher knowledge and explanations but produced no positive impact on student achievement. Caveat: the population was over 200 fourth-grade mathematics teachers, not secondary teachers. https://ies.ed.gov/use-work/resource-library/report/evaluation-report/focusing-mathematical-knowledge-impact-content-intensive-teacher-professional-development
- TNTP, The Mirage: Confronting the Hard Truth About Our Quest for Teacher Development (executive summary), 2015, together with Hill, H. C., Review of The Mirage, National Education Policy Center, September 2015. Source for the roughly $18,000 per teacher per year estimate and the flat-or-declining ratings finding, and for the methodological critique of both. Caveat: Hill argues the study’s reliance on unstable evaluation ratings and format-only PD measures makes its sweeping conclusions unsupported. https://www.nepc.colorado.edu/sites/default/files/ttr_hill_tntp_mirage.pdf
Frequently Asked Questions
Is all short professional development a waste of time?
No. A short session can introduce something, give a staff shared language, or deliver information that genuinely only needs to be delivered once. What the evidence does not support is expecting a short session to change classroom practice on its own. The federal review that applied What Works Clearinghouse standards found no statistically significant student effects in the studies with 5 to 14 total hours. So treat a one-off as awareness, not training, and judge it by whether you learned something true — not by whether your teaching changed.
If the evidence is this mixed, should districts cut professional development spending?
That is not the conclusion I would draw. The weak evidence is mostly about short, unsupported formats, and the strongest evidence — coaching — is the most expensive model. Cutting the budget would most likely cut the deep work first and leave the cheap compliance days standing. The more defensible move is reallocation: fewer initiatives, fewer teachers served at once, protected follow-up time, and honest reporting about what was measured. That is a harder political argument than a budget line, which is partly why it rarely happens.
Is coaching fair if only some teachers get it?
This is a real equity problem and it deserves a straight answer: no, not automatically. Coaching capacity is finite, so somebody decides who gets it, and those decisions can quietly track seniority, favor, or who is in trouble. If a school runs coaching, the selection rule should be written down and visible, the purpose should be stated as growth rather than remediation, and the rotation should eventually reach everyone. A coaching program nobody can explain the entry criteria for will be read as surveillance.
How does a teacher raise concerns about weak PD without sounding like a complainer?
Ask about logistics rather than value. “How many sessions is this across the semester?” and “Will anyone be coming by to see us try it?” are operational questions with clean answers, and they surface the transfer problem without putting the presenter on trial. If the answer is that there is no follow-up, you have not criticized anyone — you have learned what the session is for, and you can calibrate your effort honestly. Save the broader argument for whoever actually builds the calendar.
Does this apply to middle and high school the same way it applies to elementary?
Not identically, and the research base is honestly thinner for secondary. Several of the key studies — including all nine that met evidence standards in the 2007 federal review, and the randomized content-PD study with no achievement effect — were conducted in elementary grades. The structural logic still holds for grades 6 through 12: departmentalized teaching, 120 or more students per teacher, and a single planning period make follow-up harder to schedule, not easier. Treat the hour thresholds as orientation rather than as settled secondary findings.
What should a first-year teacher prioritize when everything is offered at once?
Classroom systems first, content pedagogy second, everything else later. A new teacher who stabilizes openings, transitions and response-to-misbehavior buys back the attention needed to learn anything else. Decline nothing that is required, but spend discretionary learning time narrowly, and find one colleague who will let you watch them and will watch you. That informal arrangement has more in common with the coaching evidence than most official programs do, and it costs nothing but nerve.
How can parents tell whether a school’s professional development is serious?
Ask what the school is working on this year and how many things are on the list. A school that can name two or three specific instructional priorities, say how teachers are being supported on them, and describe what they expect to see change is doing something real. A school that lists nine initiatives and reports how many staff attended training is reporting attendance. Neither answer tells you about any individual teacher, but the first indicates a leadership team that understands depth beats coverage.
Does teacher professional development affect how students are treated day to day?
There is suggestive evidence that it does, though I would call this the least settled claim in the article. A 2002 study of 254 teachers across grades 1 to 12 found that teachers who felt more pressure from above were less self-determined about their work and tended toward more controlling behavior with students. That is a correlational model, not proof. My own position is plainer: adults who are handed scripts and checked for compliance tend to pass that posture along, and adults trusted with judgment tend to extend it.


