{"id":9838,"date":"2026-09-21T06:41:04","date_gmt":"2026-09-21T13:41:04","guid":{"rendered":"https:\/\/shermandorn.com\/?p=9838"},"modified":"2026-09-21T06:50:49","modified_gmt":"2026-09-21T13:50:49","slug":"reasoning-about-and-interpreting-effect-sizes-how-big-are-they-really","status":"publish","type":"post","link":"https:\/\/shermandorn.com\/?p=9838","title":{"rendered":"Reasoning about and interpreting effect sizes: How big are they, really?"},"content":{"rendered":"\n<p>Half a century ago, educational psychologist Gene Glass coined a term for research synthesis, <em>meta-analysis<\/em>: \u201cthe term is a bit grand, but it is precise\u201d (Glass, 1976, p. 3). Glass and Mary Lee Smith had wondered how to combine different studies about the effectiveness of behavioral and talk therapy, and decided that one way to do so would be to scale the magnitude of the effect by the standard deviation of the study sample. This <em>effect size<\/em> borrowed from the existing methodological literature on calculating the needed sample size for a study (e.g., Cohen, 1962), and Glass and Smith used this abstract measure of difference to average the differences each study claimed talk therapy made (or didn\u2019t make) in the lives of therapy clients.\u00a0<\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><strong>The effect size as an abstract (and generalizable) measure of a study finding\u2019s magnitude<\/strong><\/h1>\n\n\n\n<p>In doing so, Glass and Smith went well beyond the prior practices of literature reviews that typically conducted narrative descriptions of study after study, comparing their designs and making a qualitative judgment about what one could learn from the research as a whole, or just counting the number of studies that found statistically significant results versus those that did not. When the original studies are quantitative, they reasoned, shouldn\u2019t the review\u2019s conclusions be as well, and go beyond narratives and simple vote-counting? The idea of meta-analysis broke well beyond the walls of psychology and now dominates the way some fields conduct literature reviews, including (and perhaps especially) medical research. Some quantitative fields such as economics are largely without meta-analysis, but it is the common expectation, especially in describing the effects of policies and practices. It has also developed well beyond its original, relatively simply process, as researchers have debated and formalized the steps of searching for relevant studies, screening them for characteristics that rule them in or out, extracting the statistical findings necessary to include in a meta-analysis, and conducting the meta-analytic statistical procedures.&nbsp;<\/p>\n\n\n\n<p>And the concept of an effect size, borrowed from research design study guidelines and embedded within meta-analysis, is now a common measure within individual studies, not just calculated after the fact by those conducting systematic literature reviews, but recommended in the last several editions of the <em>Publication Manual of the American Psychological Association<\/em>. The advantage of an effect size is its abstraction, and the ability to generalize beyond the context and sample characteristics of an individual study.&nbsp;<\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><strong>Can that magnitude be characterized in absolute terms?&nbsp;<\/strong><\/h1>\n\n\n\n<p>But there arose a challenge from that generality: how does one interpret an effect size, returning back from the abstract to the practical \u2013 or in the original motivation for Glass and Smith, the clinical? On this point the methodologist Jacob Cohen who coined the term <em>effect size<\/em> stepped back into the discussion, and in the 1980s and 1990s, Cohen asserted in textbooks and articles that there were abstract classifications of small, medium, and large effects, and that for each type of effect-size calculation, those were universal classes (e.g., Cohen, 1992). Even today, decades later, many statistical texts take Cohen\u2019s classification as the statistical gospel: there are small, medium, and large effects, and thou shalt not use any division but Cohen\u2019s.&nbsp;<\/p>\n\n\n\n<p>The problem is that this claim of universality is not backed up by empirical evidence. As Kraft (2020) observed with effect sizes in K-12 academic achievement, many effect sizes that Cohen would have classified as small turn out to be important in the lives of children, and are much more in the middle of the range in K-12 research. Kraft was explicit that his different categorization of effect sizes was based on research specific to the context of K-12 academic achievement in the United States.&nbsp;<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Contextualizing effect sizes<\/strong><\/h2>\n\n\n\n<p>Kraft further pointed out that the magnitude of an effect size is not the only relevant issue. Cost and implementation issues matter \u2013 an effect size that is the result of a relatively inexpensive and easily feasible practice is often a better choice than a policy or practice whose effect size that may be a little larger but is expensive either in direct cost or in the tradeoffs that a K-12 school district, college, private organization, or state might need to implement it.&nbsp;<\/p>\n\n\n\n<p>This leaves both researchers and consumers of research with a serious challenge: if there are no trustworthy absolute ways to grasp what an effect size means, and if one needs to weigh potential benefits against the costs, what does one <em>do<\/em> with this abstract thing called an effect size? In K-12 education in the past few decades, a practice has spread to interpret effect sizes in ways that practitioners and families might understand: time in school. If one can equate a certain effect size into the growth one might expect in a month or year of formal education, can that help?&nbsp;<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Awkward (and less awkward) translations<\/strong><\/h2>\n\n\n\n<p>I have seen practitioners hold onto this interpretation as the best guide in a miasma of abstractions. But it can also create misunderstandings of the magnitude of a potential benefit (or harm). In 2015, I wrote <a href=\"https:\/\/shermandorn.com\/?p=8079\">a blog entry<\/a> explaining that the \u201cmonths of learning\u201d translation often misstated the accuracy of an effect-size estimate, and more critically it removed the comparative question from the value of an effect size: the benefit <em>compared to what other options<\/em>? (I also made a fun but silly argument about unit translations and \u201cconverted\u201d effect sizes into measures of the career of Yankee pitching great Manny Rivera.)&nbsp;<\/p>\n\n\n\n<p>This informal piece of mine on the topic is almost certainly the only blog entry I\u2019ve ever written that\u2019s been cited in a peer-reviewed journal article: Baird and Pane (2019). The abstract summarizes Baird and Pane\u2019s takeaway: \u201cyears\/months\/weeks of learning\u201d is the worst choice they considered for interpreting effect sizes:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>We compare years of learning to three other translation options: benchmarking against other effect sizes, converting to percentile growth, and estimating the probability of scoring above a proficiency threshold. After enumerating the desirable properties of translations, we examine each option\u2019s strengths and weaknesses. We conclude that <em>years of learning performs worst<\/em>, and percentile gains performs best, making it our recommended choice for more interpretable translations of standardized effects. (emphasis added)<\/p>\n<\/blockquote>\n\n\n\n<p>The broader point I take away is that such back-translation efforts are inherently awkward, if well-intentioned.&nbsp;<\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><strong>Some suggested rules of thumb<\/strong><\/h1>\n\n\n\n<p>So what is one to do, between the need to abstract for generality and the need to interpret for specific contexts? In many cases one does not have the type of literature base Kraft (2020) had to establish benchmarks for a relatively narrow area (K-12 academic achievement).&nbsp;<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><em>Where a natural interpretation exists, use it<\/em>. One needs effect sizes in units such as standard deviation where the scale is not universal \u2013 scores on a specific achievement test, or composite scales that average Likert-type survey items. But if one is using a measure that does not need that abstraction \u2013 especially dichotomous variables such as high school graduation or continued enrollment in college \u2013 use those measures: a percentage-point difference in graduation is interpretable by itself, and generalizable by itself.&nbsp;<\/li>\n\n\n\n<li><em>Compare an effect size with intervention studies that are in the same context, or the nearest context you can find<\/em>. If your measure is pre-post differences in self-efficacy in teaching fractions to third-graders (using a survey with Likert-type items), then look for other studies of boosting confidence of elementary math teachers, even if not of teaching fractions or of third grade. If there are no such studies, you may need to move to projects that target middle school math teaching, or high school, and possibly expand to other areas of STEM. (One more reason to look for them to write your proposal lit review!)&nbsp;<\/li>\n\n\n\n<li><em>Use observational studies across time as a metric of what happens without interventions<\/em>. Maybe there are no studies of confidence among STEM teachers with interventions. Even if not, there may be studies that surveyed teachers across time about how they feel about their jobs \u2013 maybe confidence in their own teaching, maybe how satisfied they are, and if there are such surveys a few months or a year apart, that would give you an effect size about <em>ordinary progression<\/em> that may be as close as you can get to your own study.&nbsp;<\/li>\n\n\n\n<li><em>Make reasoned arguments about the nature of your problem of practice and the difficulty of change<\/em>. You know your situated context well, and you chose your problem of practice because it is important. Is this an area where change is relatively easy but no one has tackled your PoP? Or an area where change is typically hard, and other people care about the PoP but haven\u2019t been able to be successful yet? Those two very different stories suggest different expectations for an effect size.&nbsp;<\/li>\n<\/ol>\n\n\n\n<p>Note that these four ideas are sequenced in the order of greater distance from your data, and thus an increasing burden of making a persuasive argument. You may end up with a more qualitative argument about what type of effect size we should expect from <em>any<\/em> intervention in your problem of practice rather than a firm but unjustified claim that Cohen\u2019s <em>d<\/em> = 0.15 is \u201csmall\u201d or \u201cvery small.\u201d As Kraft (2020) would note, that statement would not be true at all in K-12 academic achievement studies!&nbsp;<\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><strong>References<\/strong><\/h1>\n\n\n\n<p>Baird, M. D., &amp; Pane, J. F. (2019). Translating standardized effects of education programs into more interpretable metrics. <em>Educational Researcher<\/em>, <em>48<\/em>(4), 217-228.<\/p>\n\n\n\n<p>Cohen, J. (1962). The statistical power of abnormal-social psychological research: A review. <em>Journal of Abnormal and Social Psychology, 65<\/em>(3), 145-153.&nbsp;<\/p>\n\n\n\n<p>Cohen, J. (1992). A power primer. <em>Psychological Bulletin<\/em>, <em>112<\/em>(1), 155.<\/p>\n\n\n\n<p>Glass, G. V. (1976). Primary, secondary, and meta-analysis of research. <em>Educational Researcher<\/em>, <em>5<\/em>(10), 3-8.<\/p>\n\n\n\n<p>Kraft, M. A. (2020). Interpreting effect sizes of education interventions. <em>Educational Researcher, 49<\/em>(4), 241\u2013253. <a href=\"https:\/\/doi.org\/10.3102\/0013189X20912798\">https:\/\/doi.org\/10.3102\/0013189X20912798<\/a>&nbsp;&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Half a century ago, educational psychologist Gene Glass coined a term for research synthesis, meta-analysis: \u201cthe term is a bit grand, but it is precise\u201d (Glass, 1976, p. 3). Glass and Mary Lee Smith had wondered how to combine different studies about the effectiveness of behavioral and talk therapy, and decided that one way to [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[4],"tags":[],"class_list":["post-9838","post","type-post","status-publish","format-standard","hentry","category-research"],"jetpack_publicize_connections":[],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/pag0MB-2yG","_links":{"self":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts\/9838","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=9838"}],"version-history":[{"count":6,"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts\/9838\/revisions"}],"predecessor-version":[{"id":9844,"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts\/9838\/revisions\/9844"}],"wp:attachment":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=9838"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=9838"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=9838"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}