{"id":8079,"date":"2015-09-24T15:08:16","date_gmt":"2015-09-24T22:08:16","guid":{"rendered":"http:\/\/shermandorn.com\/wordpress\/?p=8079"},"modified":"2015-09-24T19:39:04","modified_gmt":"2015-09-25T02:39:04","slug":"weeksdays-of-learning-is-well-intended-bad-interpretative-factoid","status":"publish","type":"post","link":"https:\/\/shermandorn.com\/?p=8079","title":{"rendered":"&#8220;Weeks\/days of learning&#8221; is well-intended bad interpretative factoid"},"content":{"rendered":"<p>The Institute of Education Sciences\u00a0has released a new\u00a0<a href=\"http:\/\/ies.ed.gov\/ncee\/pubs\/20154020\/index.asp\">Evaluation of the Teacher Incentive Fund, or TIF, (after two years)<\/a>, which is generally solid research by Mathematica Policy Research, at least at a quick first read today. The main findings:<\/p>\n<ul>\n<li>Most of the experimental part of TIF was implemented by the schools.<\/li>\n<li>Some parts of the program were more difficult to implement (e.g., higher performance pay for a more limited group of educators), or more difficult to maintain.<\/li>\n<li>Part of the logic model was hard to confirm, especially the issue of educator understanding of their opportunities to earn higher pay.<\/li>\n<li>The bottom-line effects on student performance were weak: 0.04 standard deviations in math, 0.03 in reading. If you obsess\u00a0about <em>p<\/em> values, only the association with reading was statistically significant.<\/li>\n<\/ul>\n<p>I say that this is generally solid research&#8230; until you get to the part of the document where the main effect size for reading is\u00a0translated into a statement that teacher and principal performance pay\u00a0is associated with three\u00a0additional &#8220;weeks of learning&#8221; in reading. Mathematica\u00a0is\u00a0using\u00a0a common, well-intended attempt to translate the abstract concept of effect size into something a general audience can understand. This translation has become more common in the last few years.<\/p>\n<p>It is also bad interpretation. That doesn&#8217;t mean that documents should not attempt the translation for a general audience, but there are problems with just using terms like &#8220;three weeks\u00a0of learning&#8221; as naked representations. To cut to the chase:<\/p>\n<ul>\n<li>&#8220;Weeks (or days) of learning&#8221; avoids the most important part of recontextualizing effect sizes: comparing the effect size in an individual study with effect sizes from empirical research in the same domain &#8212; i.e., if you are translating your research findings for use in the real world, how does\u00a0<em>this<\/em>\u00a0intervention or policy compare with\u00a0<em>other<\/em> interventions or policies that are realistic alternatives?<\/li>\n<li>&#8220;Weeks (or days) of learning&#8221; implies more accuracy than is realistic; it is hard to spot a difference between 3 and 4 weeks of learning\u00a0(and for those tempted to publish &#8220;days of learning,&#8221; under no circumstance in the real world can research make an empirically-justified distinctions between 15 and 16\u00a0days of learning). This study does not report standard errors for the estimates, but Mathematica does report the effect sizes under different models (or sensitivity to model assumptions), and the variations easily surpass 0.01 standard deviations, or the equivalent of one week of learning. If you want to talk about weeks\u00a0of learning for this study, we need to understand that depending on the model used the inferred effect on reading for the first cohort in the second year is likely to range somewhere between 0\u00a0and 4 weeks of learning. That interval may looks odd, but it&#8217;s a better representation of the research\u00a0than the statement in the report.<\/li>\n<\/ul>\n<p>Reporters reading such findings can\u00a0ask the authors two questions before writing stories, as a consequence:<\/p>\n<ul>\n<li>What are the effect sizes of potential alternatives, either in standard-deviation units or weeks\/days of learning?<\/li>\n<li>What is the error of the estimate &#8212; or the confidence intervals, in weeks or days of learning?<\/li>\n<\/ul>\n<p><!--more-->Background: the idea of translating effect sizes into &#8220;weeks\/days of learning&#8221; is an attempt to recontextualize research that is abstract. In the growth of systematic research reviews in the past 40 years, including meta-analyses, it has become more common to publish not only the raw estimates of effects in the context of an individual study, but to convert that estimate from the scale of an individual study into a more generalized unit, standard deviations. So we now float in a research environment of effect sizes: all well and good for comparing all sorts of things, a la John Hattie&#8217;s <a href=\"http:\/\/visible-learning.org\/\">Visible Learning<\/a> project to compare effect sizes of various education techniques and policies, but hard to explain to a general audience.<\/p>\n<p>Thus, the attempt to translate effect sizes\u00a0<em>back<\/em> into a concrete unit. In the case of this report, the\u00a0Mathematica report authors\u00a0use <a href=\"http:\/\/www.ncaase.com\/docs\/HillBloomBlackLipsey2007.pdf\">a 2008 study<\/a>\u00a0that provides one estimate of generalized annual gain in learning for specific grades, and then convert the effect size they find, 0.03\u00a0standard deviation units, as follows. ((Hill, C. J., Bloom, H. S., Black, A. R., &amp; Lipsey, M. W. (2008). Empirical benchmarks for interpreting effect sizes in research. <i>Child Development Perspectives<\/i>, <i>2<\/i>(3), 172-177.)) While it is not clear which benchmark they use, it looks to be about 0.40 standard deviations per year (the grade 4-to-5 gain in the referenced study):<\/p>\n<ul>\n<li>0.03 standard deviation units \/ 0.40 standard deviations per year =<\/li>\n<li>0.075 years of learning * 36 weeks\u00a0in a school year =<\/li>\n<li>2.7 weeks of learning, rounded to 3.<\/li>\n<\/ul>\n<p>Conversion between units is an interesting exercise, and my very first lab in high school physics was a snail race: we had to estimate the speed of each of our snails in furlongs per fortnight. ((It&#8217;s about 10^-4 furlongs per fortnight&#8211;snails are not fast.)) The conversion of school effect sizes into weeks or days of learning is a similar stretch in unit conversions. It is not nearly\u00a0as ridiculous as Randall Munroe&#8217;s <a href=\"https:\/\/xkcd.com\/687\/\">comic about abusing dimensional analysis<\/a>, but it is best understood as a counter-intuitive chain of reasoning, a chain of reasoning with some plausibility but not inherent logic.<\/p>\n<p>I have two primary concerns with the &#8220;weeks\/days of learning&#8221; translation. The first is that it misunderstands the needs of the readers. If you are a teacher, principal, or a policymaker you\u00a0<em>might<\/em> want to know a graspable figure for &#8220;how much&#8221; from the evidence but the relevant decision is different: Among my realistic choices in this domain, what is best? Translating a single effect size into weeks of learning is useless for that purpose, unless this report had\u00a0<em>also<\/em> translated effect sizes of competing options into weeks of learning.<\/p>\n<p>The second concern is with the implication of accuracy, more than the evidence can bear. If you follow the logic of this specific translation, a week of learning represents 0.011 standard deviations. To say that incentive pay\u00a0is associated with 3 weeks of additional learning reading rather than 2 or 4\u00a0weeks of learning, that means we need to be able to trust that the research really could make distinctions down to one hundredth of a standard deviation. In few\u00a0real-world education studies can one make that claim credibly, and as noted above, the sensitivity analyses reported in the appendices make clear that model assumptions create changes of more than 0.01 standard deviations &#8212; you would need to add standard errors in as well.<\/p>\n<p>Fortunately, these flaws are easily remediated\u00a0<em>if<\/em> either study authors or reporters specify\u00a0the relevant answers: 3 weeks of learning in comparison with (in this case unspecified) realistic alternatives, or a range of 0 to 4 weeks of learning (from the sensitivity analyses).<\/p>\n<hr align=\"center\" width=\"25%\" \/>\n<p>That ends my serious criticism of days and weeks of learning. And now we get to have some fun. Because I have two non-serious criticisms of the concept of weeks or days of learning.<\/p>\n<p>One is that we don&#8217;t know what\u00a0<em>kind<\/em> of school week\u00a0this study is describing. Is incentive pay\u00a0associated with three weeks\u00a0where everyone is focused, and things are happening on all cylinders? Or are we talking about the weeks\u00a0right before Christmas break, when no one is paying attention? Are these weeks with lots of subs? Or three weeks\u00a0with a series of schoolwide\u00a0assemblies where teachers never have time to get into a topic? Maybe &#8212; and I hate to bring this up, but you\u00a0<em>know<\/em> it&#8217;s a possibility &#8212; these are three weeks\u00a0when everyone is sick in turn, and on too many days, Johnnie came in despite being sick because his mother didn&#8217;t have any sick days at work, and he upchucked right beside Betsy&#8217;s desk at 9:30, the school splits the janitor with another school and didn&#8217;t have one that day, and it took about 60 minutes to find someone who could spare the time to find the key to the closet with mops, clean the mess up, find a fan to blow out the air into the hallway (and the fourth graders passing by the room on their way to the playground instantly made a HUGE complaint about what they smelled), and then find the class and let Ms. Deronde know she could bring the kids inside. With luck, no one else threw up in the meantime.\u00a0I don&#8217;t know about you, but if this study is saying that performance pay\u00a0is associated with three of THOSE weeks,\u00a0we don&#8217;t want any part of that.<\/p>\n<p>My other complaint &#8212; yes, there&#8217;s another one &#8212; is that we don&#8217;t know that weeks\u00a0of learning is the right measure. Should we stop the unit conversion with weeks\u00a0of learning? Let&#8217;s see what else we could do. ((For convenience, I have put all of the calculations and sources used in a <a href=\"https:\/\/docs.google.com\/spreadsheets\/d\/1Krpz6jh_iFMf4u5Tqq4zuGZs1VSM0hSUeCslUYblO7g\/edit#gid=2030041491\">public Google Sheet<\/a>.)) A week\u00a0is a unit of time, and we know that light travels at 186,282 miles per second, so a week\u00a0is equivalent to 1.9 billion\u00a0miles of learning. That&#8217;s not very useful to children, but I know that the average home run in the big leagues is 397 feet, so a week\u00a0of learning is also equivalent to 24.6\u00a0billion home runs of learning. Not bad! Let&#8217;s imagine that we&#8217;re asking what performance pay\u00a0means in New York City &#8212; I mean, if an intervention or policy can make it in New York, it can make it anywhere, right? Everyone in New York loves Manny Rivera, and he allowed only 71 home runs in his career. So a week\u00a0of learning is equivalent to 346\u00a0million Manny Riveras of learning. Or, for this study, 935 million Manny Riveras.<\/p>\n<p>Doesn&#8217;t every child deserve 935\u00a0million Manny Riveras of learning? If you don&#8217;t support performance pay, you are denying every child her or his\u00a0right to 935\u00a0million Manny Riveras of extra learning\u00a0in reading. ((Pedants may object on grammatical grounds. Perhaps it should be 935 million Mannys Rivera of learning.))<\/p>\n<p>But maybe we should get out of the Big Apple. Manny Rivera&#8217;s official stats give a weight of 195 pounds, so a week\u00a0is also equivalent to 67.5\u00a0billion pounds. But pounds of what? ((An alternative is to translate each Manny Rivera into his $90 million net worth. A week\u00a0of learning is thus worth 31\u00a0quadrillion dollars, or $84 quadrillion dollars <em>per child<\/em>\u00a0for the estimated effect of performance pay in reading. Puts Raj Chetty&#8217;s calculations to shame, if you ask me.))<\/p>\n<p>Bull. Pounds of bull. You can take Sherman out of the University of South Florida, but I <a href=\"http:\/\/www.gousfbulls.com\/\">had to see what the equivalent of 346\u00a0million Manny Riveras in bull would be<\/a>. At about a ton a bull, it&#8217;s 33.8 million bulls. Every week\u00a0of learning is the equivalent of just under 34\u00a0million bulls.<\/p>\n<p>And I know what you&#8217;re thinking: what are these bulls doing? ((Oh, I understand what those of you with one-track minds are thinking. Okay, for those of you who are wondering, an adult bovine produces approximately 65 pounds of manure per day. You can do the calculations from there.)) They&#8217;re running in Pamplona! It turns out that while there are 12 animals who participate each of eight days in the\u00a0annual Pamplona Encierros, half of them are steer. With 33.8\u00a0million bulls and 48\u00a0bulls per Running of the Bulls in the San Fermin Festival, each week\u00a0is worth approximately 700,000 San Fermin Festivals with their bull runs. Or, for the effect size in reading, that&#8217;s about two million San Fermin Festivals.<\/p>\n<p>And why would you\u00a0<em>ever<\/em> deny young children a chance to observe the Pamplona Running of the Bulls two million\u00a0times? It&#8217;s criminal to even think about! Oh, I know what some of you are thinking: Sherman, many\u00a0young children would never want to see the Running of the Bulls once, let alone two million\u00a0times. People are gored! Some have died!<\/p>\n<p>Nonsense. The Running of the Bulls is just like Youtube cat videos, if the cats could chase, maim, and kill people.<\/p>\n<hr align=\"center\" width=\"25%\" \/>\n<p>Okay, I lied. I am not ending with the ridiculous. Because <a href=\"http:\/\/blogs.edweek.org\/edweek\/teacherbeat\/2015\/09\/so_are_those_federal_performan.html\">Steve Sawchuk<\/a>\u00a0wrote a story on the study and quoted my &#8220;I am not impressed&#8221; tweet (to put it precisely, &#8220;Meh&#8221; &#8212; and Morgan Polikoff&#8217;s response suggesting a Likert-type scale for effect sizes), one final thought:<\/p>\n<p>Whether or not you think performance pay is a good or bad idea, this is hard work both in politics\u00a0and in practice. This is the type of hard work that is very common\u00a0with complicated policy ideas. The reason why it is important to compare effect sizes against alternatives in the same domain is because it is important to consider\u00a0the\u00a0return on the\u00a0<em>political<\/em> investment in policies. If you have the choice between several policies, all of which are hard to accomplish, you may want to consider where to invest time and political capital. Or, if there are a number of smaller-scale alternatives that are easier to pull off (simultaneously) and each of which have at least the magnitude of effect as the Almost Impossible, it would make a great deal of sense to push for\u00a0the several Difficult Lifts with Demonstrable Effects over the One Almost Impossible Lift with Minimal Effects.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The Institute of Education Sciences\u00a0has released a new\u00a0Evaluation of the Teacher Incentive Fund, or TIF, (after two years), which is generally solid research by Mathematica Policy Research, at least at a quick first read today. The main findings: Most of the experimental part of TIF was implemented by the schools. Some parts of the program [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":false,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[11,4,18],"tags":[],"class_list":["post-8079","post","type-post","status-publish","format-standard","hentry","category-education-policy","category-research","category-the-academic-life"],"jetpack_publicize_connections":[],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/pag0MB-26j","_links":{"self":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts\/8079","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=8079"}],"version-history":[{"count":41,"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts\/8079\/revisions"}],"predecessor-version":[{"id":8125,"href":"https:\/\/shermandorn.com\/index.php?rest_route=\/wp\/v2\/posts\/8079\/revisions\/8125"}],"wp:attachment":[{"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=8079"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=8079"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/shermandorn.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=8079"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}