Warning: Trying to access array offset on true in /home/shermand/www/www/wp-content/themes/hybrid/library/functions/styles.php on line 77

Dan Weisberg presentation at Hillsborough Community College, October 6

Yesterday, The New Teacher Project Vice President Dan Weisberg spoke at one of the Florida Department of Education What's Working forums, and because it was in Tampa I was able to attend. He spoke the same day that TNTP released its 12-page booklet suggesting six guidelines for teacher evaluation. And, not that this should have been a surprise, Weisberg's talk was about teacher evaluation. In the audience for most of the two hours were a number of legislators (Rep. Kelli Stargel, Rep. Bill Heller, Sen. Ronda Storms) or aides (I think an aide to Sen. Nancy Detert was in the audience) as well as teachers, parents, higher-ed faculty, and several union staff members. In his 30-minute talk, Weisberg provided a summary of the argument that appeared in The Widget Effect and a description of some places, including DC and New Haven, that are moving in what and other TNTP officials think is a more productive direction.

Early in his talk, Weisberg repeated the oversimplified claim that three years of "effective" or "ineffective" teachers can make the difference between closing or maintaining the achievement gap. Apart from the circular reasoning of defining effectiveness by the same test scores that are used to claim that effectiveness labels are important, that slide suggested to me that this was probably a canned presentation for the most part rather than something targeted at Florida, apart from a transitory reference to Hillsborough's experimentation with teacher evaluation. There was also no acknowledgment in the presentation of recent research on the limits of value-added measures, and the slides all had "2009" in the bottom left-hand corner. So it was substantially based on a months-old canned presentation. To balance that, I know that few in the audience (including the legislators) would have had the widget-effect argument burned in their brain, and this was a general presentation. One surprise to me in the presentation: Weisberg emphasized the value of having observers who were independent of school principals and thus providing a facially-fair way of evaluating teachers, and in the Q&A (see below), he acknowledged that any of the systems he was describing as better than the modal status quo were still in their infancy. 

So, to the discussion after the presentation. The following are my sketchy notes of the Q&A, where Weisberg answered questions along with Commissioner Eric Smith and K-12 Chanceller Frances Haithcock. The Q&A was based entirely on questions generated by the in-person audience rather than from any questions submitted before the forum:

Q from faculty member at St. Petersburg College: Are you looking for a state-level evaluation system or guidance?

A from Commissioner Smith: Under Race to the Top, it really is district work, not a state-imposed plan. I think that's one of the great attributes of RTTT.

Q: Is there any attempt to inform the training ground for teachers in the university system to build a pipeline for more effective teachers?

A from Weisberg: I wouldn't say that's first generation work. That's an important component. … it provides a roadmap. To be fair to the ed schools, there is no such roadmap [right now].

A from Smith: It is something we're going at aggressively. What's really important… we've been publishing the performance of graduates from various institutions of schools of education and how well they have done in learning gains [of the graduates' students]. … This feedback is extraordinarily powerful. (Sherman's comment: Theoretically, it can be useful. In practice, FDOE is hundreds of miles away from the target. The state department of education measures are based on former state accountability officer Jerry Richardson's 10-year-old still-used yes-no classification of learning gains and a single-teacher assignment of students to "teachers responsible for a subject.") 

Q from a recent Hillsborough Teacher of the Year: Having looked at all the variations, if you could create the utopian evaluation system, what would it look like?

A from Weisberg: We've got our ideas about that… here's what I think, though we have strong opinions about what the performance descriptors should be. We should all readily admit that we don't know with any degree of precision of what practices translate directly to student achievement. … I'm not sure we're ever going to get there. Everyone's looking for the magic bullet, the formula… We wouldn't say that about surgeons or lawyers or conductors of symphonies. They've all got their own styles, and we know that about teaching. This is less about the holy grail and more about the process of continuing to refine it. … One of the things we can't do now is… what NFL scouts do, reverse-engineer.

Q from a parent in Pinellas: [After discussing the Schott Foundation report tagging Pinellas COunty as the worst large district in graduating African American males] I want to commend Commissioner Smith and the chancellor for standing up and beside you to consider another way we can turn this crisis around. … but we don't have principals and superintendents coming here. How do you, Commissioner Smith, get a buy-in from these administrators to aid the teachers' development? (This was one of the questions that appeared to be using the forum as an opportunity to ask officials about a different topic.) 

A from Smith: We have some evidence we're doing better in the state but nowhere close to where we need to be, and I dare say that's true of the nation. As a nation we've been fighting the fight of equity… for a number of decades. The greatest equity is the quality of the classroom teacher we bring. …

Q from Leo Haggerty, a member of the Hillsborough teachers union executive board: Vanderbilt's study say that teacher bonuses don't raise kids' test scores. That tells me this is a two-pronged effort, one to improve teachers and what is being done to improve parent and student areas where they are more accountable or where they become more effective?

A from Smith: I've read the study… It doesn't matter what you give me as an incentive if I'm not skilled and not prepared and not capable… bonuses don't make education happen. Bonuses can be reinforcers and motivators. The real question… is around evaluation.

Q: As a teacher-educator, I wholly support having highly qualified, highly effective teachers… and I applaud the state's efforts to begin… to hold all preparation to the same standard. I wonder if we look at that, you were talking about a roadmap. I would counter that we do have a roadmap, and that's our research base, particularly out of Stanford University that go far beyond test scores. (Sherman's comment: I think this was a reference to Linda Darling-Hammond's work.) 

A from Haithcock: I don't know the research you're talking about. The research we're looking at closely is the Harvard MET program. Look at the value-added that's been used in places… and see if they can equate that with a number of observation and evaluation strategies so that you have an equated system. The first information is coming out October 14 or later this month. I'm not real sure if that's what you're talking about, but we think this is probably the primary look at it.

A from Weisberg: I'm familiar with Darling-Hammond's work. There are a number of systems that have attempted to … correlate with better student outcomes. … I do think that the study the Chancellor's talking about will end up being very powerful.

Q: I applaud your vision and what you'd like to come out of the evaluation system. Those of us working in the field, my district is Hernando County. Right now we have a testing and accountability office with one person. At this moment, the state statute, RTTT, right now, this year, we have to have that data. Most teachers don't teach FCAT subjects. But they have to have the data tied to FCAT. … we're going to end up with so many teachers "needing improvement."

A from Haithcock: [Haithcock discussed general policy stuff but did not address the small-district problem or the fact that most teachers aren't in FCAT-tested subjects. The only part of her answer that was responsive was a mention of some county consortia that are supposed to develop additional assessments that aren't at the state level. Given that the teacher's question was a crucial question, I was disappointed and surprised at Haithcock's fudging on the answer.]

Q from mentor-teacher in Hillsborough: Once that [ineffectiveness] has been diagnosed, where do we go from there?… what type of time frame do we give … I guess I'm concerned about the tail end of the process.

A from Weisberg: We're working in a system where maybe you as a teacher have that feeling in the pit of your stomach but no one's told you that. … in our view, it would be a bad result to simply identify ineffective teachers and exit them summarily. Many of them will become effective teachers if given the right feedback, the right support. The time element is critical. First of all, you need to reintroduce the human element. These are human beings… you want to give them a legitimate opportunity to show that they can be an effective teacher. But if you say that if that process can take more than a year, … that teacher is going to get another group of kids. I know that as a parent, I certainly wouldn't want to opt to have my kid in a classroom with a teacher that has been identified by a fair system as ineffective. So that's hard.

A from Smith: Ditto. You do the value-added on the student achievement side, on a three-year average. The question is how quickly can you identify through the observation process a weakness…. [and] intervene.

Q from Lisa Johnston, a special ed teacher in Pinellas: We're aware that the key factor in the success of fundamental schools is 100% parental involvement. Is there going to be an assessment piece that talks about parent involvement?

A from Smith: There are some things that are in our control, and some are not in our control. The role of public education is to intervene in a child's life, where the child may not have the advantages that others have. … to create a life for themselves that their parents may not be experiencing. So… what do you control about that? It needs to be recognized. What I would see is a flexing of resource… talent… services. … All parents, regardless of the degree of involvement… are sending their kids to us to see significant learning gains.

Q from another faculty member from St. Petersburg College, Sue Blanchard: Right now, Florida uses high-stakes testing as the data that's sent out. .. I speak for the special educators and wonder… well, not all of our students take them, but we still need quality teachers for them. General ed teachers will be disinclined to work for inclusion. Is there something in the formula… the differences that children who come with extra challenges… that the teachers' effectiveness formula would be acknowledged, that extra work?

A from Haithcock: I think that's what we should be smart enough to develop…. the value-added model in itself… it's not a new concept. There are now a number of varieties of it. It's still not a seasoned instrument in the discussions we're having now. It's much improved… over the value tables [used in MAP, an attempt at a state-defined merit-pay system that few districts have participated in].

A from Weisberg: Why is it that high stakes tests… should have any impact on a teacher's evaluation? First, I haven't heard anyone talk about making value-added… more than one component of an evaluation. It is a critical component. … Why do you need that? If you don't have objectively measured student outcomes be one of the bases for performance evaluation of the people in the system, …. it is just too easy to say we did the best we could and we'll call it a day.

A from Smith: The issue of students with disabilities have come up a lot in our conversations int he last six months. I don't know any parents who are more involved in learning gains than parents of special-needs children… we are outperforming a majority of the states in our gains in NAEP with the special-needs population.

Q from Barbara Woolmarth, parent and reading coach and a member of the Florida Education Association governance board: I think it's extremely important… we're struggling as a district to figure out what other measures we can use to adequately and fairly evaluate teachers of students with disabilities. The two points you didn't bring up yet, around that thinking, are that especially with secondary-school students, many different instructors touch each student in the process of a three-year gain. How [do you] effectively control for that? But especially for our kids who are not taking standardized assessments [but instead are using] alternative assessment, we may see a "just-right moment" and intervention, and the need to be toileted may be a significant gain. But it may not be a continuous trend line [and instead looks like a plateau and then] a roller-coaster ride. To compare those teachers' results to other teachers' results is unfathomable.

A from Haithcock: [Haithcock refers to the IEP as an assessment system and then to the "teacher of record" issue.] Fortunately, we have a Gates grant with several other states… that are working on these very specific things… to do a much better job of defining the [teacher] of record. That information should be ready by the end of this calendar year.

A from Smith: There is a fair degree of differentiation between states on alternative assessments. We are also one of the lead states working to develop a new generation of alternative assessments for the core standards. Again, that'll be work that is ongoing, federally-funded. … as we develop the technique, the quality of support for special-needs teachers shall improve.

Q from Frank Whorter, a special education teacher and behavior specialist, and a former Pasco teacher of the year: The improvements in Florida have "been carried on the backs" of teachers.

A from Smith: We're working on Race to the Top, and a fairly well-defined time schedule for that. We've been working on the schedule of that with FEA [the Florida Education Association], and hopefully [we can] come to some local resolution of that and phasing it in over the four years. This has been something that's been worked out collaboratively.

Q from another special education teacher: You can tell we're pretty passionate about that. … Finland teaches students how to think. We teach students how to pass the test. … (The teacher then talks about co-teaching and class sizes, the teacher-mentor evaluations, etc…. looked like the teacher started with the topic at hand and then wandered. There wasn't much of an answer from the folks on the stage, though officials from the state department and I think Hillsborough talked to her afterwards–I think she was rightly upset that the peer evaluators in Hillsborough are telling co-teachers that they cannot work together during evaluations even though that's how they work.)

Q from Linda Boyd: Will evaluations be just between my principal and me, or …

Q from Thibault: Dan, want to talk about the L.A. Times…?

A from Weisberg: [Weisberg recounted the Times series and the publishing of database.] That may have been good journalism, but it's not good policy. … I think that's just a destructive thing. (Then he turned to Boyd and said that the culture has to change.) Teachers are treated as individual professionals and excellence is recognized and celebrated… yes, that means that you're going to have differentiation not just on some slide [in a presentation such as here] but within a school.

At this point, I got to ask my question. I had the facilitator return the presentation on the screen in front to a slide that showed the District of Columbia IMPACT weightings of different evaluation components for teachers in tested subjects and out of tested subjects. They were two pie charts, one showing test scores with a majority of the component weighting, and another showing test scores as a clear minority of the component weighting  I said something like the following: I just want to point out briefly that whatever you think of the DC evaluation system, if Senate Bill 6 had been signed in the spring [instead of being vetoed by Charlie Crist], no district in the state would have had the flexibility to use something like the DC split in weightings for evaluation. Dan Weisberg talked about how these new systems are experimental, and we're supposed to learn from them as they develop. The question is for Dan Weisberg, what is the dividing line between statutory language or regulations that are appropriate guidance and statutory language or state regulations that become a straightjacket?

A from Weisberg: I don't have a great answer… The real level where that's important is at the school and district level, and I don't think anyone's figured out what to do with that.

A from Smith: I don't really want to reargue Senate Bill 6, though it feels like I've been doing that for several months; it would have been a 4-year rollout… But there is a bigger issue. In my 38 years, I've never seen lots of money given to design a better system [as in Race to the Top]. … I think… the worst situation is if we do a milquetoast solution or walk away… if not, others may tell us what to do from on high.

Q from George-Ann Jones, teacher in Pasco and a building rep for the teachers union in Pasco County (just north of Hillsborough and Pinellas): When I was in Pueblo/Denver in the 1990s and early 2000s, … we were all very much involved in creating standards-driven education. We were very much involved in project-based education… Why are we starting this whole intervention with teachers rather than starting it with the management up above?

A from Smith: Our RTTT grant speaks directly to that. A lot of the conversation is around teachers, and we haven't spent a lot of time talking publicly about evaluation of assistant principals, principals, central office, and so forth. That's a parallel path, (along with) schools of education, teacher preparation programs, administrator preparation programs, … I think Chancellor Haithcock is going to pester me to death because she talks to me about leadership almost every day.

Q from Michelle Rolard, a mentor-teacher in Hillsborough: Here in Hillsborough we put into place a new evaluation system… I think that it's important for us to recognize that once those teachers have been identified, the piece that needs to go with that as well is the support for those teachers and a certain amount of time… one thing we're doing in Hillsborough County is the idea of support. Any effective responses in districts?

A from Weisberg: Yes… one of the things you can see once you have a meaningful distribution of teachers is which principals, coaches, APs, are doing a really good job of helping teachers. … one thing you can't do is layer a theoretically good evaluation system on top of the current system without changing priorities. One thing you have to tackle is, what are you doing to reorder your priorities?

Q from Otis Anthony, a district-level officer in the Polk County school system: Minority teachers are highly underrepresented. Has anyone given any thought to any impact these new evaluation systems may have on these numbers? Are you concerned about that at all?

A from Smith: Using this information to help the quality of teacher preparation programs, especially in historically-black colleges and universities, should help not only helping them to get into the profession but succeed.

And that's where the open forum ended, or at least my notes. The St. Pete Times's Tom Marshall wrote a blog entry (and an almost identical short article) on the forum. If there are any spelling differences between my report and his writing, please assume he's correct and I'm wrong. Some additional thoughts:

  • How to evaluate teachers of students with disabilities is obviously on the radar. What is not clear from the discussion is whether much discussion is focusing on English language learners, though a Q&A generally goes where the questions lead, and special educators were very well represented in those asking questions. 
  • At least in public, the commissioner and K-12 chancellor are talking about value-added measures and the development of new assessments as if they are adequate technical solutions to the political and human problems of evaluation. I know that this forum was on the general structure of teacher evaluation rather than the technical details (okay, weeds) of value-added measures, but I think anything that smacks of a Panglossian attitude on measures is a tactical and strategic mistake. If I had been up there on the stage, I would have said something like, "There is a healthy discussion about the value and limits of value-added or growth models. What's important here is to have components in an evaluation system that pay attention to student outcomes and to have an overall evaluation system that is structured in a way that is both meaningful and fair. How we get from one to the other should acknowledge the limits of technical tools but not ignore their potential." 
  • I suspect Smith's response to my question was aimed not at me (I'm just a faculty member at a university) but instead at some teachers a few rows behind me and the legislators in the room. To the teachers, I'm inferring from his comments that he meant, "You've got a limited amount of time, and four years at the outside, before any pig-headed resistance becomes an excuse for a heavy-handed approach." And to the legislators, I think he was pleading for more than six months before they try the heavy-handed approach. (Yes, I know what district those teachers are from. No, you don't get to learn which one from me; if you were there, you'll know.) 
  • For the most part, K-12 teachers asked important, focused questions (and questions that generated far more substantive responses than the questions from higher-ed faculty, me included). 

7 responses to “Dan Weisberg presentation at Hillsborough Community College, October 6”

  1. Glen McGhee

    “At least in public, the commissioner and K-12 chancellor are talking about value-added measures and the development of new assessments as if they are adequate technical solutions to the political and human problems of evaluation.”

    Yes, this was the message from the first VAM meeting in Sept. One questioner asked how measures developed for students could then be applied to teachers. The speaker replied that they could be, even if they were not as fine-grained as he would like them to be. I sat behind Martinez, and handed him a WSJ article by the “numbers guy” that showed just how unstable student measures were over a one year period, since he specifically asked questions about that.

  2. john thompson

    Sherman,

    How can Weisberg say the following and then issue that silly report they just issued?

    “We should all readily admit that we don’t know with any degree of precision of what practices translate directly to student achievement. … I’m not sure we’re ever going to get there. Everyone’s looking for the magic bullet, the formula… We wouldn’t say that about surgeons or lawyers or conductors of symphonies. They’ve all got their own styles, and we know that about teaching. This is less about the holy grail and more about the process of continuing to refine it. …

    In the study the TNTP wrote that teachers would be evaluated based on:

    “nearly all students at all skill levels master the lesson objective” using a checklist that is “specific, student-center-centered” and “leaves little room for inference …”

    Would he use such a checklist for a surgeon or a conductor?

    Even the 16 superintendents in the Washington Post Op Ed acknowledged the current impossibility of that sort of evaluation in classes of 25-30 with some kids reading at 4th grade and others reading from Tolstoy.

    I guess the didn’t know they have teachers where the range is lower than 4th grade and higher than 30 in a class.

    Also, do you think the Weisberg understands the issue of discipline and why that means that principals should be the last person to evaluate an actual teacher’s test score growth vs targets. Smith apparently sees it as a virtue that districts don’t have to require an independent evaluator. I’d hope that would be a deal-breaker

  3. CCPhysicist

    You have to take this document seriously. After all, it numbers its pages with leading zeros even though the page numbers never get into double digits. That said …

    I don’t understand the criticism from John Thompson. It is extremely clear on page 04 that they think teachers should not be evaluated on teacher behavior or routine or the paper version of the lesson plan or their bulletin boards, which is perfectly consistent with the statement quoted from Weisberg.

    I would use a checklist of patient outcomes to evaluate a surgeon. (I’d also include one of surgeon behavior, starting with “does surgeon use a checklist”, but that is a related issue since it impacts on patient outcomes.) Checklists are a great way to ensure that you don’t overlook something in one case that you focus on in another. One example is a grading key, what I now know is called a “rubric”, and a process that normally ensures that I check the same things on every paper.

    PS – I think if the principal had to personally go through the “student centered” progress reports for those 30 kids you describe with the teacher and an outside objective peer evaluator, the principal might get fired.

  4. john thompson

    CCPhysicist,

    Page four makes it clear that the TNTP is advocating something much worse that evaluating behavior. It is using a theory of learning that died with Descartes, claiming that we can objectively identify what each student has learned and evaluate teachers accordingly.

    I have often voiced support for checklists of teachers’ and surgeons’ behaviors for the reasons that you cited.

    But what would you consider a reasonable checklist of surgeons’ outcomes? If you use mortality stats, then we’ll just deny treatment to the sickest patients. That’s why even advocates of data-DRIVEN accountability for medicine have conceded that its best to focus on data-INFORMED accountability.

    Your PS further supports my argument. No principal would do himself what he would fire teachers for not doing. The purpose of giving teachers an impossibly long checklist is not to make them do everything on the checklist. The purpose is establishing control over the teacher. Under D.C.’s IMPACT or the TNTP scheme, administrators would not fire everyone. But they would require everyone to meet goals that were impossible to meet. They gives leverage to administrators to use or to abuse the system. The meeting of evaluation targets in inner city schools would be impossible so teachers’ futures would be based solely on the judgment of evaluators. And sense evaluators would be subject to meeting impossible goals, they would be pressured to protect themselves by blaming teachers.

    What they want, I believe, is to drive Baby Boomers out of the profession, or to at least shut us up. They don’t want to be bothered by our objections to their theories. They do not want to listen to our hard-earned experience.

    I’m asking an honest question, though, about what Weisberg believes. His TNTP has a long history of bait and switch, of saying things to the press and contradicts what they’ve written. They often write things that are contradicted by their own footnotes. They assert things, apprently based on PR strategies, that are manifestly untrue. They write inaccurate things about the union, and then issue a “reconciliation,” but not a retraction or an apology.

    Above all, they propose policies that make sense only if you grandiose beliefs about the reliability of standardized testing. In short, it is easy to see the legacy of Michelle Rhee in the TNTP. I’m legitimately interested in whether Weisberg’s speech means they are moving away from the slash and burn absolutism of the orgainization’s founder.

  5. J.F. Lesoine

    Circular reasoning all too often finds its way into politics. Reviewing teachers to help them improve as educators is not a bad idea, but there are so many variables from the real world that affect performance metrics that it is difficult to provide a fair and unbiased measure of educational effectiveness.

  6. john thompson

    Sherman,

    Were it up to me I’d go with the Toledo Plan.

    I’m not being flippant when I wish we could univent VAMs. They not only endanger evaluations and they not only threaten inner city schools more than affluent schools, but they threaten data-informed decision-making. We’d be much better off investing the money and time in diagnostic assessments and data system to combat truancy.

    But I’d support the Grand Bargain where peer review committees interpet data for evaluation purposes. I’d think that even diagnostic data would be useful in evaluations as long as it handled by peer review committees. I’d accept the New Haven Plan also.

    Let’s focus on what we want from evaluations. The #1 benefit would be efficiently removing the bottom 5 to 10% per year. Most of them are laughable apparent. Data would be most helpful in those cases in supplementing or complementing evalutions.

    And that’s a key point. Primitive data systems should never DRIVE accountability. We can’t allow primitive test scores to indict teachers as ineffective so they’ll have to then prove themselves innocent.

    Using data to help teachers improve is great. But that’s an argument for diagnostic data, not lowest common denominator standardized tests.

    And that’s why the TNTP would be the worst case scenario. If they believe what they wrote, its an invitation to create more standardized testing-driven schools.

    And that get’s me back, Sherman, to my question. What do Weisberg et. al really believe? Am I wrong, or are they trying to bring back an epistemology that has been obsolete for hundreds of years?

    Given their other spin, do they really believe that the way to attract and retain teaching talent is to expand high-stakes standardized testing?

    So to answer your question, I don’t think its a valid question. Our failure to evaluate comes from a lot of complicated reasons that are completely unrelated to standardized testing. The teacher-bashers wouldn’t be appeased by rubrics.

    Even though I think the dichotomy that you and the TNTP propose is a straw man. But if I HAD to chose between , I’d flip a coin between #1 or #3 and here’s why. Under any of the three scenarios, you are giving up on improving urban schools. The second scenario which is the backbone of the TNTP’s system is a desecration of the principles of education. It would drive integrity out of schools. And it would just make urban schools worse.

    It would be hard for #3 to be as bad as #1 in the short run. Unless you can create systems that are worthy of respect, those systems would quickly degenerate back to #1.

    If you were offering #4 with an equally flawed system ussing rubrbics and checks and balances, that would be significantly better.