It's teacher hunting season!
Showing posts with label value-added. Show all posts
Showing posts with label value-added. Show all posts

Friday, January 18, 2013

UPDATED: One Thousand Evaluation Petition-Signing Teachers Can't Be Wrong --The Real Story Behind the Evaluation Talks Collapse

UPDATES AT END: GOTHAM SCHOOLS LINKS WITH MORE ANALYSIS OF EVALUATIONS TALKS COLLAPSE - MULGREW CONTINUES TO MISS POINT - SPITEFUL NYC GOV'T SENDS INSPECTORS AFTER THE UFT The New York City are news was abuzz since 2:00 pm with news that talks between New York City mayor Michael Bloomberg and United Federation of Teachers Michael Mulgrew over teacher evaluations had broken down.

Some main points lost in the discussion:
UFT president Mulgrew has basically changed his posture on the question of the teacher evaluation system.

Developments seemed all pointed towards go for the high-stakes test based evaluation system (20 percent of a teacher's rating from state tests, 20 percent from local --read, city-- assessments). Then in a mid-December delegate assembly of the union, the MORE caucus advocated a democratic vote by the rank and file members of the union. The Daily News was rare among news outlets to catch the significance of the action, albeit, in a December 29, 2012 editorial, and without naming the active party involved:
Far more pertinent, at a union delegate assembly, a motion opposing Mulgrew’s authority to reach an evaluation deal with Chancellor Dennis Walcott — demanding instead that the matter be put to the membership — won a stunning 30% of the votes. A union president accustomed to 95% support then ran scared. Before that, by some accounts, Mulgrew and Walcott appeared to be progressing toward a deal even though Mulgrew veered far and wide, wanting to discuss even next year’s school closures.

But then, last week, he demanded that the city and the union must first settle how the new system will be implemented and rolled out, and how teachers will be trained in it . . . .

(Read more at the Daily News.)
Published yesterday morning, January 17, 2013, before the early afternoon announcement of the evaluation talks breakdown, "Potent Mix of Politics Shapes Current Education Debate" in the New York Times Schoolbook, Tim Clifford, a New York City teacher penned another distinctively accurate representation of the MORE strategic position in Mulgrew's changing posture. He pointed out that MORE's activities around the evaluation have Mulgrew nervous, as his Soviet-style 91 percent victory could be difficult to replicate unless he make some moves to co-opt the mass energy pushing back against the value-added test-based evaluation system. Truthfully, the hand-writing was clear at the start of the week, with MORE's announcement of a rally (for a membership vote) outside UFT headquarters, set for half an hour before the delegate assembly. Here is the latter part of Clifford's Times article, with the crucial election year elements included:
It’s likely that both the city and the UFT want an evaluation deal. For Bloomberg, this could be his last chance effect a major change in how teachers are hired and fired after several failed attempts to get rid of LIFO (Last In First Out) rules for excessing and layoffs. Yet he has insisted that the deal must include a means of holding teachers’ “feet to the fire” by making evaluations public, which is not required by state law. For its part, the UFT had a hand in crafting the new Annual Profession Performance Review (APPR) in the first place, helping limit efforts to make standardized test scores count for more than 25% of a teacher’s grade.

But there are other underlying political factors that may hinder an agreement. Foremost among these is the upcoming UFT election. Last time, Michael Mulgrew, then basically an unknown among teachers, won a staggering 91 percent of the vote as the protégé of outgoing president Randi Weingarten, facing no meaningful opposition. This time around, a new caucus has been gaining traction. This caucus, called MORE (Movement of Rank and File Educators), opposes any teacher evaluation agreement based on standardized test scores, which critics argue have a wide margin of error and other problems.

MORE’s candidate for president, Julie Cavanagh, is a well-spoken, well-regarded educator who is beginning to make a dent in Mulgrew’s hold on leadership. MORE’s recent resolution to have members vote on any evaluation deal, rather than union delegates mostly loyal to Mulgrew, garnered a significant amount of support. Said Cavanagh: “It is unacceptable that he (Mulgrew) does not recognize the truth: That the highest decision-making body of this union is its rank and file members. We should decide if ‘we as a union’ accept this: Because we are our union.”

Membership unrest in Chicago due to evaluations led to the ouster of the leadership there and conferred near hero status among unionists to Karen Lewis, who stood up to education reformers; the same could happen here if teachers are dissatisfied with the evaluation deal. And lest the potential mayoral candidates feel too sanguine, the story of Adrian Fenty, who was booted out as mayor of Washington, D.C. due largely to his support for test-obsessed Michelle Rhee, should act as a cautionary tale.

Complicating matters further is the teachers’ contract. The current deal expired in October, 2009, and the UFT did not receive the 4 + 4 percent over two years that other city workers got at the time. There is pressure on the union to settle a contract right away by tying evaluations to a new contract with higher wages, but there is also considerable sentiment that the UFT should wait out the Bloomberg era and try to get a favorable deal from the next mayor.

If Bloomberg and Mulgrew fail to come to terms on a contract, pressure will brought to bear on the current crop of mayoral hopefuls as to what kind of contract, with what kinds of wage increases, they’d be willing to sign. Democratic candidates are sure to vigorously court the UFT’s endorsement but by doing so they may risk losing financial support from Bloomberg, who will likely try to keep his reforms intact.

Other issues face the schools as we enter a new year. A bus strike is upon us. The city is looking to close 26 more schools, and is certain to be met with a fight. Governor Andrew Cuomo stepped into the fray in his State of the State address, calling for a teacher “bar” exam, as well as a longer school day and year that could add 300 hours to the school year without a clear means of financing those initiatives, which easily would cost billions of dollars in an age when school budgets have been cut every year for the last four years.

While the outcomes may not be certain, one thing is: 2013 promises to be a contentious year in education in New York. Whoever wins their political battles this year will likely affect the city’s schools well into the foreseeable future.
This writer is pleased that Mulgrew is speaking truth to power, calling mayor Bloomberg's assertions lies, however, the wish remains that he would directly and comprehensively reject the illogic undergirding these tests. Click again to the original, now classic, Gary Rubinstein statistical analyses of New York student test performances.

Why critical thinking is important: (Critical thinking ... something missing in the Common Core and the other new trends ... hmm.)
This chart of Value-Added Measures, from an article by math instructor Gary Rubinstein, demonstrates how no real correlation can be drawn between the scores of students in one year with generally the same students in another year. (Actually, good fortune of down servers preventing access to the original Rubinstein post brings us to another mathematician's quite scattered plotting of test results. See below.)


* * *

Moreover, this writer points out that Mulgrew is still enthusiastically defending the indefensible: see "MULGREW TELLS DELEGATES SCUTTLED NEW EVALUATION SYSTEM WOULD BE GREATEST THING SINCE SLICED BREAD" at the ICE-UFT blog.
This blog said the following last week: "The UFT is willing to concede on almost everything but Bloomberg's people may make it so humiliating that President Mulgrew would not even get a fig-leaf out of this. On the other hand, the Union could demand real safeguards (a right to grieve any unfair evaluations) so the DOE would reject any agreement." We were almost completely right except it looks like it was the mayor and not the DOE that inserted the poison pills. The fig-leaf was the two year sunset clause and the expedited arbitration if procedures weren't followed.
Trust me these were not great gains.

What happens next? I see the UFT going over the mayor's head to the state to try to get the system into law. What should people be doing? Call, email or talk to your union representatives, particularly Unity Chapter Leaders, and tell them you want no part of this and the real fight in Albany and Washington DC is to change the law so that no part of any teacher's rating is based on junk science.


* * *

UNITY-UFT STANDING IN THE WAY OF A MEMBERSHIP VOTE

Yes, there is no new evaluation system. But the UFT leadership remains tarnished for blocking a membership-wide vote on the system. For, contracts are voted on by the membership. The evaluation system has contract-like effects and significance.

Additionally unsettling is that the Unity leadership used its staff director lecture the delegates back in December about what the democratic representation scope is for the delegates. Certainly, a great error in principal. This excerpt from the ICE-UFT blog's report of that earlier Delegate Assembly:
Leroy Barr was called on to refute Kit's points. Leroy said that the membership elects Delegates and Chapter Leaders to represent members and the DA has a proud history of these duly elected representatives doing their job.


UPDATES:
Back to the analysis of the talks breakdown, Mulgrew still won't own up to the reality that his tentative evaluation agreement --yes, it is debatable as to whether there was some deal ready in the middle of the night-- was morbidly flawed, given that it was resting on illogical premises of the VAM testing models. His complaints have been around secondary, yet still important, side-issues. (The same can be said for Leo Casey writing at EdWize. He still has not rejected the fundamentally flawed VAM basis for the evaluations.) It's understandable that they will not own up to the essential core flaw of the high-stakes test-based evaluation system, for they has to save face.

The MORE caucus on its website, "Post-Mortem: The Non-Deal Between the UFT and DOE," has cited three critical reasons behind the failure of the evaluation deal, with discussion of each reason:
--Reason #1: Race to the Top is Bad Policy
--Reason #2: A Growing Backlash against Education Reform
--Reason #3: High-Handed and Un-Democratic School Leadership

We cannot entirely rest and must be watchful. There are rumors that the UFT might try to squeak in a deal in the next few weeks.

THE PERSISTENT PROBLEM OF THE UFT'S TOP-DOWN, UN-DEMOCRATIC MANNER
It is being argued, and rightly so, that the UFT ought to release the Memorandum of Understanding (MOU), so that the UFT members may see what teacher evaluation agreement almost was agreed to. And just as the Great Powers' problem of secret negotiations and World War I, there is a major problem for democracy when the issues in the MOU are being kept from the members, as though we are young children.

New York City sends inspectors after the United Federation of Teachers after the breakdown of evaluations talks, UFT president Mulgrew tweets.

Right-minded teachers ought to oppose this harassment, this intimidation of our union.

Sunday, January 13, 2013

Value Added Medicine?: Movement on Doctor's Compensation Comes to NYS

Teachers are under attack. A major tool is the evaluation system, which improperly relies heavily on commercialized standardized test results. The test results do not take into account what some economists call externalities, factors external to the immediate teaching process? Is the student prepared? Does he or she attend class regularly and with attentiveness? Is the school tone, set by school administration, problematic? Has the administrator assigned the teacher a group of students with more challenged circumstances than students in other classes in the school? Does poverty impact on the student's life?

Despite the illogic of holding test results central to teachers' performance, this method is being steamrolled through, by policy makers, test-promoters and media pundits.

In New York State, outcome-based evaluations are being used to evaluate doctors. Doctors will be rated according to the patient's outcome after he or she leaves her or his office. The same poverty concerns enter into consideration. Does the patient live in a high pollution area? Does the patient's job have high stress? Does the patient experience classism, racism, sexism or homophobia? Does police frisking create stress in the individual's life? Are there affordable fruits and vegetables available in a supermarket in easy access to the patient's home? Are there affordable gym facilities available to the patient? Then, there are the lifestyle issues? Does the patient floss her or his teeth? Does the patient have a preference for a healthier diet? Does the patient exercise? Does the patient take recreational drugs or smoke? Does the patient bicycle without a helmet or reflective gear?

Is all of this the doctor's responsibility? Will doctors be responsible for the effects that externalities play upon the patient's life? Will the doctors avoid working with high-poverty clients? Will doctors find it incumbent to their income to push aside patients that neglect to handle his or her life with proper life-style choices?

The New York Times yesterday opened its front page story, "New York City Ties Doctors’ Income to Quality of Care," on the topic with:
In a bold experiment in performance pay, complaints from patients at New York City’s public hospitals and other measures of their care — like how long before they are discharged and how they fare afterward — will be reflected in doctors’ paychecks under a plan being negotiated by the physicians and their hospitals.

The proposal represents a broad national push away from the traditional model of rewarding doctors for the volume of services they order, a system that has been criticized for promoting unnecessary treatment. In the wake of changes laid out in the Affordable Care Act, public and private hospitals are already preparing to have their income tied partly to patient outcomes and cost containment, but the city’s plan extends that financial incentive to the front line, the doctors directly responsible for treatment. It also shows how the new law could change longstanding relationships, giving more power to some of the poorest and most vulnerable patients over doctors who run their care.
The article closed with quotes from various doctors citing the difficulties with the approach. Indeed, Dr. David Himmelstein, professor at the City University of New York and a visiting professor at Harvard Medical School noted, “The consequences in a complex system like a hospital for giving an incentive for one little piece of behavior are virtually impossible to foresee.”
But Dr. Himmelstein said there were still hazards in the city’s plan. He said that when primary-care doctors in England were offered bonuses based on quality measures, they met virtually all of them in the first year, suggesting either that quality improved or — the more likely explanation, in his view — “they learned very quickly to teach to the test.”

“I think the most interesting finding is, things that were not measured, in a few studies, appeared to have gotten a bit worse,” Dr. Himmelstein said. For instance, patients were not as likely to stick with the same doctor, possibly because they were encouraged to see whichever doctor was available — speed was one quality measure — rather than the doctor who might know them best. In another example, while the doctors reported that they had controlled blood pressure in virtually all their patients, a random survey showed no downward trend in blood pressure or strokes.

There could have been any number of ways of outsmarting the system, he said: “If you take blood pressures three times and report the lowest, is that lying or merely tipping the numbers in your favor?”

Dr. Himmelstein also said doctors could try to avoid the sickest and poorest patients, who tend to have the worst outcomes and be the least satisfied. But physicians within the public hospital system have little ability to choose their patients, Mr. Aviles said. He added that he did not expect the doctors to act so cynically because, “in the main, physicians are here because they are attracted to that very mission of serving everybody equally.”

Sunday, September 23, 2012

Research Shows: Value-Added Teacher Accountability Not Ready for Prime Time

In Alternet research cited by supporters of the Chicago Public Schools' quest for value-added modeling (VAM) actually deflate pro-test arguments: "researchers were unable to compare the long term results of high value-added teachers with results of teachers who excelled in other ways that might, conceivably, have even larger impacts on long term outcomes."

Other research concluded: "It concluded that holding teachers accountable for growth in the test scores (called “value-added”) of their students is more harmful than helpful to children’s educations."

What does help? [This point is not the focus of the Alternet article.] A diverse curriculum and attention to children's basic health needs. These concerns are among the objectives of the Chicago Teachers Union strikers, and a fuller discussion can be read in their report, "The Schools Chicago's Children Deserve."

Writing in Alter Net, Richard Rothstein of the Economic Policy Institute raises another concern, so far, little addressed, even among observers that are critical of the link of testing to teacher evaluation: will principals, when they are informed of teachers' value-added scores, use that information in a biased manner against the teachers to skew their own observations of teachers?
"But according to the Chicago district proposal, the observations will be conducted by principals, who will know the value-added scores of teachers they are observing. How principals will be influenced by this knowledge cannot be known—will they tend to give high ratings to teachers with high value-added scores in order not to call attention to possible flaws in their observational skills, will they tend to offset value-added conclusions in order to save favored teachers who have low value-added, or will they tend to sink unfavored teachers with high value-added?"

Rothstein wondered how widespread teacher discontentment is. One specific kind of discontent, that over test-based teacher evaluations, is sure to grow exponentially, as legislation mandating such evaluations spreads like wildfire across the United States. This legislation has been spurred on by President Barack Obama and Education Secretary Arne Duncan's signature reform, Race to the Top. As of October 2011, when the National Council on Teacher Quality published a study, these 17 states and the District of Columbia are using "student achievement" as an "objective" role in assessing teacher performance: Arizona, Colorado, Delaware, D.C., Florida, Idaho, Illinois, Indiana, Louisiana, Maryland, Michigan, Minnesota, Nevada, New York, Ohio, Oklahoma, Rhode Island, and Tennessee.

Richard Rothstein debunks assumptions in the drive for value-added modeling teacher accountability in "Is 'Teacher Accountability' Ready for Prime-Time?," from Alter Net, September 17, 2012
Economic Policy Institute / By Richard Rothstein
Is 'Teacher Accountability' Ready for Prime-Time?
Though Rahm Emanuel wants to put student test scores at the center of teacher evals, there's little proof such measuring sticks make any sense.
September 17, 2012

It was bound to happen, whether in Chicago or elsewhere. What is surprising about the Chicago teachers’ strike is that something like this did not happen sooner.

The strike represents the first open rebellion of teachers nationwide over efforts to evaluate, punish and reward them based on their students’ scores on standardized tests of low-level basic skills in math and reading. Teachers’ discontent has been simmering now for a decade, but it took a well-organized union to give that discontent practical expression. For those who have doubts about why teachers need unions, the Chicago strike is an important lesson.

Nobody can say how widespread discontent might be. Reformers can certainly point to teachers who say that the pressure of standardized testing has been useful, has forced them to pay attention to students they previously ignored, and could rid their schools of lazy and incompetent teachers.

But I frequently get letters from teachers, and speak with teachers across the country who claim to have been successful educators and who are now demoralized by the transformation of teaching from a craft employing skill and empathy into routinized drill instruction using scripted curriculum. They are also demoralized by the weeks and weeks of the school year now devoted to gamesmanship—test preparation designed not to teach literacy or mathematics but only to make it seem that students can perform in an artificial setting better than they actually do.

I suspect, but cannot prove that the latter group of teachers is more numerous and that teachers in the discontented group are more likely to be seasoned, experienced, and successful. I suspect that teachers in the group supportive of standardized testing are more likely to be young, frequently hired outside the usual teacher training stream, and conditioned to think of education as little more than test preparation.

The research evidence is weighty in support of the discontented view; two years ago, EPI assembled a group of prominent testing experts and education policy experts to assess the research evidence on the use of test scores to evaluate teachers. [Eva L. Baker, Paul E. Barton, Linda Darling-Hammond, Edward Haertel, Helen F. Ladd, Robert L. Linn, Diane Ravitch, Richard Rothstein, Richard J. Shavelson, and Lorrie A. Shepard, "Problems with the Use of Student Test Scores to Evaluate Teachers"] It concluded that holding teachers accountable for growth in the test scores (called “value-added”) of their students is more harmful than helpful to children’s educations. Placing serious consequences for teachers on the results of their students’ tests creates rational incentives for teachers and schools to narrow the curriculum to tested subjects, and to tested areas within those subjects. Students lose instruction in history, the sciences, the arts, music, and physical education, and teachers focus less on development of children’s non-cognitive behaviors—cooperative activities, character, social skills—that are among the most important aims of a solid education.

Recently, however, some have made claims to the contrary that there are great benefits to holding teachers accountable for standardized test scores. One study, sponsored by the Gates Foundation, administered a higher quality test of reasoning and critical thinking skills to students who had also taken their state’s high stakes standardized test of basic skills. The Gates researchers found that teachers whose students had high value-added scores on the standardized basic skills test also tended to have high value-added scores on the test of reasoning (i.e., teachers’ value-added on the two tests were positively correlated). This was a potentially important finding because it suggested that narrowing the curriculum as a consequence of high stakes testing is not something about which we should be concerned. If we know that teachers who are effective at teaching basic skills are also effective at developing reasoning skills, then we can hold teachers accountable only for basic skills and be confident that their students are getting both.

But although the teacher results were correlated, they were only weakly correlated. True, more teachers who had high value-added scores on a basic skills test also had high value-added scores on a test of reasoning, but it wasn’t many more. If you fired teachers who did poorly at teaching basic skills you would get rid of many teachers who did poorly at developing reasoning skills, but you would also get rid of many teachers who did well at developing reasoning skills. The first group (those who did poorly) would be larger than the second group (those who did well), but not much larger.

The second highly publicized study, done by a group of Harvard researchers, concluded that teachers whose students had high value-added test scores were also those whose students had better long term adult outcomes—better earnings, for example. This was a potentially important finding because it suggested that these tests had not become ends in themselves, but rather that success for students on these tests made the students more likely to be successful as adults, and if you put pressure on teachers to increase their students’ test scores you would also be putting pressure on these teachers to improve their students’ adult success. And that would be a good thing.

The flaw here is that the researchers were unable to compare the long term results of high value-added teachers with results of teachers who excelled in other ways that might, conceivably, have even larger impacts on long term outcomes. For example, the researchers could not say whether teachers who are more effective at developing their students’ cooperative behavior, or reasoning skills (and we know from the Gates study that only sometimes are these the same teachers who are more effective at teaching basic skills) might have students who have even better adult outcomes—like earnings. If this were the case (and we have no reason to believe it one way or the other), then getting teachers to shift their attention from teaching reasoning or cooperative behavior to standardized test preparation might be lowering their students’ future earnings, not raising them.

In short, the two recent studies most heavily promoted by supporters of the Chicago district’s plan to evaluate teachers in part by their students’ test scores do not confirm that the district’s position is wise. It may be, but it also may do great harm.

The Chicago district, and other promoters of teacher evaluation based in large part on student test scores, have become aware of these problems. And so they now emphasize that they support evaluating teachers by “multiple measures”—not only their students’ test scores but by the performance by students of assigned tasks under the supervision of experts, by observation of teachers by their principals, and sometimes (for high school students, for example) by student reports of teacher effectiveness.

This is a fine balanced approach in theory, but is very difficult to implement in practice. For example, when the Gates Foundation study also showed a correlation between a teacher’s value-added test scores and a rated observation by instructional experts, it conducted this experiment by providing the experts with videotapes of teachers conducting instruction. The experts watching (and evaluating) the videotapes did not know the value-added scores of the teachers on the tapes, so the two measures (value-added scores and expert observation) were independent. But according to the Chicago district proposal, the observations will be conducted by principals, who will know the value-added scores of teachers they are observing. How principals will be influenced by this knowledge cannot be known—will they tend to give high ratings to teachers with high value-added scores in order not to call attention to possible flaws in their observational skills, will they tend to offset value-added conclusions in order to save favored teachers who have low value-added, or will they tend to sink unfavored teachers with high value-added? One thing of which we can be certain: Armed with knowledge of teacher value-added scores, it will be much harder for principals to observe and evaluate teachers objectively. In times past, when student test scores did not have high stakes for schools or teachers, principals with knowledge of test results could use this knowledge constructively to guide their observations; principals would visit classrooms where test scores were poor to see if they could determine something being done poorly, or visit classrooms where test scores were good to see if they could learn what was being done right. With high stakes now attached to the test, such constructive evaluation is less likely.

With a rush to implement test-based accountability before these systems have been tested experimentally, or even thought-through carefully, the Chicago district proposal is in some respects silly. What about teachers who don’t teach math or reading and so who don’t have standardized value-added scores? Or those who have not had students for a full year, or who have not been teaching the same subject for sufficient time to have value-added scores? The district proposes to evaluate these using their school-wide average value-added scores. Perhaps, with this proposal, the district is acknowledging that a teacher’s impact on a student is not only the result of her own efforts, but of the school’s entire teacher corps, working collaboratively. But if so, then individual teacher evaluation-by-test-score makes no sense (even if individual teacher data happen to be available), and student growth data should be used only to evaluate schools as a whole.

The impact of teachers’ practice on each other is apparent. Should, for example, a fifth grade teacher’s value-added score be adjusted if her students had come from a class the year before with a fourth grade teacher whose value-added score was unusually high or low? With similar students, a fifth grade teacher will have an easier (or perhaps harder) time if her students had a more effective teacher in the fourth grade. It could be easier if students had a more effective teacher the previous year, because the skills students learn in one year can give them an advantage in learning in subsequent years. Or a teacher’s job could be harder if students had a more effective teacher the previous year, because students who learn more in one year will have less room to grow in the next year. Nobody, no educational theorist or practical policy maker has an answer to this problem, and the Chicago proposal ignores this obvious source of distortion.

News reports suggest that the Chicago strike may be settled soon. The district’s latest proposal is that value-added test score data could ultimately make up 25 percent of a teacher’s evaluation, with student growth on some yet-to-be-defined “performance” tasks—writing an essay, for example—comprising another 15 percent. This is far better than what some districts around the country are attempting to do, with standardized test score data making up half of a teacher’s evaluation. The union will likely agree to something close to the district proposal, with an appeals process that is stronger than what the district has thus far proposed.

Although the Chicago teacher evaluation system will not be the worst in the country, it will still rely on methods that are not yet ready for prime time. Whether the Chicago strike slows down the rush in other places to implement a terribly flawed system, or the settlement encourages other places to try it, remains to be seen.

Richard Rothstein is a research associate at the Economic Policy Institute and senior fellow at the Chief Justice Earl Warren Institute on Law and Social Policy at U.C. Berkeley.

Wednesday, September 12, 2012

Valerie Strauss: Why Rahm Emanuel and The New York Times are Wrong about Teacher Evaluation

Posted at 12:53 PM ET, 09/12/2012 at the Washington Post
Why Rahm Emanuel and The New York Times are wrong about teacher evaluation
By Valerie Strauss, The Answer Sheet column
You know things are going very badly for public school teachers when The New York Times editorial board calls a bad teacher evaluation system a “sensible policy change.”

The Times ran an editorial on Wednesday that smacked Chicago teachers for striking against a school reform package pushed by Mayor Rahm Emanuel, a former chief of staff of President Obama. It says in part:
Teachers’ strikes, because they hurt children and their families, are never a good idea. The strike that has roiled the civic climate in Chicago— and left 350,000 children without classes — seems particularly senseless because it is partly a product of a personality clash between the blunt mayor, Rahm Emanuel, and the tough Chicago Teachers Union president, Karen Lewis. Beyond that, the strike is based on union discontent with sensible policy changes — including the teacher evaluation system required by Illinois law — that are increasingly popular across the country and are unlikely to be rolled back, no matter how long the union stays out. The Washington Post editorial board has also supported test-based teacher evaluation, including in this piece on the strike.
The Post editorial says that “the system developed by Chicago officials, on which they offered to work with the union, is careful to measure student growth. That means teachers aren’t blamed if their students start out behind but instead are evaluated on their ability to make progress during the year they have responsibility.” Well, a single test is hardly the way to tell whether a child has made progress. How many of you have had a headache or been sick or emotionally upset and bombed a test? There are ways to measure how students are achieving without a standardized test.

Think what you want about the Chicago teachers strike. But that doesn’t change this:

The Times can say that using standardized test scores to evaluate teachers is a sensible policy and Obama can say it and Education Secretary Arne Duncan can say it and Emanuel can say it and so can Bill Gates (who has spent hundreds of millions of dollars to develop it) and governors and mayor from both parties, and heck, anybody can go ahead and shout it out as loud as they can.

It doesn’t make it true.

Can all these very smart people be wrong? Yes, according to many experts on assessment who have done extensive research on the subject.

These experts have said over and over and over that the method by which test scores are factored into an evaluation of how effective a teacher is are dramatically unreliable and unfair. Some say it will destroy the teaching profession because it will identify effective teachers as ineffective and ineffective teachers as effective. Some bad teachers will be fired but some good ones will too. Others will leave in disgust.

That’s what happened, for example, in New York City when Carolyn Abbott, who teaches mathematics to seventh- and eighth-graders at the Anderson School, a citywide gifted-and-talented school on the Upper West Side of Manhattan, learned that her “value-added” score made her the worst eighth grade teacher in the entire city. The score of course didn’t reflect that her students already scored near 100 percent proficiency and were doing advanced math — but the formula didn’t care.

The value-added formulas actually compare how students are predicted to perform on the state ELA and math tests, based on their prior year’s performance, with their actual performance, as Teachers College Professor Aaron Pallas wrote here. Teachers whose students do better than predicted are said to have “added value”; those whose students do worse than predicted are “subtracting value.” By definition, he wrote, about half of all teachers will add value, and the other half will not.

No, Abbott’s case wasn’t an aberration. Lots of scores are wrong. Yet state after state insists on foisting this on teachers and even principals.

This isn’t just about the adults. Kids suffer when good teachers are said to be bad and bad teachers are said to be good, and especially when standardized tests have such high stakes that teachers feel forced to tailor their teaching to the test.

And think about this: If teachers are evaluated on test scores, there has to be standardized test for every class. What happens when tests have such high stakes? Kids learn how to pass tests rather than how to solve problems and think creatively.

Even the former education commissioner of Texas, Republican Robert Scott, recognized this and said earlier this year that all of this testing is a “perversion” of what a quality education should be.

Of course teachers should be evaluated — and evaluated better than they have been in most places for decades. Yes, bad teachers should be removed from the classroom, sooner rather than later. There are ways to do this that is fair, and it is already being done in places such as the high-achieving Montgomery County Public School system in Maryland.

In fact, The New York Times’ own columnist Michael Winerip wrote about that evaluation system last year in this story, which noted that then-Superintendent Jerry Weast had rejected $12 million in Race to the Top money because it required districts to use test scores to evaluate teachers. Weast was quoted as saying: “We don’t believe the tests are reliable. You don’t want to turn your system into a test factory.” Most Washington D.C.-area superintendents still think it’s a bad idea, including Weast’s successor, Joshua Starr. (Unfortunately Winerip doesn’t write about education anymore for The Times.)

In a Q & A with area superintendents, I quoted Loudoun County Public Schools Superintendent Edgar Hatrick as saying, “It is troubling that we will now take tests we’re not sure are good measures of student performance and extrapolate teacher performance from student scores.”

The Illinois law calls for an evaluation system in which at least 20 percent is based on student standardized test scores. Emanuel wants fully half of the evaluation to be based on the test scores.

(Incidentally, former Washington D.C. schools chancellor Michelle Rhee instituted a teacher evaluation system a few years ago that had 50 percent of individual assessments linked to student test scores — in courses where standardized tests were given — but her successor, Kaya Henderson, just dropped it down to 35 percent because of problems with the system.)

Read this from a letter that scores of researchers from 16 universities throughout the Chicago metropolitan area sent to Emanuel warning against the “value-added” system of teacher involvement, which uses complicated formulas to factor test scores into an evaluation:
As university professors and researchers who specialize in educational research, we recognize that change is an essential component of school improvement. We are very concerned, however, at a continuing pattern of changes imposed rapidly without high-quality evidentiary support.
The new evaluation system for teachers and principals centers on misconceptions about student growth, with potentially negative impact on the education of Chicago’s children. We believe it is our ethical obligation to raise awareness about how the proposed changes not only lack a sound research basis, but in some instances, have already proven to be harmful.
Professors in Georgia sent a letter to their governor against value-added evaluations. More than 1,500 New York principals and more than 5,400 teachers, parents, professors, administrators and citizens have signed an open letter blasting that state’s value-added evaluation system, the letter which can be found here.

The National Research Council, the research arm of the National Academies, which include the National Academy of Sciences, the National Academy of Engineering and the Institute of Medicine, issued a major report last year on this issue that said:
The standardized test scores that have been trumpeted to show improvement in the schools provide limited information about the causes of improvements or variability in student performance.This would be true, presumably, for any school system that use standardized tests as a measure of achievement.
Mathematicians have said that value-added models are hardly ready for prime time as teacher evaluation tools, and that includes John Ewing, president of Math for America, a nonprofit organization dedicated to improving mathematics education in U.S. public high schools. In this post he said:
The most common misuse of mathematics is simpler, more pervasive, and (alas) more insidious: mathematics employed as a rhetorical weapon—an intellectual credential to convince the public that an idea or a process is “objective” and hence better than other competing ideas or processes. This is mathematical intimidation. It is especially persuasive because so many people are awed by mathematics and yet do not understand it—a dangerous combination.

The latest instance of the phenomenon is valued-added modeling (VAM), used to interpret test data. Value-added modeling pops up everywhere today, from newspapers to television to political campaigns. VAM is heavily promoted with unbridled and uncritical enthusiasm by the press, by politicians, and even by (some) educational experts, and it is touted as the modern, “scientific” way to measure educational success in everything from charter schools to individual teachers.

Yet most of those promoting value-added modeling are ill-equipped to judge either its effectiveness or its limitations. Some of those who are equipped make extravagant claims without much detail, reassuring us that someone has checked into our concerns and we shouldn’t worry. Value-added modeling is promoted because it has the right pedigree — because it is based on “sophisticated mathematics.” As a consequence, mathematics that ought to be used to illuminate ends up being used to intimidate.
That’s what value-added is doing: Intimidating teachers, and many of them believe that it will be the end of the teaching profession. Who would want to go into a profession where a big part of the evaluation system is faulty?

Follow The Answer Sheet every day by bookmarking www.washingtonpost.com/blogs/answer-sheet.

By Valerie Strauss | 12:53 PM ET, 09/12/2012, the Answer Sheet column at the Washington Post

Sunday, April 29, 2012

If it’s not valid, reliability doesn’t matter so much! More on VAM-ing & SGP-ing Teacher Dismissal

Here is one powerful critique, skewering the logic of the VAM reasoning in teacher "assessment" from School Finance 101: Data and thoughts on public and private school funding in the U.S. Because of its workman/woman virtues I have included the entire post.
If it’s not valid, reliability doesn’t matter so much! More on VAM-ing & SGP-ing Teacher Dismissal [note: VAM: value-added measurement; SGP: student growth percentile] Posted on April 28, 2012
This post includes a few more preliminary musings regarding the use of value-added measures and student growth percentiles for teacher evaluation, specifically for making high-stakes decisions, and especially in those cases where new statutes and regulations mandate rigid use/heavy emphasis on these measures, as I discussed in the previous post.
========
The recent release of New York City teacher value-added estimates to several media outlets stimulated much discussion about standard errors and statistical noise found in estimates of teacher effectiveness derived from the city’s value-added model. But lost in that discussion was any emphasis on whether the predicted value-added measures were valid estimates of teacher effects to begin with. That is, did they actually represent what they were intended to represent – the teacher’s influence on a true measure of student achievement, or learning growth while under that teacher’s tutelage. As framed in teacher evaluation legislation, that measure is typically characterized as “student achievement growth,” and it is assumed that one can measure the influence of the teacher on “student achievement growth” in a particular content domain.

A brief note on the semantics versus the statistics and measurement in evaluation and accountability is in order.

At issue are policies involving teacher “evaluation” and more specifically evaluation of teacher effectiveness, where in cases of dismissal the evaluation objective is to identify particularly ineffective teachers.

In order to “evaluate” (assess, appraise, estimate) a teacher’s effectiveness with respect to student growth, one must be able to “infer” (deduce, conjecture, surmise…) that the teacher affected or could have affected that student growth. That is, for example, given one year’s bad rating, the teacher had sufficient information to understand how to improve her rating in the following year. Further, one must choose measures that provide some basis for such inference.

Inference and attribution (ascription, credit, designation) are not separable when evaluating teacher effectiveness. To make an inference about teacher effectiveness based on student achievement growth, one must attribute responsibility for that growth to the teacher.

In some cases, proponents of student growth percentiles alter their wording [in a truly annoying & dreadfully superficial way] for general public appeal to argue that:

SGPs are a measure of student achievement growth. Student achievement growth is a primary objective of schooling. Therefore, teachers and schools should obviously be held accountable for student achievement growth.

Where accountable is a synonym for responsible, to the extent that SGPs were designed to separate the measurement of student growth from attribution of responsibility for it, then SGPs are also invalid on their face for holding teachers accountable. For a teacher to be accountable for that growth it must be attributable to them and one must be using a method which permits such inference.

Allow me to reiterate this quote from the authors of SGP:

“The development of the Student Growth Percentile methodology was guided by Rubin et al’s (2004) admonition that VAM quantities are, at best, descriptive measures.” (Betebenner, Wenning & Briggs, 2011)

I will save for another day a discussion of the nuanced differences between statistical causation and inference and causation and inference as might be evaluated more broadly in the context of litigation over determination of teacher effectiveness. The big problem in the current context, as I have explained in my previous post, is created by legislative attempts to attach strict timelines, absolute weights and precise classifications to data that simply cannot be applied in this way.

Major Validity Concerns
We identify [at least] 3 categories of significant compromises to inference and attribution and therefore accountability for student achievement growth:

The value-added estimate (or SGP) was influenced by something other than the teacher alone The value-added (or SGP) estimate given one assessment of the teacher’s content domain produces a different rating than the value-added estimate given a different assessment tool The value-added estimate (or SGP) is compromised by missing data and/or student mobility, disrupting the link between teacher and students. [the actual data link required for attribution]

The first major issue compromising attribution of responsibility for or inference regarding teacher effectiveness based on student growth is that some other factor or set of factors actually caused the student achievement growth or lack thereof. A particularly bothersome feature of many value-added models is that they rely on annual testing data. That is, student achievement growth is measured from April or May in one year to April or May in the next, where the school year runs from September to mid or late June. As such, for example, the 4th grade teacher is assigned a rating based on children who attended her class from September to April (testing time), or about 7 months, where 2.5 months were spent doing any variety of other things, and another 2.5 months were spent with their prior grade teacher. Let alone the different access to resources each child has during their after school and weekend hours during the 7 months over which they have contact with their teacher of record.

Students with different access to summer and out-of-school time resources may not be randomly assigned across teachers within a given school or across schools within a district. And students who had prior year teachers who may have checked out versus the teacher who delved into the subsequent year’s curriculum during the post-testing month of the prior year may also not be randomly distributed. All of these factors go unobserved and unmeasured in the calculation of a teacher’s effectiveness, potentially severely compromising the validity of a teacher’s effectiveness estimate. Summer learning varies widely across students by economic backgrounds (Alexander, Entwisle & Olsen, 2001) Further, in the recent Gates MET Studies (2010), the authors found: “The norm sample results imply that students improve their reading comprehension scores just as much (or more) between April and October as between October and April in the following grade. Scores may be rising as kids mature and get more practice outside of school.” (p. )

Numerous authors have conducted analyses revealing the problems of omitted variables bias and the non-random sorting of students across classrooms (Rothstein, 2011, 2010, 2009, Briggs & Domingue, 2011, Ballou et al., 2012). In short, some value-added models are better than others, in that by including additional explanatory measures, the models seem to correct for at least some biases. Omitted variables bias is where any given teacher’s predicted value is influenced partly by factors other than the teacher herself. That is, the estimate is higher or lower than it should be, because some other factor has influenced the estimate. Unfortunately, one can never really know if there are still additional factors that might be used to correct for that bias. Many such factors are simply unobservable. Others may be measurable and observable but are simply unavailable, or poorly measured in the data. While there are some methods which can substantially reduce the influence of unobservables on teacher effect estimates, those methods can typically only be applied to a very small subset of teachers within very large data sets.[2] In a recent conference paper, Ballou and colleagues evaluated the role of omitted variables bias in value-added models and the potential effects on personnel decisions. They concluded:

“In this paper, we consider the impact of omitted variables on teachers’ value-added estimates, and whether commonly used single-equation or two-stage estimates are preferable when possibly important covariates are not available for inclusion in the value-added model. The findings indicate that these modeling choices can significantly influence outcomes for individual teachers, particularly those in the tails of the performance distribution who are most likely to be targeted by high-stakes policies.” (Ballou et al., 2012)

A related problem is the extent to which such biases may appear to be a wash, on the whole, across large data sets, but where specific circumstances or omitted variables may have rather severe effects on predicted values for specific teachers. To reiterate, these are not merely issues of instability or error. These are issues of whether the models are estimating the teacher’s effect on student outcomes, or the effect of something else on student outcomes. Teachers should not be dismissed for factors beyond their control. Further, statutes and regulations should not require that principals dismiss teachers or revoke their tenure in those cases where the principal understands intuitively that the teacher’s rating was compromised by some other cause. [as would be the case under the TEACHNJ Act]

Other factors which severely compromise inference and attribution, and thus validity, include the fact that the measured value-added gains of a teacher’s peers – or team members working with the same students – may be correlated, either because of unmeasured attributes of the students or because of spillover effects of working alongside more effective colleagues (one may never know) (Koedel, 2009, Jackson & Bruegmann, 2009). Further, there may simply be differences across classrooms or school settings that remain correlated with effectiveness ratings that simply were not fully captured by the statistical models.

Significant evidence of bias plagued the value-added model estimated for the Los Angeles Times in 2010, including significant patterns of racial disparities in teacher ratings both by the race of the student served and by the race of the teachers (see Green, Baker and Oluwole, 2012). These model biases raise the possibility that Title VII disparate impact claims might also be filed by teachers dismissed on the basis of their value-added estimates. Additional analyses of the data, including richer models using additional variables mitigated substantial portions of the bias in the LA Times models (Briggs & Domingue, 2010).

A handful of studies have also found that teacher ratings vary significantly, even for the same subject area, if different assessments of that subject are used. If a teacher is broadly responsible for effectively teaching in their subject area, and not the specific content of any one test, different results from different tests raise additional validity concerns. Which test better represents the teacher’s responsibilities? [must we specify which test counts/matters/represents those responsibilities in teacher contracts?] If more than one, in what proportions? If results from different tests completely counterbalance, how is one to determine the teacher’s true effectiveness in their subject area? Using data on two different assessments used in Houston Independent School District, Corcoran and Jennings (2010) find:

[A]mong those who ranked in the top category (5) on the TAKS reading test, more than 17 percent ranked among the lowest two categories on the Stanford test. Similarly, more than 15 percent of the lowest value-added teachers on the TAKS were in the highest two categories on the Stanford.

The Gates Foundation Measures of Effective Teaching Project also evaluated consistency of teacher ratings produced on different assessments of mathematics achievement. In a review of the Gates findings, Rothstein (2010) explained:

The data suggest that more than 20% of teachers in the bottom quarter of the state test math distribution (and more than 30% of those in the bottom quarter for ELA) are in the top half of the alternative assessment distribution.(p. 5)

And:
In other words, teacher evaluations based on observed state test outcomes are only slightly better than coin tosses at identifying teachers whose students perform unusually well or badly on assessments of conceptual understanding.(p. 5)

Finally, student mobility, missing data, and algorithms for accounting for that missing data can severely compromise inferences regarding teacher effectiveness. Corcoran (2010) explains that the extent of missing data can be quite large and can vary by student type:

Because of high rates of student mobility in this [Houston] population (in addition to test exemption and absenteeism), the percentage of students who have both a current and prior year test score – a prerequisite for value-added – is even lower (see Figure 6). Among all grade four to six students in HISD, only 66 percent had both of these scores, a fraction that falls to 62 percent for Black students, 47 percent for ESL students, and 41 percent for recent immigrants.” (Corcoran, 2010, p.20- 21)

Thus, many teacher effectiveness ratings would be based on significantly incomplete information, and further, the extent to which that information is incomplete would be highly dependent on the types of students served by the teacher.

One statistical resolution to this problem is imputation. In effect, imputation creates pre-test or post-test scores for those students who weren’t there. One approach is to use the average score for students who were there, or more precisely for otherwise similar students who were there. On its face imputation is problematic when it comes to attribution of responsibility for student outcomes to the teacher, if some of those outcomes are statistically generated for students who were not even there. But not using imputation may lead to estimates of effectiveness that are severely biased, especially when there is so much missing data. Howard Wainer (2011) esteemed statistician and measurement expert formerly with Educational Testing Service (ETS) explains somewhat mockingly how teachers might game imputation of missing data by sending all of their best students on a field trip during fall testing days, and then, in the name of fairness, sending the weakest students on a field trip during spring testing days.[3] Clearly, in such a case of gaming, the predicted value-added assigned to the teacher as a function of the average scores of low performing students at the beginning of the year (while their high performing classmates were on their trip), and high performing ones at the end of the year (while their low performing classmates were on their trip), would not be correctly attributed to the teacher’s actual teaching effectiveness, though it might be attributable to the teacher’s ability to game the system.

In short, validity concerns are at least as great as reliability concerns, if not greater. If a measure is simply not valid, it really doesn’t matter whether it is reliable or not.

If a measure cannot be used to validly infer teacher effectiveness, cannot be used to attribute responsibility for student achievement growth to the teacher, then that measure is highly suspect as a basis for high stakes decisions making when evaluating teacher (or teaching) effectiveness or for teacher and school accountability systems more generally.

References & Additional Readings

Alexander, K.L, Entwisle, D.R., Olsen, L.S. (2001) Schools, Achievement and Inequality: A Seasonal Perspective. Educational Evaluation and Policy Analysis 23 (2) 171-191 Ballou, D., Mokher, C.G., Cavaluzzo, L. (2012) Using Value-Added Assessment for Personnel Decisions: How Omitted Variables and Model Specification Influence Teachers’ Outcomes. Annual Meeting of the Association for Education Finance and Policy. Boston, MA. http://aefpweb.org/sites/default/files/webform/AEFP-Using%20VAM%20for%20personnel%20decisions_02-29-12.docx Ballou, D. (2012). Review of “The Long-Term Impacts of Teachers: Teacher Value-Added and Student Outcomes in Adulthood.” Boulder, CO: National Education Policy Center. Retrieved [date] from http://nepc.colorado.edu/thinktank/review-long-term-impacts Baker, E.L., Barton, P.E., Darling-Hammong, L., Haertel, E., Ladd, H.F., Linn, R.L., Ravitch, D., Rothstein, R., Shavelson, R.J., Shepard, L.A. (2010) Problems with the Use of Student Test Scores to Evaluate Teachers. Washington, DC: Economic Policy Institute. http://epi.3cdn.net/724cd9a1eb91c40ff0_hwm6iij90.pdf Betebenner, D., Wenning, R.J., Briggs, D.C. (2011) Student Growth Percentiles and Shoe Leather. http://www.ednewscolorado.org/2011/09/13/24400-student-growth-percentiles-and-shoe-leather Boyd, D.J., Lankford, H., Loeb, S., & Wyckoff, J.H. (July, 2010). Teacher layoffs: An empirical illustration of seniority vs. measures of effectiveness. Brief 12. National Center for Evaluation of Longitudinal Data in Education Research. Washington, DC: The Urban Institute. Briggs, D., Betebenner, D., (2009) Is student achievement scale dependent? Paper presented at the invited symposium Measuring and Evaluating Changes in Student Achievement: A Conversation about Technical and Conceptual Issues at the annual meeting of the National Council for Measurement in Education, San Diego, CA, April 14, 2009. http://dirwww.colorado.edu/education/faculty/derekbriggs/Docs/Briggs_Weeks_Is%20Growth%20in%20Student%20Achievement%20Scale%20Dependent.pdf Briggs, D. & Domingue, B. (2011). Due Diligence and the Evaluation of Teachers: A review of the value-added analysis underlying the effectiveness rankings of Los Angeles Unified School District teachers by the Los Angeles Times. Boulder, CO: National Education Policy Center. Retrieved [date] from http://nepc.colorado.edu/publication/due-diligence. Budden, R. (2010) How Effective Are Los Angeles Elementary Teachers and Schools?, Aug. 2010, available at http://www.latimes.com/media/acrobat/2010-08/55538493.pdf. Braun, H, Chudowsky, N, & Koenig, J (eds). (2010) Getting value out of value-added. Report of a Workshop. Washington, DC: National Research Council, National Academies Press. Braun, H. I. (2005). Using student progress to evaluate teachers: A primer on value-added models. Princeton, NJ: Educational Testing Service. Retrieved February, 27, 2008. Chetty, R., Friedman, J., Rockoff, J. (2011) The Long Term Impacts of Teachers: Teacher Value Added and Student outcomes in Adulthood. NBER Working Paper # 17699 http://www.nber.org/papers/w17699 Clotfelter, C., Ladd, H.F., Vigdor, J. (2005) Who Teaches Whom? Race and the distribution of Novice Teachers. Economics of Education Review 24 (4) 377-392 Clotfelter, C., Glennie, E. Ladd, H., & Vigdor, J. (2008). Would higher salaries keep teachers in high-poverty schools? Evidence from a policy intervention in North Carolina. Journal of Public Economics 92, 1352-70. Corcoran, S.P. (2010) Can Teachers Be Evaluated by their Students’ Test Scores? Should they Be? The Use of Value Added Measures of Teacher Effectiveness in Policy and Practice. Annenberg Institute for School Reform. http://annenberginstitute.org/pdf/valueaddedreport.pdf Corcoran, S.P. (2011) Presentation at the Institute for Research on Poverty Summer Workshop: Teacher Effectiveness on High- and Low-Stakes Tests (Apr. 10, 2011), available at https://files.nyu.edu/sc129/public/papers/corcoran_jennings_beveridge_2011_wkg_teacher_effects.pdf. Corcoran, Sean P., Jennifer L. Jennings, and Andrew A. Beveridge. 2010. “Teacher Effectiveness on High- and Low-Stakes Tests.” Paper presented at the Institute for Research on Poverty summer workshop, Madison, WI. D.C. Pub. Sch., IMPACT Guidebooks (2011), available at http://dcps.dc.gov/portal/site/DCPS/menuitem.06de50edb2b17a932c69621014f62010/?vgnextoid=b00b64505ddc3210VgnVCM1000007e6f0201RCRD. Education Trust (2011) Fact Sheet- Teacher Quality. Washington, DC. http://www.edtrust.org/sites/edtrust.org/files/Ed%20Trust%20Facts%20on%20Teacher%20Equity_0.pdf Hanushek, E.A., Rivkin, S.G., (2010) Presentation for the American Economic Association: Generalizations about Using Value-Added Measures of Teacher Quality 8 (Jan. 3-5, 2010), available at http://www.utdallas.edu/research/tsp-erc/pdf/jrnl_hanushek_rivkin_2010_teacher_quality.pdf Working with Teachers to Develop Fair and Reliable Measures of Effective Teaching. MET Project White Paper. Seattle, Washington: Bill & Melinda Gates Foundation, 1. Retrieved December 16, 2010, from http://www.metproject.org/downloads/met-framing-paper.pdf. Learning about Teaching: Initial Findings from the Measures of Effective Teaching Project. MET Project Research Paper. Seattle, Washington: Bill & Melinda Gates Foundation. Retrieved December 16, 2010, from http://www.metproject.org/downloads/Preliminary_Findings-Research_Paper.pdf. Jackson, C.K., Bruegmann, E. (2009) Teaching Students and Teaching Each Other: The Importance of Peer Learning for Teachers. American Economic Journal: Applied Economics 1(4): 85–108 Kane, T., Staiger, D., (2008) Estimating Teacher Impacts on Student Achievement: An Experimental Evaluation. NBER Working Paper #16407 http://www.nber.org/papers/w14607 Koedel, C. (2009) An Empirical Analysis of Teacher Spillover Effects in Secondary School. 28 (6 ) 682-692 Koedel, C., & Betts, J. R. (2009). Does student sorting invalidate value-added models of teacher effectiveness? An extended analysis of the Rothstein critique. Working Paper. Jacob, B. & Lefgren, L. (2008). Can principals identify effective teachers? Evidence on subjective performance evaluation in education. Journal of Labor Economics. 26(1), 101-36. Sass, T.R., (2008) The Stability of Value-Added Measures of Teacher Quality and Implications for Teacher Compensation Policy. National Center for Analysis of Longitudinal Data in Educational Research. Policy Brief #4. http://eric.ed.gov/PDFS/ED508273.pdf McCaffrey, D. F., Lockwood, J. R, Koretz, & Hamilton, L. (2003). Evaluating value-added models for teacher accountability. RAND Research Report prepared for the Carnegie Corporation. McCaffrey, D. F., Lockwood, J. R., Koretz, D., Louis, T. A., & Hamilton, L. (2004). Models for value-added modeling of teacher effects. Journal of Educational and Behavioral Statistics, 29(1), 67. Rothstein, J. (2011). Review of “Learning About Teaching: Initial Findings from the Measures of Effective Teaching Project.” Boulder, CO: National Education Policy Center. Retrieved [date] from http://nepc.colorado.edu/thinktank/review-learning-about-teaching. Rothstein, J. (2009). Student sorting and bias in value-added estimation: Selection on observables and unobservables. Education Finance and Policy, 4(4), 537–571. Rothstein, J. (2010). Teacher Quality in Educational Production: Tracking, Decay, and Student Achievement. Quarterly Journal of Economics, 125(1), 175–214. Sanders, W. L., Saxton, A. M., & Horn, S. P. (1997). The Tennessee Value-Added Assessment System: A quantitative outcomes-based approach to educational assessment. In J. Millman (Ed.), Grading teachers, grading schools: Is student achievement a valid measure? (pp. 137-162). Thousand Oaks, CA: Corwin Press. Sanders, William L., Rivers, June C., 1996. Cumulative and residual effects of teachers on future student academic achievement. Knoxville: University of Tennessee Value- Added Research and Assessment Center. Sass, T.R. (2008) The Stability of Value-Added Measures of Teacher Quality and Implications for Teacher Compensation Policy. Urban Institute http://www.urban.org/UploadedPDF/1001266_stabilityofvalue.pdf McCaffrey, D.F., Sass, T.R., Lockwood, J.R., Mihaly, K. (2009) The Intertemporal Variability of Teacher Effect Estimates. Education Finance and Policy 4 (4) 572-606 McCaffrey, D.F., Lockwood, J.R. (2011) Missing Data in Value Added Modeling of Teacher Effects. Annals of Applied Statistics 5 (2A) 773-797 Reardon, S. F. & Raudenbush, S. W. (2009). Assumptions of value-added models for estimating school effects. Education Finance and Policy, 4(4), 492–519. Rubin, D. B., Stuart, E. A., and Zanutto, E. L. (2004). A potential outcomes view of value-added assessment in education. Journal of Educational and Behavioral Statistics, 29(1):103–116. Schochet, P.Z., Chiang, H.S. (2010) Error Rates in Measuring Teacher and School Performance Based on Student Test Score Gains. Institute for Education Sciences, U.S. Department of Education. http://ies.ed.gov/ncee/pubs/20104004/pdf/20104004.pdf.