Showing posts with label standardized tests. Show all posts
Showing posts with label standardized tests. Show all posts

Tuesday, November 4, 2014

The problem with tests that are not standardized

This was written by Alfie Kohn who writes and speaks on parenting and education. Kohn tweets here and his website is here. This post was originally found here.

by Alfie Kohn

I’m baffled by the number of educators who are adamantly opposed to standardized testing yet raise no objection to other practices that share important features with such testing.

For starters, consider those lists of specific, prescriptive curriculum standards to which the tests are yoked. Here we find the same top-down control and one-size-fits-all mentality that animate standardized testing. Yet from the early days of the “accountability” movement right down to current efforts to impose the Gates-funded Common Core from coast to coast, an awful lot of people give the standards (and the whole idea of uniform standards) a pass while frowning only at the exams used to enforce them.[1]

Example #2: Elaborate rubrics used to judge students’ performance represent another form of standardized assessment that’s rarely recognized as such. The point is to break down something, such as a piece of writing, into its parts so that teachers, and sometimes the students themselves, can rate each of them, the premise being that it’s both possible and desirable for all readers to arrive at the same number for each criterion. Rubrics are borne of a demand to quantify and an impulse to simplify. One result, argues Maja Wilson, is that “the standardization of the rubric produces standardized writers.”[2] But, again, even many teachers who are outraged by standardized tests don’t blink when standardization is smuggled in through the back door. Some insist, against all evidence to the contrary, that there’s no problem as long as one uses a good rubric.

It’s my third example, though, on which I’d like to linger. When teachers test their students, the details of those tests will differ from one classroom to the next, which means these assessments by definition are not standardized and can’t be used to compare students across schools or states. But they’re still tests, and as a result they’re still limited and limiting.

As with rubrics (and grades), there’s a reflexive tendency to insist that we just need better tests, or that we ought to just modify the way they’re administered (for example, by allowing students to retake them). And, yes, it’s certainly true that some are worse than others. Multiple-choice tests are uniquely flawed as assessments for exactly the same reason that multiple-choice standardized tests are: They’re meant to trick students who understand the concepts into picking the wrong answer, and they don’t allow kids to generate, or even explain, their responses. Multiple-choice exams can be clever but, as test designer Roger Farr of Indiana University ultimately concluded, there is no way “to build a multiple choice question that allows students to show what they can do with what they know.”

We can also concede that some reasons for giving tests are more problematic than others. There’s a difference between using them to figure out who needs help — or, for more thoughtful teachers, what aspects of their own instruction may have been ineffective — and using them to compel students to pay attention and complete their assignments. In the latter case, a test is employed to pressure kids to do what they have little interest in doing. Rather than address possible deficiencies in one’s curriculum or pedagogy (say, the exclusion of students from any role in making decisions about what they’ll learn), one need only sound a warning about an upcoming test — or, in an even more blatant exercise of power, surprise students with a pop quiz — to elicit compliance.

Even allowing for variation in the design of the tests and the motives of the testers, however, the bottom line is that these instruments are typically more about measuring the number of facts that have been crammed into students’ short-term memories than they are about assessing understanding.[3] Tests, including those that involve essays, are part of a traditional model of instruction in which information is transmitted tostudents (by means of lectures and textbooks) so that it can be disgorged later on command. That’s why it’s so disconcerting to find teachers who are proud of their student-centered approach to instruction, who embrace active and interactive forms of learning, yet continue to rely on tests as the primary, or even sole, form of assessment in their classrooms.

While some of their questions may require problem-solving skills, tests, per se, are artificial pencil-and-paper exercises that measure how much students remember and how good they are at the discrete skill of taking tests. That’s how it’s possible for a student to be a talented thinker and yet score poorly. Most teachers can, without hesitation, name several such students in their classes when the exams are designed by Pearson or ETS, but may fail to see that the same thing applies in the case of performance on tests they design themselves.

Not only do tests assess the intellectual proficiencies that matter least, however — they also have the potential to alter students’ goals and the way they approach learning. The more you’re led to focus on what you’re going to have to know for a test, the less likely you are to plunge into a story or engage fully with the design of a project or experiment. And intellectual immersion can be all but smothered if those tests are given, or even talked about, frequently. Learning in order to pass a test is qualitatively different from learning for its own sake.[4]

***

Many years ago, the eminent University of Chicago educator Philip Jackson interviewed 50 teachers who had been identified as exceptional at their craft. Among his findings was a consistent lack of emphasis on testing, if not a deliberate decision to minimize the practice, on the part of these teachers.[5]

The first reason for this, I think, is that exemplary educators understand that tests are not a particularly useful form of assessment. Second, though, these teachers learned at some point that they didn’t need tests. The most impressive classrooms and curricula are designed to help the teacher know as much as possible about how students are making sense of things. When kids are engaged in meaningful, active learning — for example, designing extended, interdisciplinary projects — teachers who watch and listen as those projects are being planned and carried out have access to, and actively interpret, a continuous stream of information about what each student is able to do and where he or she requires help. It would be superfluous to give students a test after the learning is done. We might even say that the more a teacher is inclined to use a test to gauge student progress, the more that tells us something is wrong — perhaps with the extent of the teacher’s informal and informed observation, perhaps with the quality of the tasks, perhaps with the whole model of learning. If, for example, the teacher favors direct instruction, he or she probably won’t have much idea what’s going on in the students’ minds. That will lead naturally to the conclusion that a test is “necessary” to gauge how they’re doing.[6]

Assessment literally means to sit beside, and that’s just what our most thoughtful educators urge us to do. Yetta Goodman coined the compound noun “kidwatching” to describe reading with each child to gauge his or her proficiency. Marilyn Burns insists that one-on-one conversations tell us far more about students’ mathematical understanding than a test ever could — since all wrong answers aren’t alike. Of course this assumes that we’re really interested in kids’ understanding, not merely their level of phonemic awareness or ability to apply an algorithm. The less ambitious one’s educational goals, the more likely that a test will suffice — and that the wordstesting and assessing will be used interchangeably.

One can fill a bookshelf with accounts of other forms of authentic assessment: portfolios, culminating projects, performance assessments, and what the late Ted Sizer called “exhibitions of mastery”: opportunities for students to demonstrate their proficiency not by recalling facts on demand but by doing something: constructing and conducting (and explaining the results of) an experiment, creating a restaurant menu in a foreign language, turning a story into a play. In other words, when some form of evaluation is desired after, rather than during, the learning, tests stillaren’t necessary or even particularly helpful. They needn’t be used for “summative,” let alone for “formative,” assessment.

Many of us rail against standardized tests not only because of the harmful uses to which they’re put but because they’re imposed on us. It’s more unsettling to acknowledge that the tests we come up with ourselves can also be damaging. The good news is that far superior alternatives are available.


NOTES

1. See my essay “Beware of the Standards, Not Just the Tests,”Education Week, September 26, 2001. This phenomenon is even more pronounced in Canada. Its education system is completely decentralized; each province controls its own policies. Despite the considerable variation in the amount of testing from one to the next, however, all of the provinces have very specific grade-by-grade curricula that every teacher is expected to teach. Objections to this level of control, with the concomitant diminution of autonomy for teachers, are rarely heard — even in provinces where there is outspoken resistance to testing.

2. Maja Wilson, Rethinking Rubrics in Writing Assessment(Heinemann, 2006), p. 39.

3. A spate of recent studies that attracted considerable attention in the popular press argues that frequent tests (including self-tests) are more effective than other forms of studying. But the outcome measure in these studies is almost always limited to the number of facts that are correctly recalled on later tests. Rather than offering an argument in favor of conventional assessment, these experiments actually illuminate how words like “learning” and “achievement” — as used by researchers and journalists alike — often mean little more than the successful, and presumably temporary, process of memorizing facts. For a close look at one such study, see this essay.

4. I recently made this point — about how the anticipation of being tested can distract students from engaging with ideas — in a Twitter post that was retweeted more than 400 times. This degree of popularity led me to suspect I had been misunderstood. I followed up with a clarification that all tests have this effect, not just standardized tests. The retweet rate dropped off by 90 percent.


5. Philip W. Jackson, Life in Classrooms (Teachers College Press, 1968/1990).

6. Frank Smith once wrote, “A teacher who cannot tell without a test whether a student is learning should not be in the classroom.” I see what he means, but his formulation strikes me as a bit harsh. Teachers need help to learn how to assess without tests, and they need support and encouragement to eliminate a practice that is still used by most of their colleagues and widely expected by administrators, parents, and the students themselves. Moreover, the barrier to gauging how successfully students are learning often lies not with the teacher but with features of the school structure, such as classes that are too large or periods that are too short. That’s an argument for organizing to change these problematic policies, not for continuing to test.

Thursday, August 23, 2012

What do standardized test scores tell us?

Today I am going to continue my critique of the Fraser Institute's Report Card on Alberta's High Schools for 2011.

Consider this chart that I created based on information from the Fraser Report:


Here are some interesting details:
  • 5 out of the top 20 schools have reported 0% special needs with the highest being 19%. Every single school in the bottom 20 reported a special needs population with the least being 4.9% and the most being 100%.
  • 12 out of the top 20 schools have an average parent income over $100,000 and 6 of them were over $200,000. In the bottom 20, not one school has an average parent income over $100,000 while half are below $60,000.
  • There are outliers. Bawlf is the only school in the top 20 with an average parent income below $50,000, and there are three schools in the bottom 20 who have an average parent income over $90,000; however, two of those three schools report that over 20% of their population is special needs.
Those in favor of ranking schools via their standardized test scores like to say that it provides parents with the information they need to choose a school for their children. At first glance this looks like it makes a lot of sense -- many people see standardized test scores as the public's window into the quality of our schools. But what if standardized test scores aren't telling us what we think they are telling us? What if standardized test scores tell us less about in-school factors and more about out-of-school factors? In fact, this is exactly the case. Socio-economic status is by far the strongest predictor of student performance on standardized tests.

In Alfie Kohn's book The Case Against Standardized Testing, Kohn explains what standardized testing  really tells us:
The main thing they tell us is how big the students' houses are. Research has repeatedly found that the amount of poverty in the communities where schools are located, along with other variables having nothing to do with what happens in classrooms, accounts for the great majority of the difference in test scores from one area to the next. To that extent, tests are simply not a valid measure of school effectiveness. (Indeed, one educator suggested that we could save everyone a lot of time and money by eliminating standardized tests and just asking a single question: "How much money does your mom make? ... OK, you're on the bottom.") Only someone ignorant or dishonest would present a ranking of schools' test results as though it told us about the quality of teaching that went on in those schools when, in fact, it primarily tells us about socio-economic status and available resources. Of course, knowing what really determines the score makes it impossible to defend the practice of using them as the basis for high-stakes decisions.
When some hear the argument that poverty matters, they like to declare that poverty isn't destiny and that socio-economic status isn't everything. Some will say that within a given school, a group of students of the same status will have variations in the scores. To this Kohn replies:
Sure. And among people who smoke three packs of cigarettes a day, there are going to be variations in lung cancer rates. but that doesn't change the fact that smoking is the factor most powerfully associated with lung cancer.
In Edmonton, Todd Rogers from the University of Alberta conducted research on the variables that affect student performance on Alberta's Provincial Achievement Tests. Rogers found that "by far, the strongest predictor of student performance on achievement tests is socio-economic status (SES)."

In Calgary, Hugh Lytton and Michael Pyryt came to similar conclusions: "Social class factors explain about 45 per cent of the variation in achievement test results. The correlation between income level and achievement test scores is very strong."

Both studies were summarized by the Alberta Teachers' Association News in 1997.

In his book Measuring Up: What Educational Testing Really Tells Us, Daniel Koretz writes about a friend of his that ran a large testing program who often received calls from parents asking him for how they could use standardized test scores to select the best school for their children. Often these phone calls were disappointing for parents because they wanted a method that was simple and free from ambiguity and complexity. Koretz's friend shared an example of when a parent simply wanted a list of the schools with the highest test scores. After trying to explain that test scores shouldn't really be used that way, Koretz's friend lost his patience and told the parent, "If all you want is high average test-scores, tell your realtor that you want to buy into the highest-income neighbourhood you can manage. That will buy you the highest average score you can afford."

Real accountability is about transparency but there is nothing transparent about how standardized testing reduces learning to the convenience of a number or a rank. We are mistakenly led to believe that standardized test scores tell us about school quality when really it is an echo-chamber for affluence and opportunity. Mark Twain may have summarized all this up nicely when he said:
It ain't what you don't know that gets you in trouble. It's what you know for sure that just ain't so.

Wednesday, March 17, 2010

Standardized Test Scores: At best unhelpful and at worst harmful

A large body of research shows that the standardized test scores are effective predictors... of the size of homes that surround a school! Studies have shown that 50%-90% of the factors that influence standardized test score results are external from the learning that occurs in the classroom - effectively making standardized test scores a great measurement for the affluence of a school's population.

Because focusing on standardized tests end up measuring what matters least, they actually end up encouraging the worst kinds of teaching and learning environments. And so there are two feasible reactions to higher standardized test scores. One is "so what!". This implies an understanding that the successes and failures illustrated by rising and lowering test scores says nothing about the quality of education a school provides its students. The second reaction to high scores is "uh-oh!" This reaction implies a kind of deep concern for what kinds of real learning the school had to sacrifice in order to achieve these higher scores.

For more on Joe Bower's views on this topic, take a look at these blog posts:


Standardized Testing is Dumbing Down Our Schools
Multiple Choice Tests Suck
Accountability and George W. Bush
The People's Republic of Standardization
High Stakes Testing's Kryptonite
Bastardized Accountability
 
For more information about booking Joe Bower for a lecture or workshop, please contact by e-mail: joe.bower.teacher@gmail.com


Return to Joe's list of presentations

Tuesday, January 19, 2010

Peoples' Republic of Standardization


Yong Zhao's recently released book Catching Up or Leading the Way is a must read for educators and policy makers who want to see where our current high stakes testing regimes will take us. Zhao does a masterful job of showing how China has long had an obsession with standardized testing. As far back as AD 605, the Sui Dynasty instituted a Civil Exam called the keju. It was a high stakes gateway to the ruling class that, during its 1,300 year history, proved to be one of the only ways of gaining social promotion. The keju’s importance has risen to astronomical heights – so much so that many have come to see it as China’s fifth grand invention after the compass, gun powder, paper and movable type. Yong Zhao shows the obsessive importance of the keju in his book Catching Up or Leading the Way:



Passing the exams was considered one of the most important accomplishments in a person’s life. Indeed, the two happiest moments for an individual in China were said to be the wedding night and seeing one’s name on the list of people who passed the keju. It was the pursuit of a lifetime for many. With no age limit or limit on how many times one could try, historical records show that some persisted in taking the tests into their 70s. The most famous case took place in 1699, when an individual took the test at the age of 102. Chinese literature has many stories – romantic, sad, happy, and bizarre – about individuals who studied for the keju or about their long journey’s to the sites where the keju was held.


Zhao goes on to explain that even though the keju was never an education system, but a political one, it has infected the Chinese classrooms with its obsessive preparation and narrow focus. For 1300 years, teaching and learning in China has been hijacked by standardized testing.

Because the Confucian classics were the core content of the keju, rote memorization became the most popular kind of learning. Most test questions involved reading an excerpt from the original classics and identifying the missing phrases.

Zhao does a remarkable job of giving a short history lesson that illuminates the long term effects of this kind of education. China began using the movable type printing technique 400 years before Guttenberg. They used the magnetic compass perhaps as much as a century before the rest of the world. And of course, the Chinese were the inventors of gun powder. At one time, China showed some very strong initiative and creativity, but that was centruies ago. Something seemed to happen around the 15th century - China ceased to be cutting edge.

There are probably many answers to why China has suffered from this lack of innovation, but Zhao makes a compelling argument that the keju is certainly a prime suspect. Because the keju placed so much emphasis on such a narrow list of skills, a kind of 'talent cleansing' occurred. People with an alternate skill-set were discriminated and discouraged from pursuing their interests in science and technology - those individuals and China as a whole have suffered the long-term consequences.

Traditional China's keju can be found in its modern day reincarnation the gaokao. The gaokao is like the American's Standard Aptitude Test (SAT) on steroids and ecstasy. College admissions in China are solely and entireley dependent on performing well on the gaokao.

Together the keju and gaokao have contributed to a widely recognized problem in Chinese education: gaofen dineng which literally means high scores but low ability. The sheer number of stories and examples of how wide spread gaofen dineng has become in China has lead many Chinese to actually associate gaofen dineng with their entire education system.

In Canada and the United States, most recognize the term valedictorian as a title 'earned' by those top performers in their graduating class. In ancient times, China bestowed the top performer of the keju with the title of zhuangyuan. Today, zhuangyuans, those who achieve the highest scores on the gaokao, become instant celebrities. They have their '15 minutes in the spotlight'. The problem is that the research is showing that their importance and success isn't lasting much longer than that 15 minutes. Zhuangyuans who become distinguished leaders, accomplished engineers or creative entrepeneurs are the exception and not the rule. For the most part, these zhuangyuans excel on the tests and disapear into obscurity - leading many to question why the tests were so important in the first place.

There is a real paradox revolving around this whole China story. Canada and the United States are looking across the Pacific Ocean, and we are envious. We aspire to be more like the Chinese - we want more standardization and more accountability through high-stakes testing. And yet the Chinese are looking across the Pacific and wish they were more like us. They are envious of our creativity, ingenuity and individualism. There is a real 'grass is greener over there' scenario going on here. The scary realization we need to make here is that they are right - we have it right, but we are squandering more and more of it every time we longingly look east.

Yong Zhao summarizes very nicely in his book Catching Up or Leading the Way that in the end, it makes very little sense for developed countries like the US and Canada to fret over 'catching up' to developing nations like China. We got where we are by leading the way, and for some reason we are now turning around to follow those who are trying to catch ujp with us.