STANDARDS-BASED ASSESSMENT
Halaman 104-115
Halaman 104-115
STANDARDS-BASED
ASSESSMENT
In the previous chapter, you saw that a standardized test is an
assessment instrument for which there are uniform procedures for
administration, design, scoring, and reporting. It is also a procedure that, through
repeated administrations and ongoing research, demonstrates criterion and
construct validity. But a third, and perhaps the most important, element of
standardized testing is the presupposition of an accepted set of standards on
which to base the procedure. This feature of an edu- cational and business
world caught up in a frenzy of standardized measurement is perhaps the most
complex, and is the subject of this chapter.
A history of standardized testing in the United States reveals that during
most of the decades in the middle of the twentieth century, standardized tests
enjoyed a popularirty and growth that was almost unchallenged. Standardized
instruments brought with them convenience, efficiency, and an air of empirical
science. In schools, for example, millions of children could be led into a
room, seated, armed with a lead pencil and a score sheet, and almost instantly
assessed on their achieve ment in subject-matter areas in their curricula.
Standardized test advocates' utopian dream of quickly and cheaply assessing
students across the country soon became a political issue, and would-be office
holders to this day promise to"reform" education with tests, tests,
and more tests.
Toward the end of the twenticth
century, such claims began to be challenged on all fronts (sce Medina &
Neill, 1990; Kohn, 2000), and at the vanguard of those challenges were the
teachers of those millions of children. Teachers saw not only possible inequity
in such tests but a disparity between the content and tasks of the tests and
what they were teaching in their classes.
within each content area being
measured. For example, most departments of education at the state level in the
United States have now specified (or are in the process of spec- ifying) the
appropriate standards (that is, criteria or objectives) for each grade level (kindergarten
to grade 12) and each content area (math, language, sciences, arts).
The construction of such standards
makes possible a concordance between standardized test specifications and the
goals and objectives of educational pro- grams. And so, in the broad domain of
language arts, teachers and educational administrators began the painstaking
process of carefully examining existing cur- ricular goals, conducting needs
assessments among students, and designing appro- priate assessments of those
standards. A subfield of language arts that is of
increasing importance in the United States, with its millions of non-native
users of English, is English as a Second Language (ESL), also known as English
for Speakers of Other Languages (ESOL), English Language Learners (ELLS), and
English Language Development (ELD).
(Note: The once popular term Limited English
Proficient (LEP) has now been discarded because of the negative connotation of
the word limited.)
ELD
STANDARDS
The process of designing and conducting appropriate periodic reviews of
ELD stan- dards involves dozens of curriculum and assessment specialists,
teachers, and researchers (Fields, 2000; Kuhlman, 2001). In creating such
"benchmarks for accountability" (O'Malley & Valdez Pierce, 1996),
there is a tremendous responsi- bility to carry out a comprehensive study of a
number of domains:
- Literally thousands of categories of language ranging from phonology at one end of a continuum to discourse, pragmatics, functional, and sociolinguistic elements at the other end;
- Specification of what ELD students' needs are, at thirteen different grade levels, for succeeding in their academic and social development;
- A consideration of what is a realistic number and scope of standards to be included within a given curriculum;
- A separate set of standards (qualifications, expertise, training) for teacbers to teach ELD students successfully in their classrooms; and
- A thorough analysis of the means aallable to assess student attainment of those standards.
Standards-setting is a global challenge. In many non-English-speaking
coun- tries, English is now a required subject starting as early as the first
grade in some countries and by the seventh grade in virtually every country worldwide.
In Japan and Korea, for example, a "communicative" curriculum in
English is required from third grade onward. Such mandates from ministries of
education require the spec- ification of standards on which to base curricular
objectives, the teachability of which has been met with only limited success in
some areas (Chinen, 2000; Yoshida, 2001; Sakamoto, 2002).
The Listening and Speaking standards for English-language learners
(ELLS) identify a student's competency to understand the English language and
to produce the language orally. Students must be prepared to use English
effectiveły in social and academic settings. Listening and speaking skills
provide one of the most important building blocks for the foundation of second
language acquisition. These skills are essential for developing reading and
writing skills in English; however, to ensure that ELLS acquire proficiency in
English listening, speaking, reading, and writing, it is important that
students receive reading and writing instruction in English while they are
developing fluency in oral English.
ELD
ASSESSMENT
The development of standards obviously implies the responsibility for
corre assessing their attainment. As standards-based education became more
accepted in 1990s, many school systems across the United States found that the
standardized te of past decades were not in line with newly developed
standards. Thus began the in active process not only of developing standards
but also of creating standards-bas assessments. The comprehensive process of
developing such assessment in Calife still continues as curriculum and
assessment specialists design, revise, and vali numerous tests (Morgan &
Kuhlman, 2001; Stack et al., 2002; see also the web http://www.cde.
ca.gov/statetests/celdt/celdt.html).
the California English Language
Development Test (CELDT) was developed. The CELDT is a battery of instruments
designed to assess the attainment of ELD stan- dards across grade levels. (For
reasons of test security, specifications for this test are not available to the
public.)
The process of
administering a comprehensive, valid, and fair assessment of ELD students
continues to be perfected. Stringent budgets within departments of education
worldwide predispose many in decision-making positions to rely on tra- ditional
standardized tests for ELD assessment, but rays of hope lie in the exploration
of more student-centered approaches to learner assessment. Stack, Stack, and
Fern (2002), for example, reported on a portfolio assessment system in the San
Francisco Unified School District called the Language and Literacy Assessment
Rubric (LALAR), in which multiple forms of evidence of students' work are
collected. Teachers observe students year-round and record their observations
on scannable forms. The use of the LALAR system provides useful data on
students' performance at all grade levels for oral production, and for reading
and writing performance in elementary and middle school grades (1-8). Further
research is ongoing for high school levels (grades 9-12).
CASAS
AND SCANS
At the higher levecls of education (colleges,
community colleges, adult schools language schools, and workplace settings),
standards-based assessment systems have also had an enormous impact. The
Comprehensive Adult Student Assessmen System (CASAS), for example, is a program
designed to provide broadly based assessments of ESL curricula across the
United States. The system includes more than 80 standardized assessment
instruments used to place learners in programs diagnose learners' needs,
monitor progress, and certify mastery of functiona basic skills. CASAS
assessment instruments are used to measure functiona reading, writing,
listening, and speaking skills, and higher-order thinking skills CASAS scaled
scores report learners' language ability levels in employment and adult life
skills contexts.
A
similar set of standards compiled by the U. S. Department of Labor, no known as
the Secretary's Commission in Achieving Necessary Skills (SCANS), ou lines
competencies necessary for language in the workplace. The compctencie cover
language functions in terms of
- resources (allocating time, materials, staff, etc.),
- interpersonal skills, teamwork, customer service, etc.
- information processing, evaluating data, organizing files, etc.,
- systems (eg., understanding social and organizational systems), and
- technology use and application.
These five
competencies are acquired and maintained through training in the basic skills
(reading, writing, listening, speaking); thinking skills such as reasoning and
cre ative problem solving: and personal qualities, such as selfesteem and
sociability
TEACHER
STANDARDS
In addition to the movement to create
standards for learning, an equally strong move ment has emerged to design
standards for teaching. Cloud (2001,p. 3) noted that a stu- dent's
"performance [on an assessment] depends on the quality of the
instructional program provided, ... which depends on the quality of
professional development." Kuhlman (2001) emphasized the importance of
teacher standards in three domains:
- linguistics and language development.
- culture and the interrelationship between language and culture.
- planning and managing instruction .
How
to assess whether teachers have met standards remains a complex issue. Can
pedagogical expertise be assessed through a traditional standardized test? In
the first of Kuhlman's domains-linguistics and language development-knowledge
can perhaps be so evaluated, but the cultural and interactive characteristics
of effective teaching are less able to be appropriately assessed in such a
test. TESOL's standards committee advo- cates performance-based assessment of
teachers for the following reasons:
- Teachers can demonstrate the standards in their teaching
- Teaching can be assessed through what teachers do with their learners in their classrooms or virtual classrooms (their performance).
- This performance can be detailed in what are called "indicators": examples of evidence that the teacher can meet a part of a standard.
- The processes used to assess teachers need to draw on complex evidence of performance. In other words, indicators are more than simple "how to" statements.
- Performance-based assessment of the standards is an integrated system. It is neither a checklist nor a series of discrete assessments.
- Each assessmentr within the system has performance criteria against which the performance can be measured.
- Performance criteria identify to what extent the teacher meets the standard. Student learning is at the heart of the teacher's performance.
The standards-based
approach to teaching and assessment presents the prot sion with many
challenges, However thorny those issues are, the social consequen of this
movement cannot be ignored, especially in terms of student assessment.
THE
CONSEQUENCES OF STANDARDS-BASED AND STANDARDIZED TESTING
A couple of decades ago I had the pleasure and
challenge of serving on the TOEI Rescarch Committee. Among other things, it was
a good opportunity to hear so of the "inside" stories about the TOEFL.
One of those stories, as told by Rus Webster (personal communication),
illustrates the high-stakes nature of this globa marketed standardized test.
A
ring of enterprising "business" persons organized a group of pretend
te takers to take the TOEFL in an early time zone on a given day. (In those
days the te were administered everywhere on the same day across a number of
time zones. TOEFL. administrations ended in some East Asian countries as much
as 8 to 14 hr before they began in the United States.)
The
task of cach test -taking "spy" was not to pass the TOEFL, but to
memoriz subset of items, including the stimulus and all of the multiple-choice
options, a immediately upon leaving the exam to telephone those items to the
central or nizers. As the memorized subsections were calied in, a complete form
of the TOI was quickly reconstructed The organizers had employed expert
consultants to g erate the correct response for cach item, thereby recreating
the test items and th correct answers! For an outrageous price of many
thousands of dollars, prearran buyers of the results were given copies of the
test items and correct responses va few hours to spare before entering a test
administration in the Western Hemisph.
The story of how
this underhanded group of entrepreneurs were caught brought to justice is a
long tale of blockbuster spy-novel proportions involving the and, eventually,
international investigators. But the story shows the huge gate-keep role of
tests like the TOEFL and the high price that some were willing to pay to access
to a university in the United States and the visa that accompanied it.
The
widespread global acceptance of standardized tests as valid procedure:
assessing individuals in many walks of life brings with it a set of
consequences fall under the category of consequential validity discussed in
Chapter 2. Som those consequences are positive. Standardized tests offer high
levels of practic and reliability and are often supported by impressive construct
validation stu They are therefore capable of accurately placing tens and
hundreds of thousan test-takers onto a norm-referenced scale with high
reliability ratios (most ran between 80 and 90 percent).
Are
the institutions that produce and utilize highstakes standardized tests jus-
tified in their decisions? An impressive array of research would seem to say
yes. Consider the fact that correlations between TOEFI. scores and academic
perfor- mance in the first year of college are impressively high (Henning &
Cascallar, 1992).
But several nagging, persistent issues emerge
from the arguments about the con- sequences of standardized testing. Consider
the following interrelated questions:
- Should the educational and business world be satisfied with high but not per- fect probabilities of accurately assessing test-takers on standardized instru- ments? In other words, what about the small minority who are not fairly assessed?
- Regardless of construct validation studies and correlation statistics, should fur- ther types of performance be elicited in order to get a more comprehensive picture of the test-taker?
- Does the proliferation of standardized tests throughout a young person's life give rise to test-driven curricula, diverting the attention of students from cre- ative or personal interests and in-depth pursuits?
- Is the standardized test industry in effect promoting a cultural, social, and political agenda that maintains existing power structures by assuring opportu- nity to an elite (wealthy) class of people?
Test Bias
It is no secret that standardized tests
involve a number of types of test bias.That bias comes in many forms: language,
culture, race, gender, and learning styles (Medina & Neill, 1990). The
National Center for Fair and Open Testing, in its bimonthly newsletter Fair
Test, every year offers dozens of instances of claims of test bias from
teachers, parents, students, and legal consultants (see their website:
www.fairtest. org). For example, reading sclections in standardized tests may
use a passage from a literary piece that reflects a middle-class, white,
Anglo-Saxon norm. Lectures used
for
listening stimuli can easily promote a biased sociopolitical view. Consider the
fol- lowing prompt for an essay in "general writing ability"on the
IELTS:
You rent a house
through an agency The heating system has stopped working. You phoned the agency
a week ago, but it has still not been mended. Write a letter to the agency.
Explain the situation and tell them what you want them to do about it.
While this task
favorably illustrates the principle of authenticity, a number of cul- tural and
economic presuppositions are evident in such a prompt, calling into ques- tion
its potential cultural bias,
In an cra when we seek to recognize the
multiple intelligences present within every student (Gardner, 1983, 1999), is
it not likely that standardized tests promote logical-mathematical and
verbal-linguistic intelligences to the virtual exclusion of the other
contextualized, integrative intelligences? Only very recently have
traditionally receptive tests begun to include written and oral production in
their test battery-a positive sign. But is it enough? It is also clear that
many otherwise "smart" people do not perform well on standardized
tests.
Test-Driven
Learning and Teaching
Yet another
consequence of standardized testing is the danger of test-drives learning and
teaching. When students and other test-takers know that one singie measure of
performance will determine their lives, they are less likely to take a pos
itive attitude toward learning. The motives in such a context are almost
exclusivels extrinsic, with little likelihood of stirring intrinsic interests.
Test-driven learning is a worldwide issue. In Japan, Korea, and Taiwan, to name
just a few countries, students approaching their last year of secondary school
focus obsessively on passing the year-end college entrance examination, a major
section of which is English (Kuba 2002).
a reward for
their schools' high performance on the state-mandated grade-level test, the
Florida Comprehensive Achievement Exam (Fair Test, 2000). The effect of this
policy was undue pressure on teachers to make sure their students excelled in
the exam, possibly at the risk of ignoring other objectives in their cur-
ricula. But a further, ultimately more serious effect was to punish schools in
lower-socioeconomic neighborhoods. A teacher in such a school might actually be
a superb teacher, and that teacher's students might make excellent progress through
the school year, but because of the test-driven policy, the teacher would
receive no reward at all.
ETHICAL ISSUES: CRITICAL LANGUAGE TESTING
Unfortunately,
too many policymakers and educators have ignored the complexities of testing
issues and the obvious limitations they should place on standardized test use.
Instead, they have been seduced by the promise of simplicity and objectivity.
The price which has been paid by our schools and our children for their
infatuation with tests is high.
Shohamy (1997, p. 2) further defines the
issue:"Tests represent a social technology deeply embedded in education,
government, and business; as such they provide the mechanism for enforcing
power and control. Tests are most powerful as they are often the single
indicators for determining the future of individuals." Test designers, and
the corporate sociopolitical infrastructure that they represent, have an
obligation to maintain certain standards as specified by their client
educational institutions. These standards bring with them certain ethical
issues surrounding the gate-keeping nature of standardized tests.
- teachers, and learners" (Shohamy, 1997, p. 3). The issues of critical language testing are numerous: Psychometric traditions are challenged by interpretive, individualized proce- dures for predicting success and evaluating ability.
- Test designers have a responsibility to offer multiple modes of performance to account for varying styles and abilities among test-takers.
- Tests are deeply embedded in culture and ideology.
- Test-takers are political subjects in a political context.
These issues are
not new. More than a century ago, British educator E Y. Edgeworth (1888)
challenged the potential inaccuracy of contemporary qualifying examinations for
university entrance. In recent years, the debate has heated up. In 1997, an
entire issue of the journal Language Testing was devoted to questions about
ethics in lan- guage testing.
One of the
problems highlighted by the push for critical language testing is the widespread
conviction, already alluded to above, that carefully constructed standand ized
tests designed by reputable test manufacturers are infallible in their
predictive validity. One standardized test is deemed to be sufficient;
follow-up measures are con sidered to be too costly:
A further problem with our test-oriented
culture lies in the agendas of those who design and those who utilize the
tests. Tests are used in some countries to deny cit- zenship (Shohamy, 1997, p.
10). Tests may by nature be culture-biased and therefore may disenfranchise
members of a nonmainstream value system. Test givers are alway in a position of
power over test-takers and therefore can impose social and political ide
ologies on test-takers through standards of acceptable and unacceptable items.
Tes promote the notion that answers to real-world problems have unambiguous
right wrong answers with no shades of gray. A corollary to the latter is that
tests presume to reflect an appropriate core of common knowledge, such as the competenc
reflected in the standards discussed earlier in this chapter.
Language tests,
some may argue, are less susceptible than general-knowledge tes to such
sociopolitical overtones. The research process that undergirds the TOEFL ge to
great lengths to screen out Western cultural bias, monocultural belief systems,
other potential agendas. Nevertheless, even the process of the selection of
conte alone for the TOEFL involves certain standards that may not be universal,
and the ve fact that the TOEFL. is used as an absolute standard of English
proficiency by most versities does not exonerate this particular standardized
test.
As a language teacher, you might be able to
exercise some influence in ways tests are used and interpreted in your own
milieu. If you are offere variety of choices in standardized tests, you could
choose a test that offers least degree of cultural bias. Better yet, you might
encourage the use of multe measures of performance (varying item types, oral
and written production. other alternatives to traditional assessment) even
though this might cost mi money.
FOR YOUR FURTHER READING
Kohn, Alfie. 2000. The case against
standardized testing. Westport, CT: Heinemann.
Kohn makes a
strong case for the unfairness of the widespread exclusive use of standardized
testing in elementary and secondary schools in the United States, His arguments
may be appropriate for any other test-driven educational institution in any
country.
Hamp-Lyons, Liz. (2001). Ethics,
fairness(es), and developments in language testing.
In Catherine Elder (Ed.), Experimenting with
uncertainty: Essays in bonour of Alan Davies (pp. 222-227).
(Studies in Language Testing #11). Cambridge: Cambridge University Press. The
political and ethical issues in testing are capsulized in this brief article.
The author addresses the question of what "fairness" is and how one
might discern when a test is unfair.
REFRENCES :
Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california.
REFRENCES :
Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california.
Tidak ada komentar:
Posting Komentar