Principles of language assessment meliputi: practicality, reliability, validity, authenticity dan washback.
Halaman 19-28
An effective test is practical.This means that it
- is not excessively expensive,
- stays within appropriate time constraints,
- is relatively easy to administer, and
- has a scoring/evaluation procedure that is specific and time-efficient.
A test that is prohibitively expensive is impractical. A test of language proficiency that takes a student five hours to complete is impractical — it consumes more time (and money) than necessary to accomplish its objective. A test that requires individual one-on-one proctoring is impractical for a group of several test-takers and only a handful of examiners. A test that takes a few minutes for a student to take and several hours for an examiner to evaluate is impractical for most classroom situations. A test that can be scored only by computer is impractical if the test takes place a thousand miles away from the nearest computer. The value and quality of a test sometimes hinge on such nitty-gritty, practical considerations.
The students arrived, test booklets were distributed, and directions were given. The proctor started the tape Soon students began to look puzzled. By the time the tenth item played, everyone looked bewildered Finally, the proctor checked a rest booklet and was horrified to discover that the wrong tape was playing; it was a tape for another form of the same test! Now what? She decided to randomly select a short passage from a textbook that was in the room and give the students a dicta- tion. The students responded reasonably well. The next 80 non-tape-based items proceeded without incident, and the students handed in their score sheets and dic- tation papers.
When the red-faced administrator and the proctor got togerher later to score the tests, they faced the problem of how to score the dictation-a more subjective process than some other forms of assessment (see Chapter 6). After a lengthy exchange, the two established a point system, but after the first few papers had been scored, it was clear that the point system needed rcvision. That meant going back to the first papers to make sure the new system was followed.
The two faculty members had barely begun to score the 80 multiple-choice items when students began returning to the office to receive their placements. Students were told to come back the next morning for their results. Later that evening, having combined dictation scores and the 80-item multiple-choice scores, the two frustrated examiners finally arrived at placements for all students.
It's easy to see what went wrong here. While the listening comprehension sec- tion of the test was apparently highly practical, the administrator had failed to check the materials ahead of time (which, as you will see below, is a factor that touches on unreliability as well). Then, they established a scoring procedure that did not fit into the time constraints. In classroom-based testing, time is almost always a crucial prac- ticality factor for busy teachers with too few hours in the day!
B. RELIABILITY
A reliabel test is consistent and dependable. If you give the same test the same test to the same student or matche students on two different occasions,the test should yield similar results. The issue of reability of a test may best be addressed by considering a number of factors that may contribute to the unreliability of a test.consider the following possibilities ( adapted from mousavi,2002,p.804): fluctuations in the student,im scoring,in test administration,and in the test itself.
1. Student-related reliability
The most common learner-related issue in reliability is caused by temporary illness. fatigue, a "bad day," anxiety, and other physical or psychological factors, which may make an "observed"score deviate from one's "true" score. Also included in this cate- gory are such factors as a test-taker's "test-wiseness" or strategies for efficient test taking (Mousavi, 2002, p. 804).
2. Rater Reliability
Human error, subjectivity, and bias may enter into the scoring process. Inter-rater reliability occurs when two or more scorers yield inconsistent scores of the same test, possibly for lack of attention to scoring criteria, inexperience, inattention, or even preconceived biases. In the story above about the placement test, the initial scoring plan for the dictations was found to be unreliable-that is, the two scorers were not applying the same standards.
rater-reliability issues are not limited to contexts where two or more scorers are involved. Intra-rater reliability is a common occurrence for classroom teachers because of unclear scoring criteria, fatigue, bias toward particular "good" and "bad" students, or simple carelessness. When I am faced with up to 40 tests to grade in only a week, I know that the standards I apply-however subliminally-to the first few tests will be different from those I apply to the last few. I may be "easier" or "harder" on those first few papers or I may get tired, and the result may be an inconsistent evaluation across all tests. One solution to such intra-rater unreliability is to read through about half of the tests before rendering any final scores or grades, then to recycle back through the whole set of tests to ensure an even-handed judg- ment. In tests of writing skills, rater reliability is particularly hard to achieve since writing proficiency involves numerous traits that are difficult to define. The carcful specification of an analytical scoring instrument, however, can increase rater relia- bility (J. D. Brown, 1991).
3.Test Administration Reability
Unreliability may also result from uhe conditions in which the test is administered. I once witnessed the administration of a test of aural comprehension in which a tape recorder played items for comprehension, but because of street noise outside the building, students sitting next to windows could not hear the tape accurately. This was a clear case of unreliability caused by the conditions of the test administration. Other sources of unreliability are found in photocopying variations, the amount of light in different parts of the room, variations in temperature, and even the condi- tion of desks and chairs.
Test Reluability
Sometimes the nature of the test itself can cause measurement errors. If a test is too long, test-takers may become fatigued by the time they reach the later items and hastily respond incorrectly, Timed tests may discriminate against students who do not perform well on a test with a time limit, We all know people (and you may be included in this category!) who "know" the course material perfectly but who are adversely affected by the presence of a clock ticking away. Poorly written test items (that are ambiguous or that have more than one correct answer) may be a further source of test unreliability.
C. VALIDITY
By far the most complex criterion of an effective test-and arguably the most impor- tant principle-is validity, "the extent to which inferences made from assessment results are appropriate, meaningful, and useful in terms of the purpose of the assess- ment" (Gronlund, 1998, p. 226). A valid test of reading ability actually measures reading ability-not 20/20 vision, nor previous knowledge in a subject, nor some other variable of questionable relevance. To measure writing ability, one might ask students to write as many words as they can in 15 minutes, then simply count the words for the final score. Such a test would be easy to administer (practical), and the scoring quite dependable (reliable).
1.Content – Related Evidence
If a test actually samples the subject matter about which conclusions are to be drawn, and if it requires the test-taker to perform the behavior that is being mea- sured, it can claim content-related evidence of validity, often popularly referred to as content validity (e.g., Mousavi, 2002; Hughes, 2003).
There are a few cases of highly specialized and sophisticated testing instru- ments that may have questionable content-related evidence of validity. It is possible to contend, for example, that standard language proficiency tests, with their context- reduced, academically oriented language and limited stretches of discourse, lack content validity since they do not require the full spectrum of communicative per- formance on the part of the learner (see Bachman, 1990, for a full discussion).
Direct testing involves the test taker in actu ally performing the target task.in andirect test learners are not performing the task itself but rather a task that is related in some way. For example, if you intend to test learners' oral production of syllable stress and your test task is to have learners mark (with written accent marks) stressed syllables in a list of written words, you could, with a stretch of logic, argue that you are indirectly testing their oral pro- duction A direct test of syllable production would have to require that students actually produce target words orally.
The most feasible rule of thumb for achieving content validity in classroom assessment is to test performance directly. Consider, for example, a listening/ speaking class that is doing a unit on greetings and exchanges that includes dis- course for asking for personal information (name, address, hobbies, etc.) with some form-focus on the verb to be, personal pronouns, and question formation. The test on that unit should include all of the above discourse and grammatical elements and involve students in the actual performance of listening and speaking.
2.Criterion-Related Evidence
A second form of evidence of the validity of a test may be found in what is called criterion-related evidence, also referred to as criterion-related validity, or the extent to which the "criterion" of the test has actually been reached. You will recall that in Chapter 1 it was noted that most classroom-based assessment with teacher- designed tests fits the concept of criterion-referenced assessment.
In the case of teacher-made classroom assessments, criterion-related evidence is best demonstrated through a comparison of results of an assessment with results of some other measure of the same criterion. For example, in a course unit whose objective is for students to be able to orally produce voiced and voiceless stops in all possible phonetic environments, the results of one teacher's unit test might be compared with an independent assessment-possibly a commercially produced test in a textbook-of the same phonemic proficiency. A classroom test designed to assess mastery of a point of grammar in communicative use will have criterion validity if test scores are corroborated either by observed subsequent behavior or by other communicative measures of the grammar point in question.
Criterion-related evidence usually falls into one of two categories: concurrent and predictive validity. A test has concurrent validity if its results are supported by other concurrent performance beyond the assessment itself. For example, the validity of a high score on the final exam of a foreign language course will be substantiated by actual proficiency in the language. The predictive validity of an assessment becomes important in the case of placement tests, admissions assessment batteries, language aptitude tests, and the like. The assessment criterion in such cases is not to measure concurrent ability but to assess (and predict) a test-taker's likelihood of future success.
3. Construct Related Evidence
A construct is any theory, hypothesis, or model that attempts to explain observed phenomena in our universe of perceptions. Constructs may or may not be directly or enmpirically meassured-their verification often requires infer- ential data."Proficiency" and "communicative competence" are linguistic constructs; self-esteenn" and "motivation" are psychological constructs. Virtually every issue in language learning and teaching involves theoretical constructs. In the field of assess- ment, construct validity asks, "Does this test actually tap into the theoretical con- struct as it has been defined?" Tests are, in a manner of speaking, operational definitions of constructs in that they operationalize the entity that is being mea- sured (see Davidson, Hudson, & Lynch, 1985).
Construct validity is a major issue in validating large-scale standardized tests of proficiency: Because such tests must, for economic reasons, adhere to the principle of practicality, and because they must sample a limited number of domains of lan- guage, they may not be able to contain all the content of a particular field or skill.
The TOEFL's omission of oral production content, however, is osten- sibly justified by research that has shown positive correlations between oral produc- tion and the behaviors (listening, reading, grammaticality detection. and writing) actually sampled on the TOEFL (see Duran et al., 1985), Because of the crucial need to offer a financially affordable proficiency test and the high cost of administering and scoring oral production tests, the omission of oral content from the TOEFL has been justified as an economic necessity. (Note: As this book goes to press, oral pro- duction tasks are being included in the TOEFL, largely stemming from the demands of the professional community for authenticity and content validity.)
4. Consquentil validity
As well as the above three widely accepted forms of evidence that may be intro- duced to support the validity of an assessment, two other categories may be of some interest and utility in your owu quest for validating classroom tests. Messick (1989), Gronlund (1998), McNamara (2000), and Brindley (2001), among others, underscore the potential importance of the consequences of using an assessment.
As high-stakes assessment has gained ground in the last two decades, onc aspect of consequential validity has drawn special attention: the effect of test prepa- ration courses and manuals on performance. McNamara (2000, p. 54) cautions against test results that may reflect socioeconomic conditions such as opportunities for coaching that are "differentially available to the students being assessed (for example, because only some families can afford coaching, or because children with more highly educated parents get help from their parents)." The social conse- quences of large-scale, high-stakes assessment are discussed in Chapter 6.
Another important consequence of a test falls into the category of wasbback, to be more fully discussed below. Gronlund (1998, pp. 209-210) encourages teachers to consider the cffect of assessments on students' motivation, subsequent performance in a course, independent learning, study habits, and attitude toward school work.
5. Face Validity
An important facet of consequential validity is the extent to which "students view the assessment as fair, relevant. and useful for improving learning" (Gronlund, 1998, p. 210), or what is popularly known as face validity. "Face validity refers to the degree to wvhich a test looks right, and appears to measure the knowledge or abili- ties it claims to measure, based on the subjective judgment of the examinees who take it, the administrative personnel who decide on its use, and other psychometri- cally unsophisticated observers" (Mousavi, 2002, p. 244).
Face validity means that the students perceive the test to be valid.
- A well-constructed, expected format with familiar tasks,
- A test that is clearly doable within the allotted time limit,
- Items that are clear and uncomplicated,
- Directions that are crystal clear,
- Tasks that relate to their course work (content validity), and
- A difficulty level that presents a reasonable challenge.
6. AUTHENTICITY
A fourth major principle of language testing is authenticity, a concept that is a little slippery to define, especially within the art and science of evaluating and designing tests. Bachman and Palmer (1996, p. 23) define authenticity as "the degree of corre- spondence of the characteristics of a given language test task to the features of a target language task," and then suggest an agenda for identifying those target lan- guage tasks and for transforming them into valid test items.
In a test, authenticity may be present in the following ways:
- The language in the test is as natural as possible.
- Items are contextualized rather than isolated. Topics are meaningful (relevant, interesting) for the learner.
- Some thematic organization to items is provided, such as through a story line or episode.
- Tasks represent, or closely approximate, real-world tasks.
The authenticity of test tasks in recent years has increased noticeably. Two or three decades ago, unconnected, boring, contrived items were accepted as a neces- sary component of testing. Things have changed. It was once assumed that large- scale testing could not include performance of the productive skills and stay within budgetary constraints, but now many such tests offer speaking and writing compo- nents. Reading passages are selected from real-world sources that test-takers are likely to have encountered or will encounter. Listening comprehension sections fea- ture natural language with hesitations, white noise, and interruptions. More and more tests offer items that are "episodic" in that thcy are sequenced to form mean- ingful units, paragraphs, or stories.
D. Weshback
A facet of consequential validity, discussed above, is "the effect of testing on teach- ing and learning" (Hughes, 2003, p. 1), otherwise known among language-testing speciàlists as washback. In large-scale assessment, washback generally refers to the effects the tests have on instruction in terms of how students prepare for tuhe test.
The challenge to teachers is to create classroom tests that serve as learning devices through which washback is achieved, Students' incorrect responses Can become windows of insight into further work. Their correct responses need to be praised, especially when they represent accomplishments in a student's inter- language. Teachers can suggest strategies for success as part of their "coaching" role. Washback enhances a number of basic principles of language acquisition: intrinsie motivation, autonomy, self-confidence, language ego, interlanguage, and strategie investment, among others. (See PLLT and TBP for an explanation of these principles.)
Ahotier vicwpoint on washback is achieved by a quick consideration of dilfer- ences between formative and summative tests, mentioned in Chapter 1. Formative tests, by definition, provide washback in the form of information to the learner on progress toward goals. But teachers might be tempted to feel that summative tests. wich provide assessment at the end of a course or program, do not need to offer much in the way of washback.
In my courses I never give a final examination as the last scheduled classroom session. I always administer a final exam during the penultimate session, then com- plete the evaluation of the exams in order to return them to students during the last Class. At this time, the students receive scores, grades, and comments on their work. and I spend some of the class session addressing material on which the students were not completely clear. My summative assessment is thereby enhanced by some beneficial washback that is usually not expected of final examinations.
Finally, washback also implies that students have ready access to you to disciss the feedback and evaluation you have given. White you almost ceriainly have known teachers with whom you wouldn't dare argue about a grade, an interactive, cooper aive, collaborative classroom nevertheless can promote an atmosphere of diaogue Detween students and teachers regarding evaluative judgments. For learning to con- tinuc, students nced to have a chance to feed back on your feedback, to seek clari- fication of any issues that are fuzzy, and to set new and appropriate goaIs Ror themselves for the days and weeks ahead.
References
Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california.
Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california.
Tidak ada komentar:
Posting Komentar