SUMMARY
Halaman 42- 65
Designing classroom language tests
The previous chapters introduced a number of building blocks for designing lan- guage tests. You now have a sese of where tests belong in the larger domain of assessment. You have sorted through differences between formal and informal tests, formative and summative tests, and norm- and criterion-referenced tests. You have traced some of the historical lines of thought in the field of language assessment. You have a sense of major current trends in language assessment, especially the present focus on communicative and process-oriented testing that seeks to transform tests from anguishing ordeals into challenging and intrinsically motivating learning expe- riences, By now, certain foundational principles have entered your vocabulary: prac- ticality, reliability, validity, authenticity, and washback. And you shoukd now possess a few tools with which you can evaluate the effectiveness of a classroom test.
In this chaapter,you will draw on those foundations and tools to begin the process of designing Tests or Revising existing tests.to start that process,you need to ask some critical questions;
1.What is tbe purpose of the test? Why am I creating this test or why was it created by someone else? For an evaluation of overall proficiency? To place students into a course? To measure achievement within a course? Once you have established the major purpose of a test, you can determine its objectives.
2. What are tbe objectives of the test? What specifically am I trying to find out? Establishing appropriate objectives involves a number of issues, ranging from rela- tively simple ones about forms and functions covered in a course unit to much more complex ones about constructs to be operationalized in the test. Incuded here are decisions about what language abilities are to be assessed.
3. How will tbe test specifications reflect botb the purpose and the objec- tives? To evaluate or design a test, you must make sure that the objectives are in- corporated into a structure that appropriately weights the various competencies being assessed. (These first three questions all center, in one way or another, on the principle of validity.)
4. How will the test tasks be selected and the separate items arranged? The tasks that the test-takers must perform need to be practical in the ways defined in the previous chapter. They should also achieve content validity by presenting tasks that mirror those of the course (or segment thereof) being assessed. Further, they should be able to be evaluated reliably by the teacher or scorer. The tasks themselves should strive for authenticity, and the progression of tasks ought to be biased for best performance.
5. What kind of scoring, grading, and/or feedback is expected? Tests vary in the form and function of fecdback, depending on their purpose. For every test, the way results are reported is an important consideration. Under some circumstances a letter grade or a holistic score may be appropriate: other circumstances may require that a teacher offer substantive washback to the learner.
a.TEST TYPES
The first task you will face in designing a test for your students is to determine the purpose for the test. Defining your purpose will help you choose the right kind of test, and it will also help you to focus on the specific objectives of the test. We will look first at two test types that you will probably not have many opportunitics to create as a classroom teacher-language aptitude tests and language proficiency tests-and three types that you will almost certainly need to create-placement tests, diagnostic tests, and achievement tests.
b. Language Aptitude Tests
One type of test-although admittedly not a very common one-predicts a person's success prior to exposure to the second language. A language aptitude test is designed to measure capacity or general ability to learn a forcign language and ulti- mate success in that undertaking. Language aptitude tests are ostensibly designed to apply to the ciassroom learning óf any language. Two standardized aptitude tests have been used in the United States: the Modern Language Aptitude Test (MLAT) (Carroll & Sapon, 1958) and the Pimsleur Language Aptitude Battery (PLAB) (Pimsleur, 1966). Both are English language tests and require students to perform a number of language-related tasks. The MLAT, for example, consists of five different tasks.
Tasks in the Modern Language Aptitude Test.
1. Nmber learning: Examinees must learn a set.
2. Phonetic script: Examinees must learn a set of correspondences between speech sounds and phonetic symbols.
3. Spelling clues: Examinees must read words that are spelled somewhat phonetically, and then select from a list the one word whose meaning is closest to the "disguised" word.
4. Words in sentences: Examinees are given a key word in a sentence and are then asked to select a word in a second sentence that performs the same grammatical function as the key word.
5. Paired associates: Examinees must quickly learn a set of vocabulary words from another language and memorize their English meanings.
C.Proficiency Tests
If your aim is to test global competence in a language, then you are, in conventional terminology, testing proficiency A proficiency test is not limited to any one course, curriculum, or single skill in the language; rather, it tests overall ability. Proficiency tests have traditionally consisted of standardized multiple-choice items on grammar, vocabulary, reading comprehension, and aural comprehension. Sometimes a sample of writing is added, and more recent tests also include oral production performance.
Proficiency tests are almost always summative and norm-referenced. They pro- vide results in the form of a single score (or at best two or three subscores, one for each section of a test), which is a sufficient result for the gate-keeping role they play of accepting or denying someone passage into the next stage of a journey. And because they measure performance against a norm, with equated scores and per- centile ranks taking on paramount importance, they are usually not equiPped to pro- vide diagnostic feedback.
A typical example of a standardized proficiency test is the Test of English as a Foreign Language (TOEFL") produced by the Educational Testing Service. The TOEFL is used by more than a thousand institutions of higher education in the United States as an indicator of a prospective student's ability to undertake academic work in an English-speaking milieu. The TOEFL consists of sections on listening comprehension, structure (or grammatical accuracy), reading comprehemsion, and written expression. The new computer-scored TOEFL announced for 2005 will also include an oral pro- duction component. With the exception of its writing section, the TOEFL (as well as many other large-scale proficiency tests) is machine-scorable for rapid turnaround and cost effectiveness (that is, for reasons of practicality). Research is in progress (Bernstein et al., 2000) to determine, through the technology of speech recognition, if oral production performance can be adequately machine-scored.
A key issue in testing proficiency is how the constructs of language ability are specified. The tasks that test-takers are required to perform must be legitimate sam- ples of English language use in a defined context. Creating these tasks and validating them with of a number of commercially available profi- ciency tests.
d. Placement Tests
Certain proficiency tests can act in the role of placement tests, the purpose of which is to place a student into a particular level or section of a language cur- riculum or school. A placement test usually, but not always, includes a sampling of the material to be covered in the various courses in a curriculum; a student's per- formance on the test should indicate the point at which the student will find mate- mial neirher ron casv nor too difficult but apprensigtely challenging.
The English as a Second Language Placement Test (ESLPT) at San Francisco State University has three parts. In Part I, students read a short article and then write a suminary essay. In Part I, students write a composition in response to an article. Part III is multiple-choice students read an essay and identify grammar errors in it: The first part of the test acts as both a test of reading com- prehension and a test of writing (a summary). The second part requires students to state opinions and to back them up, a task that forms a major component of thewriting courses. Fiunally, proofreading drafts of essays is a useful academic skill, and the exercise in error detection simulates the proofreading process.
Tcachers and adınioistrators in the ESL program at SFSU are satisfied with this test's capacity to discriminate appropriately, and they feel that it is a more authentic test than its multiple-choice, discrete-point. grammar-vocabulary predecessor. The practicality of the ESLPT is relatively low:human evaluators are required for the first two parts, a process more costly in both time and money than running the multiple- choice Part III responses through a pre-programmed scanner. Reliability problems are afso present but are mitigated by conscientious training of all evaluators of the test.
Placement tests come in many varieties: assessing comprehension and produc- tion, responding through written and oral performance, open-ended and limited responses, selection (eg multiple-choice) and gap-filling formats, depending on the narure of a program and its needs. Some programs simply use existing standardized proficiency tests because of their obvious advantage in practicality-cost, speed in scoring, and cfficient reporting of results. Others prefer the performance data avail- able in more open-ended written and/or oral production. The ultimate objective of a placement test is, of course, to correctly place a student into a course or level. Secondary benefits to consider include face validity, diagnostic information on stu- dents performance, and authenticity.
e. Diagnostic Tests
A diagnostic test is designed to diagnose specified aspects of a language. A test in pronunciation, for example, might diagnose the phonological features of English that are difficult for learners and should therefore become part of a curriculum. Usually, such tests offer a checklist of features for the administrator (often the teacher) to use in pinpointing difficulties. A writing diagnostic would elicit a writing sample from students that would allow the teacher to identify those rhetorical and linguistic features on which the course needed to focus special attention.
Diagnistic and placement tests,as we have already implied,may sometimes be indistinguishabel from each other.The san francisco state ESLPT serves dual purposes.Any placrment tests that offers information beyond simply designating a course level may also serve diagnotic purposes.
A typical diagnostic test of oral pruduction wvas created by Clifford Prator (1972) to accompany a manral of English pronunciation. Test-takers are directed to read a 150-word passage while they are tape-recorded. The test administraior then refers to an inventory of phonological items for analyzing a learners production. After multiple listenings, the administrator produces a checklist of crrors in Eive sep- arate categories, each of which has several subcategories. The main categories include.
1. stress and rhythm.
2 intonation.
3. vowels,
4.consonants,and
5. other factors.
An example of subcategories is shown in this list for the first category (stress and rhythm)
a. stress on the wrong syllable (in multi-syllabic words)
b. incorrect sentence stress
c. incorrect division of sentences into thought groups
d. failure to make smooth transitions berween words or syllables.
(Prator, 1972)
Each subcategory is appropriately referenced to a chapter and section of Prator's manual. This information can help teachers make decisions about aspects of English phonology on which to focus. This same information can help a student become aware of errors and encourage the adoption of appropriate compensatory strategies.
F. Achievement Tests
An achievement test is related directly to classroom lessons, units, or even a total curriculum. Achievement tests are (or should be) limited to particular material addressed in a curriculum within a particular time frame and are offered after a course has focused on the objectives in question. Achievement tests can also serve the diagnostic role of indicating what a student needs to contimue to work on in the future, but the primary role of an achievement test is to determinc whether couse objectives have been met-and appropriate knowledge and skills acquired-by the end of a period of instruction.
Achievement tests are often summative because they are administered at the end of a unit or term of study. They also play an important formative role. An effc- tive achievement test will offer washback about the quality of a learner's perfor- mance in subsets of the nit or course. This washback contributes to the formative nature of such tests.
The specifications for an achievement test should be determined by.
• the objectives of the lesson, unit, or course being assessed,
• the relative importance (or weight) assigned to each objective, the tasks employed in classroom lessons during the unit of time,
• practicality issucs, such as the time frame for the test and turnaround time, pur the extent to which the test structure lends itself to formative washback.
Achievement tests cange from five- or ten-minute quizzes to three-hour final exam- inations, with an almost infinite variety of item types and formats. Here is the outline for a midterm examination offered at the high-intermediate level of an intensive English program in the United States. The course focus is on academic reading and writing, the structure of the course and its objectives may be implied from the sections of the test.
g. SOME PRACTICAL STEPS TO TEST CONSTRUCTION
The descriptions of types of tests in the preceding section are intended to help you understand how to answer the first question poscd in this chapter. What is the purpose of the test? It is unlikely that you would be asked to design an aptitude test or a proficiency test, but for the purposes of interpreting those tests, it is important that you understand their nature. However, your opportunities to design placement. diagnostic, and achievement tests-especially the latter-will be plentiful. In the remainder of this chapter, we will explore the four remaining questions posed at the outset, and the focus will be on equipping you with the tools you need to create such classroom-oriented tests.
You may think that every test you devise must be a wonderfully innovative instrument that will garner the accolades of your colleagues and the admiration of your students. Not so. First, new and innovative testing formats take a lot of effort to design and a long time to refine through trial and error. Second, traditional testing techniques can, with a little creativity, conform to the spirit of an interactive, com- municative language curriculum. Your best tack as a new teacher is to work within the guidelines of accepted, known, traditional testing techniques. Slowly, with expe- rience, you can get bolder in your attempts. In that spirit, then, let us consider some practical steps in constructing classroom tests.
h. Assessing Clear, Unambiguous Objectives
In addition to knowing the purpose of the test you're creating, you need to knowv as specifically as possible what it is yoOu want to test. Sometimes teachers give tests simply because it's Friday of the third week of the course, and after hasty glances at the chapter(s) covered during those three weeks, thcy dash off some test items so that students will have something to do during the class. This is no way to approach a test. Instead, begin by taking a careful look at everything that you think your stu- dents should "know" or be able to "do, based on the material that the students are responsible for. In other words, examine the objectives for the unit you are testing.
Remember that every curriculum should have appropriately framed assessable objectives, that is, objectives that are stated in terms of overt performance by stu dents. Thus, an objective that states "Students will learn tag questions or simply names the grammatical focus "Tag questions" is not testable. You don't know whether students should be able to understand them in spoken or written langage, or whether they should be able to produce them orally or in writing. Nor do you know in what context (a conversation? an essay? an acadlemic lecture?) those linguistic forms should be used. Your first task in designing a test, then, is to determine appropriate objectives.
i. Drawing Up Test Specifications.
Test specifications for classroom use can be a simple and practical outline of your test. Cor-lare-stale sundurdized tess joce upte that are intended to be widely distributed and therefore are broadly generalized, test specifications are much more formal and detailed.) In the unit discussed above, your specifications will simply comprise (a) a broad outline of the test, (b) what skills you will test, and (c) what the items will look like. Let's look at the first two in relation to the midterm unit assessment already referred to above.
In the current example that we have bcen analyzing, your revising process is likely to result in at least four changes or additions:
1. In both interview and writing sections, you recognize that a scoring rubric will be essential. For the interview, you decide to create a holistic scale , and for the writing section you devise a simple analytic scale that captures only the objectives you have focused on.
2. In the interview questions, you realize that follow-up questions may be needed for students who give one-word or very short answers.
3. In the listening section, part b, you intend choice "c" as the correct answer, but you realize that choice "d" is also acceptable. You nced an answer that is unam biguously incorrect You shorten it to "d. Around eleven o'clock."You also note that providing the prompts for this section on an audio recording will be logis- tically difficult, and so you opt to read these items to your students.
4. In the writing prompt. you can see how some students would not use the Words so or because, which were in your objectives, so you reword the prompt:"Name one of the characters at the party in the TV sitcom we saw. Then, use the word so at least once and the word because at least once to tel why you liked or didn't like that person
Idcally, you would try out all your tests on students not in your class before actually administering the tests. But in our daily classroom teaching, the tryout phase is almost impossible. Alternatively, you could enlist the aid of a colleague to look over your test. And so you must do what you can to bring to your students an instru- ment that is, to the best of your ability. practical and reliable.
In the final revision of your test, imagine that you are a student taking the test. Go through each set of directions and all items slowly and deliberately. Time your- self. (Often we underestimate the time students will need to complete a test.) If the test should be shortened or lengthened, make the necessary adjustments. Make sure your test is neat and uncluttered on the page, reflecting all the care and precision you have put into its construction. If there is an audio component, as there is in our hypothetical test, make sure that the script is clear, that your voice and any other voices are clear, and that the audio equipment is in working order before starting the test.
K. Designing multiple choice test items.
In the sample achievement test above, two of the five componets (both of the lis tening sections) specified a multiple-choice format for items. This was a bold step to take. Multiple-choice items, which may appear to bc the simplest kind of item to construct, are extremely difficult to design correctly. Hughes (2003. pp. 76-78) cau- tions against a number of weaknesses of multiple-choice items:
• The technique tests only recognition knowledge.
• Guessing may have a considerable effect on test scores.
• The technique severely restricts what can be tested.
• It is very difficult to write successful items.
• Washback may be harmful.
• Cheating may be facilitated.
The two principles that stand out in support of multiple-choice formats are, of course, practicality and reliability. With their predetermined correct responses and time saving scoring procedures,multiple choice items offer overworked teachers the tempting possi ility of an easy and consistent process of scoring and grading.
1. Multiple-choice items are all receptive, or selective, response items in that the test-taker chooses from a set of responses (commonly called a suupply type of response) rather than creating a response. Other receptive item s include true-false questions and matching lists. (In the discussion here, the guidelines apply primarily to multiple-choice item types and not necessarity to other receptive types.
2. Every miultiple-choice item has a stem, which presents a stimulus, and sever (usually between three and five) options or alternatives to choose from.
3. One of those options, the key, is the correct response, while the others sene as distractors.
Since there will be occasions when multiple-choice items are approprate, com Sider the following four guidelines for designing multiple-choice items for both classroom-based and large-scale situations (adapted from Gronlund, 1998. pp.60-75 and J. D. Brown, 1996, pp.54-57).
1. Design each item to mcasure a specific objective.
Consider this item introduced, and then revised, in the sample test above:
Multiple-choice item, revised.
Voice : Where did George go after the party last night?
S reads : a. Yes, he did.
b. Because he was tired.
c. To Elaine's place for another party.
d. Around eleven o'clock.
Distractor (a) is designed to lure studcors who don't know how to frame thancet questions and therefore serves as an efficient distractor. But what does distracto te actually measure? In fact, the missing defnite article (the) is what J. D. Brown Crnt p. 55) calls an "unintentional clue"-a flaw that could cause the test-taker to elimi nate (e) automatically. In the process, no assessment has been made of indirect ques tions in this distractor. Can you think of a better distractor for (c) that would focus more clearly on the objective?
2. State both stem and options as simply and directly as possible.
We are sometimes tempted to make multiple choice items too wordy,a good rule of thumb is to get directly to the point.Here's an example.
you might argue that the first two sentenes of this item give it some authenticity omplish a bit of schema setting But if you simply want a student to identify ha Ne of medical professional who deals with evesight issues, those sentences are efuous, Morcover by lengthening the stem, you have introduced a potentially afounding lexical item, delerioule that could distract the student unnecessarily.
Another rule of succinctness is to remove needless redundancy from your options, In the following item, whb were is repeated in all three options. It should be placed in the stem to kep the item as succinct as possible.
3. Make certai that the intended answer is clearly the only correct one.
In the proposed unit test described earlier,the following item appeared in the original draf:
Voice : where did george go after the party last ninght?
S reads : a. Yes,he did,
b. Because he was tired.
c.To Elaine's place for another party
d.he went home around elevent o'clock
A quick consideration of the distractor (d) reveals that it is a plausible answer, along with the intended key. (c). Eliminaring unintended possible answers is often the most difficult problem of designing multiple-choice items. With only a minimum of context in each stem, a wide variety of responses may be perceived as correct.
4. Use iten indices to accept, discard, or revise items. The appropriate selection and arrangement of suitable multiple-choice items on a test can best be accomplished by measuring items against three indices: item facility (or item difficulty), item discrimination (sometimes called item differentia- tion), and distractor analysis. Although measuring these factors on classroom tests would be useful, you probably will have neither the time nor the expertise to do this for every classroom test you create, especially one-time tests. But they are a must for standardized norm-referenced tests that are designed to be administered a number of times and/or administered in multiple forms.
1. Item facility (or IF) is the extent to which an item is easy or difficult for the proposed group of test-takers. You may wonder why that is important if in your esti- mation the item achieves validity. The answer is that an item that is too easy (say 99 percent of respondents get it right) or too difficult (99 percent get it wrong) really does nothing to separate high-ability and low-ability test-takers. It is not really per- forming mch "work" for you on a test.
2. Rem discrimination (ID) is the extent to which an item differentiates be- tween high- and low-ability test-takers. An item on which high-ability students (who did well in the test) and low-ability scudents (who didn't) score equally wei would have poor ID because it did not discriminate between the two groups. Conversely, an item that garners correct responses from most of the high ability group and in- correct responses from most of the low-abiliry group has good discrimination power.
3. Dstraaiur efficiemcy is one more luportant rocasure of a tmuitiple-cheice item's valur: in a test, and one thet is related to ltcm discriminarion. The eficiency of distractors is tite extent LI Which (a) the distractors "Jure"a sifficient number of test- takers, especially lower-ability ones, And (b) those responses are samewhat evsaly iistrbuted across all distractors. Those of you who have a fear of mathematical for- mulas will be happy to read that there is no formula for calculating distructor effi riso a that an inspertion of a distripItion of resposes will u5ualiy yicld the information you need.
I. SCORING, GRADING, AND GIVING FEEDRACK Scoring.
As you design a classroom test, you must consider how the test will be scored and graded. Your scoring plaKI reflects the relative weighz that you place on each section and items in each section. The integrated-skills class that we have been using as an Example focuses on listening and speaking skils with some attentian w reading and writing. Three of your nine objectives target reading and writing skilis. How do you assigri scoring to the various components of this test?
Because oral production is a driving force in your cverall objectives, you decide weight on the speakin (oral intervicw) section than on the other chree sections. Five minutes is actuaily a long time to spend in a coe-on-one situa- tion with a student, and some sigrificant infermation can be extracted from such a session. You therefore designate 40 percent of the grade to the oral interview. You consider the listening and reading sectrions to be equally impurtant, but each ot them, especially in this multiple-choice fornat, is of less consequence than the oral nerView. Soo you give each of them a 20 percent weight.
After administering the test once, you may decide to shift some of these weights or to make other changes. You will then bave valuable information about how easy or difficuit the test was, about whether the time limit was reasonable, about your students' affective reaction to it, and about their general performance. Finaily, you will have an intuitive judgment about whether this test correctiy assessed your students. Take note of these impressions, however nonempirical they may he, and use them for revising the test in another term.
M. Grading
Your first thought might be that assigning gades to student performance on this ECSI Would be easy: just give an "A" for 90-100 percent, a "B" for 80-89 percent, and so on. Not so fast! Grading is such a thorny issue that all of Chapter 11 is devoted to the topic. How you assign letter grades to this test is a product of.
• the country, cuiture, and comext of this English classroom,
• institutional expectations (most of them unwritten)
• explicit and implicit definitions of grades that you have set forth,
• the relationship you have establisbed with this class, and
• student expectations that have been engendered in previous tests and quizzes in this class.
For the time being, then, wwe will set aside issues that deal with grading this test in particular, in favor of the comprehensive treatment of grading in Chapter 11.
n. Giving Feedback.
A section on scoring and grading would not be complete without some considera- tion of the forms in which you will offer feedback to your students, feedback that you want to become beneficial washback. in the example test that we have been refer- ring to here-which is noi unusual in the universe of possible formats for periodic classroom test consider the multitude of options.
O. EXERCISES.
[Note: (T) Individual werk; (G) Group er pair work; (C) Whole class discussion.
1. C) Coosult the MLAT website address on page 44 and obtain 35 muc mformation as you can about the MLAT. Aptitude tests propose to predict One s performance in a language course. Review the rationale supportng such testang, and then summarize the controversy surrounding aptitude tests. w aiat can you say about the validity and the ethics of aptitude testing?
2. (G) In pairs, cach assigned to one type of test (aptitude, proficiency. place- ment, diagnostic, or achievement), create a list of broad specifications for the test type you have been assigned: What are the test criteria? What kinds of items should be used? How would you samIple among a umber of possible objectives?
3. (G) Look again at the discussion of objectives (page 49). In a small group. discuss the following scenario: In the case that a teacher is faced with more objectives than are possible to sample in a test, draw up a set of guidelines for choosing which objectives to include on the test and wlich ones to exclude You might sart with considering the issue of the relative importance of all the objectives in the context of the course in question. How does one ade quately sample objectives.
References
Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california.