Rabu, 06 Mei 2020

Assessment Of Meeting 13


SUMMARY (ASSESSING LISTENING,SPEAKING)
HLM : 116 - 184 


ASSESSING LISTENING

In earlier chapters, a number of foundational principles of language assessment were introduced. Concepts like practicality, reliability, validity, authenticity, washback, direct and indirect "testitig,andformative-and sl.itpmative assessment·-are bynow part of your vocabulary. You have become acquainted with some tools for evaluating a "good" test, examined procedures for designing a classroom test, and explored the complex process of creating different kinds of test items. You have begun to absorb the intricate psychometric, educational, and political issues that intertwine in the world of standardized and standards-based testing.

Now our focus will shift away from the standardized testing juggernaut to the level at which you will usually work: the day-to-day classroom assessment of listening, speaking, reading, and writing. Since this is the level at which you will most frequently have the opportunity to apply principles of assessment, the next four chapters ofthis book will provide guidelines and hands-on practice in testing within a curriculum of English as a second or foreign language.

But first, two important caveats. The fact that the four ~nguage skills are discussed in four separate chapters should in no way predispose you to think that those slillls are or should be assessed in isolation. Every TESOL professional (see TBP, Chapter 15) will tell you that the integration ofskills is ofparamount importance in language learning. likewise, assessment is more authentic and provides more washback when skills are integrated. Nevertheless, the skills are treated independently here in order to identify prinCiples, test types, tasks, and issues associated with each one.

Second, you may already have scanned through this book to look for a chapter  on assessing grammar and vocabulary, or something in the way of a focus on form in assessment. The treatment of form-focused assessment is not relegated to a separate chapter here for a very distinct reason: there is no such thing as a test of  grammar or vocabulary that does not invoke one or more of the separate skills of  listening, speaking, reading, or writing! It's not uncommon to fmd little "grammar tests" and "vocabulary tests" in textbooks, and these may be perfectly useful instruments. But responses on these quizzes are usually written, with multiple-choice selection or ftll·in-the-blank items. In this book, we treat the various linguistic forms (phonology, morphology, lexicon, grammar, and discourse) within the context of skill areas. That way,we don't perpetuate the myth that grammar and vocabulary and other ling1.listic forms can somehow be disassociated from a mode of performance.



OBSERVING THE PERFORMANCE OF THE FOUR SKIIS

        Before focusing on listening itself, think about the two interacting concepts of performance and observation. All language users perform the acts of listening,  speaking, reading, and writing. They of course rely on their underlying competence in order to accomplish these performances.When you propose to assess someone's ability in one or a combination of the four skills, you assess that person's competence, but you observe the person'sperformance. Sometimes the performance does  not indicate true competence: a bad night's rest, illness, an emotional distraction, test anxiety, a memory block, or other student-related reliability factors could affect performance, thereby providing an unreliable measure of actual competence.

So, one important principle for assessing a learner's competence is to consider the fallibility of the results of a single performance, such as that produced in a test. As with any attempt at measurement, it is your obligation as a teacher to triangulate your measurements: consider at least two (or more) performances and/or contexts before drawing a conclusion. That could take the form of one or more of the  following designs:
  •        Several  tests that are combined to form an assessment
  •  A single test with multiple test tasks to account for learning styles and performance variables.
  •  In-class and extra-class graded work
  • Alternative forms of assessment (e.g., journal, portfolio, conference, obsen:ation, self-assessment, peeT~sessment).
Multiple measures will always give you a more reliable and valid assessment than a
single measure.


A second principle is one that we teachers often forget. We must rely as much as possible on observable performance in our assessments ofstudents. Observable means  being able to see or hear the performance ofthe learner (the senses oftouch, taste, and  smell don't apply very often to laI?-guage testing!). What, then, is obs~rvable among the  four skills of listening, speaking, reading, and writing? Table 6.1 ·Qffers an answer.  
Isn't it interesting that in the case of the receptive skills, we can observe neither the process of performing nor a product? I can hear your argument already: "But I can see that she's listening because she's nodding her head and frowning and smiling and asking relevant questions." Well, you're not observing the listening performance; you're observing the result of the listening. You can no more observe listening (or reading) than you can see the wind blowing.

THE IMPORTANCE OF liSTENING

Listening has often played second fiddle to its counterpart~ speaking. In the standardized testing industry, a number of separate oral production tests are available (fest of Spoken English, Oral ProfiCiency Inventory, and PhonePass, to name several that are described Chapter 7 of this book), but it is rare to find just a listening test. One reason for this emphasis is that listening is often implied as a component of speaking. How could y~u speak a languag~ without also listening? In addition, the  overtly observable nature of speaking renders it more empirically measurable then  listening. But perhaps a deeper cause lies in universal biases toward speaking. A good speaker is often (unwisely) valued more highly than a good listener. To determine ifsomeone is a proficient user of a language, people customarily ask, "Do you speak: Spanish?" People rarely ask, "Do you understand and speak Spanish?"

Every teacher oflanguage knows that one's oral production ability-other than monologues, speeches, reading alo~d, and the like-is only as good as one's listening  comprehension ability. But of even further impact is the likelihood that input in the  aural-oral mode accounts for a·large proportion of successful language acquisition. In a typical day, we do measurably more listening than speaking (with the exception of one or two of your friends who may be nonstop chatterboxes!).Whether in the workplace, educational, or home contexts, aural comprehension far outstrips oral production in quantifiab~e terms of time, number of words, effort, and attention.

We therefore ne.ed·-to pay close attention to listening as a mode of performance for assessment in the classroom. In this chapter, we will begin with basic prinCiples and types of listenitig, then move to a survey of tasks that can be used to assess listening. (For a review of issues in teaching listening, you may want to read Chapter 16 of TBE)

BASIC TYPES OF IJSTENING

As with all effective tests, designing appropriate assessment tasks in listening begins with the specification of objectives, or criteria. Those objectives may be classified in terms -of several types of listening performance. Think about what you do when you listen. Literally in nanoseconds, the following processes flash through your brain: 
  1. You recognize speech sounds and hold ~ temporary "imprint" of them in  short-term memory.
  2.  You simultaneously determine the type of speech event (monologue, interpersonal dialogue, transactional dialogue) that is being processed and attend to its context (who the speaker is, location, purpose) and the content of the message.
  3.  You use (bottom-up) linguistic decoding skills and/or (top-down) background schemata to bring a plausible interpretation to the message, and assign a literal and intended meaning to the utterance.
  4. In most cases (except for repetition tasks, which involve shQrt-term memory only), you delete the exact linguistic form in which the message was originally received in favor of conceptually retaining important or relevant information in long-term memory.
 Each of these stages represents a potential assessment objective:

  •   comprehending ofsurface structure elements such as phonemes,words, intonation, or a grammatical 
  • category
      understanding of pragmatic context
  • determining meaning of auditory input
  •  developing the gist, a global or comprehensive understanding
From these stages we can derive four commonly identified types of listening performance, each ofwhich comprises a category within'whiCh t01consider assessment
tasks and procedures. 

  1.  Intensive. Listening for perception of the components (phonemes, words, intonation, discourse markers, etc.) of a larger stretch of language.
  2. Responsive. Listening to a relatively short stretch oflanguage (a greeting, question, command, comprehension check, etc.) in order to make an equally short response.
  3.  Selective. Processing stretches of discourse such as short monologues for several minutes in order to "scan" for certain information.The purpose of such performance is not necessarily to look for global or general meanings, but to be able to comprehend designated information in a context of longer stretches of spoken language (such as classroom directions from a teacher, TV or radio news items, or stories). Assessm<:p.t tasks in selective listening could ask students, for example, to listen for names, numbers, a grammatical category, directions (in a map exercise), or certain facts and events.
  4.   Extensive. Listening to· develop a top-down, global understanding of spoken language. Extensive performance ranges from listening to lengthy lectures to listening to a conversation and deriving a comprehensive message or purpose. Listening for the gist, for the main idea, and making inferences are all part of extensive listening.
xMICRO- AND MACROSKII.lS OF LISTENING

A us ful way of synthesizing the above two lists is to consider a finite number of micro- and macroskills implied in the performance of listening comprehension.  Richards' (1983) list of microskills has proven useful in the domain of specifying  objectives for learning and may be even more useful in forcing test makers to carefully identify specific assessment objectives. In the following box, the skills are subdivided into what I prefer to think of as microskills (attending to the smaller bits and chunks of language, in more of a bottom-up process) and macroskills (focusing on the larger elements involved in a top-down approach to a listening task). The microand macros kills provide 17 different objectives to assess in listening.

Implied in the taxonomy above is a notion of what makes many aspects of listening difficult, or why listening is not simply a linear process of recording strings of language as they are transmitted into our brains. Developing a sense ofwhich aspects of listening performance are predictably difficult will help you to challenge your students appropriately and to assign weights to items. Consider the following list of  what makes listening difficult (adapted from Rich~ds, 1983; Dr, 1984; Dunkel, 1991):
  1. Clustering: attending to appropriate "chunks" of language-phrases, clauses, constituent.
  2.  Redundancy: recognizing the kinds of repetitions, rephrasing, elaborations, and insertions that unrehearsed spoken language often contains, and benefiting from that recognition
  3.  Reduced fonns: understanding the reduced fo~ms that may not have been a part,of an English learner's past learning experiences in classes where only formal "textbook" language has been presented.
  4. Perjonnance variables: being able to "weed out" heSitations, false starts, pauses, and corrections innat~ speech.
  5. Colloquial language: comprehending idioms, slang, reduced forms, shared cultural knowledge
  6. Rate ofdelivery: keeping up with the speed of delivery, processing automatically as the speaker continues 
  7. Stress, rhythm, and intonation: correctly understanding prosodic elements of spoken language, which is almost always much more difficult than understanding the smaller phonological bits and pieces.
  8.  Interaction: managing the interactive flow of language from listening to speaking to listening, etc.

DESIGNING ASSESSMENT TASKS: INTENSIVE LISTENING

Once you have determined objectives, your next step is to design the tasks,
including making decisions about how you will elicit performance and how you will' expect the test-taker to respond. We will look at tasks that range from intensive listening performance, such as minimal phonemiC pair recognition, to extensive comprehension of language in communicative contexts. The focus in this section is on the fllicroskills of intensive listening.

DESIGNING ASSESSMENT TASKS: RESPONSIVE liSTENING

A question-and-answer format can provide some interactivity in these lower-end listening tasks. The test-taker's response is the appropriate answer to a question. Appropriate response to a question Test-takers hear: How much time did you take to do your homework?
Test-takers read: (a) In about an hour. (b) About an hour. (c) About $10. (d) Yes, I did.

The objective of this item is recognition of the wh-question bow much and its appropriate response. Distractors are chosen to represnt common learner errors:
(a) responding to how much vs. how much longer; (c) confusing how much in reference to time vs. the more frequent reference to money; (d) confusing a wb-question with a yes/no question.

None of the tasks so far discussed have to be framed in a multiple-choice format. They can be offered in a more open-ended framework in which test-takers write or speak the response',The above item would then look like this: Open-ended response to a question
Test-takers hear: How much time did you take to do your homework? Test-takers write or speak: If open-ended response formats gain a small amount of authenticity and creativity,they of course suffer some in their practicality, as teachers must then read students'  responses and judge their appropriateness, which takes time.

DESIGNING ASSESSMENT TASKS: SELECTIVE IlSTENING

A third type of listening performance is selective listening, in-which the test-taker listens to a limited quantity of aural input and must discern within it some specific information. A number of techniques have been used 'that require selective listening.

ListeningCloze

Listening cloze tasks (sometit11es called cloze dict~tions or partial dictations) require the test-taker to listen to a story. fllonologue,or conversation and simultaneously read the written text in which selected words or phrases have been deleted. Cloze procedure is most commonly associated with reading only In its generic form, the test consists of a passage in which every nth word (typically every seventh word) is deleted and the test-taker is asked to. supply an appropriate word. In a listening cloze task, te~t-takers see a transcript of the passage that they are listening to and flU· in the blanks with the words or phrases that they hear.

One pOlential weakness of listening cloze techniques is that they may simply become reading comprehen~ion tasks. Test-takers who are asked to listen to a story with periodic deletions in the written version may not need to listen at all, yet may still be able to respond with the appropriate word or phrase. You can guard against  this eventuality if the blanks are items with high information load that cannot be easily predicted simply by reading the passage. In the example below (adapted from Bailey, 1998, p. 16), suc~ a shortcoming was avoided by only the criterion of  Numbers.

Other listening cloze tasks may focus Qn a gategory such as verb tenses, articles, two-word verbs, prepositions, or transition words/phrases. Notice two important structural differences between listening cloze tasks and standard  reading cloze. In a listening cloze, deletions are governed by the objective of the  test, not by mathematical deletion of every nth word; and more than one word may be deleted, as in the above example.

Listening cloze tasks should normally use an exact word method of scoring, in which you accept as PQnse only the, actual word or phrase that was spoken and consider other appropriate words as incorrect. (See Chapter 8 for further discussion of these two methods.) Such stringency is warranted; your objective is, after all, to test listening comprehenSion, not grammatical or lexical expectancies.

Information Transfer

 Selective listening can also be assessed through an infor:mation transfer technique in which aurally processed information must be transferred to a visual representation, such as labeling a diagram, ideniifying an element in a picture, completing a form, or showing routes on a map.

At the lower end of the scale of linguistic complexity, simple picture-cued items are sometimes efficient rubrics for assessing certain selected information.

Sentence Repetition

The task of simply repeating a sentence or a partial sentence, or sentence repetition, is also used as an assessment of listening comprehension. As in a dictation (discussed below), the test-taker must retain a stretch of language long enough to reproduce it. and then' must respond with an oral repetition of that stimulus. Incorrect listening comprehension, whether at the phonemic or discourse level, may be manifested in the correctness of the repetition. A miscue in repetition is scored as a miscue in listening. In the case ofsomewhat longer sentences, one could argue that the ability to recognize and retain chunks of language as well as threads of meaning might be assessed through repetition. In Chapter 7, we will look closely at PhonePass, a commercially produced test that relies largely on sentence repetition to assess both oral production and listening comprehension.

Sentence repet~tion is far from a flawless listening assessment task. Buck (2001,p.79) noted that such tasks "are not just tests of listening, but tests of gc;.neral oral skills." Further, this task may test only recognition ofsounds, and it can easily be contaminated by lack ofshort-term ~emory ability, thus invalidating it as an assessment of comprehension alone. And the teacher may never be able to distinguish a listening comprehension error from an oral production ~rror. Therefore,sentence repetition tasks should be used with caution.

DESIGNING ASSESSMENT TASKS: EXTENSIVE LISTENING

Drawing a clear distinction between any two of the categories of listening referred to here is problematic, but perhaps the fuzziest division is between selective and extensive listening. As we gradually move along the continuum from smaller to larger stretches of language, and from micro- to macroskills of listening, the probability of using more extensiveJistening_tasks_jrrcl"eases. Some important questions about designing assessments at this level emerge. 
  • Can listening performance be distinguished from cognitive processing factors such as memory, associations, storage, and recall?
  •  As assessment procedures become more communicative, does the task take into account test-takers' ability to use grammatical expectancies, lexical collocations, semantic interpretations, and pragmatic competence?
  •   Are test tasks themselves correspondingly content valid and authentic-that is, do they mirror real-world language and context?\
  • As assessment tasks beco~e more and more open-ended, they more closely resemble pedagogical tasks, which leads one to ask what the difference is  between assessment and teaching tasks. The answer is scoring: the former imply specified scoring procedures, while the latter do not.

We will try to address these questions as ,ve look at a number of extensive or quasie}tensive listening comprehension tasks.

Dictation

Dictation is a widely researched genre of assessing listenit:lg comprehension. In a dictation, test-takers hear a passage, typically of 50 to 100 words, recited three times: first, at normal speed; then, with long pauses between phrases or natural word groups, during which time test-takers write down what they have just heard; and finally, at normal speed once more so they can check their work and proofread. Here is a sample dictation at the intermediate level of English.

Dictations have been used as assessment tools for decades. Some readers still cringe at the thought of having to render a correctly spelled, verbatim version of a paragraph or story recited by the teacher. Until research on integrative testing was published (see Oller, 1971), dictations were thought to be not much more than glorified spelling tests. However, the required integration of listening and writing in a dictation, along with its presupposed knowledge of grammatical and discourse expectancies, brought this technique back into vogue..;. Hughes (1989), Cohen (1994), Bailey (1998), and Buck (2001) all defend the plausibility of dictation as an integrative test that requires some sophistication in the language in order to process and write down all segments correctly. Thus, I include dictation here under the rubric of extensive tasks, although I am more conlfortable with labeling it quasi extensive.

The difficulty of a dictation task can be easily manipulated by the length of the word groups (or bursts, as they are technically called), the length of the pauses, the speed at which the text is read, and the complexity of the discourse, grammar, and vocabulary used in the passage. Scoring is another matter. Depending on your context and purpose in administering a dictation, you will need to decide on.
scoring criteria for several possible  kinds of errors:

  1.          spelling error only, ,but the word appears to have been heard correctly
  2.      spelling 'and/or obvious misrepresentation of a word, illegible word
  3.      grammatical error (For example, test-taker hears I can~t do it, writes I can do it.)
  4.          skipped word or phrase
  5.          permutation of words
  6.          additional words not in the original
  7.         replacement of a word with an appropriate synonym


Dictation seems to provide a reasonably valid method for integrating listening and writing skills and for tapping into the cohesive elements of language implied in short passages. However, a word of caution lest you assume that dictation provides a quick and easy method of assessing extensivelisterung comprehens~on. If the bursts in a dictation are relatively long (more than five-word segments), this method places a certain amount ofload on memory and processing ofmeaning (Buck, 2001, p. 78). But only a moderate degree of cognitive processing is required, and claiming that dictation fully assesses the ability to comprehend pragmatic or illocutionar elements'of language, context, inference, or senlantics may be going too. far._Finally, one can easily question the authenticity of dictation: it is rare in the real world for people to write down more than a few chunks of information (addresses, phone numbers, grocery lists, directions, for example) at a time.

Despite these disadvantages, the practicality of the administration of dictations, a moderate degree of reliability in a well-established scoring system, and a strong correspondence to other language abilities speaks well for the inclusion of dictation among the possibilities for assessing extensive (or quasi-extensive) listening comprehension.

Communicative Stimulus-Response Tasks

Another-and more authentic-example of extensive listening is found in a popular genre of assessment. task in which the test-taker is presented with a stimulus monologue or conversation and then is asked to respond to a set of compreh~slions. sucntiSki--(as you saw in Chapter 4 in the discussion of standardized testing)  are corrimonly used i.fl commercially produced profiCiency tests. The monologues, lectures. and brief conversations used in such tasks are sometimes a little contrived, and certainly the subsequent multiple-choice questions don't mirror communicative, real-life situations. But with some care and creativity, one can create reasonably authentic stimuli, and in some rare cases the response mode (as shown in one example below) actually approaches complete authenticity. Here is a typical example of such a task.

Authentic Listening Tasks

Ideally, the language assessment field would have a stockpile of listening test types that are cognitively demanding. communicative, and authentic, not to mention interactive by means of an integration with speaking. However, the nature of a test as a of performance and a set of tasks with limited time frames implies an equally linlited capacity to mirror all the real-world contexts of listening perfonnance. "There is no such thing as a communicativet," stated Buck (200 1, p. 92). "Every test requires some comPofieiits-oicomIDiiiifcatlve language ability, and no test covers them all.Similarly, with the notion of authenticity, every task shares some characteristics with target-language tasks, and no test is completely authentic."

Beyond the rubrics of intensive, responsive, selective, and quasi-extensive communicative contexts described above, can we assess aural comprehension in a truly communicative context? Can we, at this end of the range of listening tasks, ascertain from test-takers that they have processed the main idea(s) of a lecture, the gist of a story, the pragmatics of a conversation, or the unspoken inferential data present in most authentic aural input? Can we assess a test-taker's comprehension of humor, idiom, and metaphor? The answer is a cautious yes, but not without some concessions to practicality.· And the answer is a more certain yes if we take the liberty of stretching the concept of assessment to extend beyond tests and into a broader framework of  tna!iy Here are some possibilities. ­
  1. Note-taking. In the academic world, classroom lectures by professors are common features of a non-native English-user's experience. One form of a midterm examination at the American Language Institute at San Francisco State University (Kahn, 2002) uses a IS-minute lecture as a stimulus. One among several response formats includes note-taldng by the test-takers. These notes are evaluated by the teacher on a 30-point system, as follows:
  2. Editing. Another authentic task provides both a written and a spoken stimulus, and requires the test-taker to listen for discrepancies. Scoring achieves relatively high reliability as there are usually a small number of specific differences that must be identified. Here is the way the task proceed.
One potentially interesting set of stimuli for such a task is the description of a political scandal frrst from a newspaper with a political bias, and then from a radio broadcast from an "alternative" news station. Test-takers are not only forced to listencarefully to differences but are subtly informed about biases in the news.

3.  Interpretive tasks. One of the intensive listening tasks described above was paraphrasing a story or conversation. An interpretive task extends the stimulus material to a longer stretch of discourse and forces the test-taker to infer a response. Potential stimuli include
  • ·         song lyrics,
  •         [recited] poetry, .
  • ·         radio/television news reports, and
  • ·         an oral account of an experience.

Test-takers are then directed to interpret the stimulus by answering a few questions
(in open-ended form). Questions might be:

·         "Why was the Singer feeling sad?"
·         "What events might have led up to the reciting of this poem?"
·          "What do you think the political activists might do next, and why?"
·         "What do you think the storyteller felt about the mystorious disappearance of  her necklace?"

This kind of task moves us away from what might traditionally be considered a test
toward an informal assessment, or possibly even a pedagogical technique or activity. But the task conforms to certain time limitations, and the questions can be quite specific, even though they ask the test-taker to use inference.\Ylhile reliable scori.tlg may be an issue (there may be more than one correct interpretation), the authenticity of the interaction in this task and potential washback tothe student surely give it some prominence among communicative assessment procedures.
4.       Retelling. In a related task, test-takers listen to a story or news event and simply retell it, or summarize it, either orally (on an audiotape) or in writing. In so doing, test-takers must identify the gist, main idea, purpose, supporting points, and/or conclusion to show full comprehension. Scorillg is partially predetermined by specifying a minimu number of elements that must appear in the retelling. Again reliability may suffer, and the time and effort needed to read and evaluate the response lowers practicality. Validity, cognitive processing, communicative ability, and authenticity are all well incorporated into the task.

A ftfth category of listening comprehension was hinted at earlier in the chapter: interactive listening. Because such interaction presupposes a process of speaking in concert with listening, the interactive nature of listening will be addressed in the next chapter. Don't forget that a significant proportion of realworld listening performance is interactive. With the exception of media input, speeches, lectures, and eavesdropping, many of our listening efforts are directed toward a two-way process of speaking and listening in face-ta-face conversations.

EXERCISES
[Note: (I) Individual work; (G) Group or pair work; (C) Whole-class discussion.]
  1.            (C) In Table 6.1 on page 118, it is noted that one cannot actually observe listening and reading performance. Do you agree?-:Anddo you agree that there isn't even a product to observe for speaking, listening, and reading? How, then, can one infer the competence of a test-taker to speak, listen, and read a language?
  2.           (C) Given that we spend much more time listening than we do speaking, why are there many more tests of speaki1l:g than listeninG.
  3.        (G) Look at the list of micro- and macroskills of listening on pages 121-122. In pairs, each assigned to a different skill (or two), brainstorm some tasks that assess those skills. Present your fmdings to the rest ot-the class.
  4.         (G) Eight characteristics of listening that make listening "difficult" are listed on page 122. In pairs, each asSigned to an assessment task itemized in this chapter, decide which of the eight factors, in order of significance; contribute to the potential difficulty of the items. Report back to the class.
  5.           (G) Divide the basic types of listening among groups or pairs, one type for each. Look at the sample assessment teclmiques provided and evaluate them according the five principles (practicality, reliability, validity [face and content], authenticity, and washback). Present your critique to the rest of the class.
  6.                (G) In the same groups as in #5 above and with the same type of listening, design some other item types, different from the one(s) provided here, that assess the same type of listening performance.
  7.                 (G) With a linguistic objective assigned to each pair or group, construct a listening cloze test for two-word verbs, verb tenses, prepositions, transition words, articles, and/or other grammatical categorieS,
  8.            (I/C) On page 131, you are reminded that dictations are considered by some assessment specialists to be integrative (requiring the integration of listening, writing, reading [proofreading], along with attendant grammatical and discourse abilities). Is this a valid claim? Justify your response.


FOR YOUR FURTIlER READING

Buck, Gary. (2001). Assesstng listening. Cambridge: Cambridge University Press.
One of a series of very useful ref~rence books on assessing specific skill areas published by Cambridge University Press, this. one gives an overview of research and pedagogy on listening comprehension and demonstrates many different assessment procedures in common use.

Richards, Jack C. (1983). Listening comprehension: Approach, design, procedure.
TESOL Quarterly, 17, 219-239.
Even though Richards published this article in 1983, it still provides a standard backdrop for teaching listening skills. While formal assessment is not directly addressed, informal assessmeht is implied in its pedagogical focus on practical classroom techniques.

Mendelsohn, David J. (1998). Teaching listening. Annual Review of Applied
Linguistics, 18, 81-101.
Mendelsohn's overview of research ort teaching 1i~tening proVides an excellent foundation for understanding assessment tasks. Me focuses on a  strategy based approach to teaching listening and adds an annotated bibliography of professional resource books.


ASSESSIIG SPEAKING

From a pragmatic view of language performance, listening and speaking are almost always closely interrelated. While it is possible to isolate some listening performance types (see Chap"ter 6),'it is very difficult to isolate oral-production tasks that do not directly involve the interaction of aural comprehension. Only in limited contexts of speaking (monologues, speeches, or telling a story and reading aloud) can we assess oral language without the aural participation of an interlocutor.
While speaking is a productive skill that can be directly and empirically observed, those observations are invariably colored by the accuracy and effectiveness Of a test-ta}{er'$ list~I1iI1g skill, which necessarily compromisestlieh"rellability and validity of an oral production test. How do you know for certain that a speaking score is exclusively a measure of oral production without the potentially frequent clarifications of an interlocutor? TItis interaction of speaking and listening challenges the designer of an oral production test to tease apart, as much as possible, the factors accounted for by aural intake.

Another challenge is the design of elicitation techniques. Because most speaking is the product of creative construction oflinguistic strings, the speaker makes choices of leXicblf,srructure, and discourse:-If-your-goal is to-have test-takers demonstrate certain spoken grammatical categories, for example, the stimulus you design must elicit those grammatical categories in ways that prohibit the test-taker from avoiding or paraphrasing and thereby dodging production of the target form.
All of these issues will be addressed in this chapter as we review types of spoken language and micro- and macros kills of speaking, then outline numerous tas~s for assessing speaking.

BASIC TYPES OF SPEAKING

In Chapter 6, we cited four categories of listening performance assessment tasks. A similar taxonomy emerges for oral production.

  1.  bnitative. At one end of a continuum of types of speaking performance is the ability to simply parrot back (imitate) a word or phrase or possibly a sentence. While this is a purely phonetic level of oral production, a number of prosodiC, lexical, and grammatical properties of language may be included in the criterion performance.We are interested only in what is traditionally labeled "pronunciation"; no inferences are made about the test-taker's ability to understand or convey meaning or to participate in an interactive conversation. The only role of listening here is in the short-term storage of a ptonlpt, just long enough to, allow the speaker to retain the short stretch of language that must be imitated.
  2. Intensive. A second type of speaking frequently employed in assessment  contexts is the production of short stretches of oral language designed to demonstrate competence in a narrow band of grammatical, phrasal, lexical, or phonological relationships (such as prosodic elements-intonation, stress, rhythm, juncture). The speaker must be aware of semantic properties in order to be able to respond, but interaction with an interlocutor or test administrator is minimal at best. Examples of intensive assessment tasks include directed response tasks, reading aloud, sentence and dialogue completion; limited picture-cued tasks ill:~luding simple sequences; and translation up to the simple Sentence level.
  3. Responsive. ,Responsive assessment tasks include interaction and test com- v prehension but at the somewhat limited level of very short conversations, standard greetings and small talk, simple requests and comments, and the liken
  4.  Interactive. The difference between responsive and interactive" speaking is in the length and complexity of the interaction, which sometimes includes mUltiple exchanges and/or multiple participants. Interaction can take the two forms of transactional language, which has the purpose of exchanging specific information, or interpersonal exchanges, which have the purpose of maintaining social relationships. (In tfie three dialogues cited above, A and B were transactional, and C was interpersonal.) In interpersonal exchanges, oral production can become pragmatically complex with the need to speak in a casual register and use colloquial language, ellipSis, slang, humor, and other sociolinguistic conventions.
  5.  Extensive (monologue). Extensive oral production tasks include speeches, oral presentations, and story-telling, during which the opportunity for oral interaction from listeners is either highly limited (perhaps to nonverbal responses) or ruled out altogether. Language style is frequently more deliberative (planning is involved) and" formal for extensive tasks, but we cannot rule out certain informal monologues" such as casually delivered spe"ech (for exatPple, my vacation in the mountains, a recipe for outstanding pasta primavera, recounting the plot of a novel or movie).

5.
MICRO- AND MACROSKUJS OF SPEAKING


In Chapter 6, a list of listening micro- and macroskills enumerated the various components of listening that make up criteria for assessment. A similar list of speaking skills can be drawn up for the same purpose: to serve as a taxonomy of skills from which you 'will select one or several that will become the objective(s) of an assessment task. The microskills refer to producing the smaller chunks of language such as phonemes, mofQ!!emes, words, collocations, and phrasal units. The macroskills imply thespeaker's focus on die-larger eiements: flu~ng, dis-course, function, style, cohesion, nonverbal c?mmunication, and strategic. ? Ptions. The niicro-and macros kills total roughly 16 different objectives to assess in speaking.


There is such an array of oral production tasks that a complete treatment is almost impossible within the confines of one chapter in this book. Below is a consideration of the most common techniques with brief allusions to related tasks. As already noted in the introduction to this chapter, consider three important issues as you set out to design tasks:
1.No speaking task is capable of isolating the single skill of'oral production. Concurrent involvement of the additional performance of aural comprehension, and  possibly reading, is usually necessary.
2.EliCiting the specific criterion you have designated for a task can be tricky because beyond the word level, spoken language offers a number of productive options to test-takers.
3.Because of the above two characteristics of oral production assessment, it is important to carefully specify scoring procedures for a response so that ultimately you achieve as high a reliability index as possibie.

DESIGNING ASSESSMENT TASKS: IMITATIVE SPEAKING

You may be surprised to see the inclusion ofsimple phonological imitation in a consideration of assessment of oral production. After all, endless repeating of words, phrases, and sentences was the province of the long-since-discarded Audiolingual Method, and in an era of communicative language teaching, many believe that nonmeaningful imitation ofsounds is fruitless. Such opinions-have faded in recentyears as we discovered that an overemphasis on fluency can sometimes lead to the decline of accuracy in speech. And so we have been paying more attention to pronunciation, especially' suprasegmentals, in an attempt to help learners be more comprehensible.

An occasional phonologically focused repetition task is warranted as long as repetition tasks are not allowed to occupy a dominant role in an overall oral production assessment, and as long as you artfully avoid a negative washback effect. Such tasks range from word level to sentence level, usually with each item focusing on. a specific phonological criterion. In a simple repetition task, test-takers repeat the stimulus, whether it is a pair ofwords, a sentence, or perhaps a question (to test for intonation production).

PHONEPASS® TEST

An example of a popular test that uses imitative (as well as intensive) production tasks is PhonePass, a widely used, commercially available speaking test in many countries. Among a number of speaJdng tasks on the test, repetition of sentences (of 8 to 12 words) occupies a prominent role. It is remarkable that research on the PhonePass test has supported the construct validity of its repetition tasks not just for a testtaker's phonological ability but also for discourse and overall oral production ability (fownshend et al., 1998; Bernstein et aI., 2000; Cascallar & Bernstein, 2000).

The PhonePass test elicits computer-assisted oral production over a telephor~,e. Test-takers. read aloud, repeat sentences, say words, and answer questions. With a downloadable test sheet as a reference, test-takers are directed to telephone a designated number and listen for directions. The test has five sections.

DESIGNING ASSESSMENT TASKS: INTENSIVE SPEAKING

At the intensive level, test-takers are prompted to produce short stretches of discourse (no more than a sentence) through which they demonstrate linguistic ability at a specified level of language. Many tasks are "cued" tasks in that they lead the testtaker into a narrow band of possibilities.

Parts C and D of the PhonePass test fulfill the criteria of intensive tasks as they elicit certain expected forms of language. Antonyms like high and low, happy and sad are prompted so that the, automated scoring mechanism anticipates only one word. The either/or task of Part D fulfills the same criterion. Intensive tasks may also be described as limited response tasks (Madsen, 1983), or mechanical tasks (Underhill, 1987), or what classroom pedagogy would label as controlled responses.

Directed Response Tasks

In this type of task, the test administrator elicits a particular grammatical form or a transformation of a sentence. Such tasks are clearly mechanical and not communicative, but they do require minimal processing ofmeaning in order to produce the correct grammatical output.

Read-Aloud Tasks

Intensive reading-aloud tasks include reading beyond the sentence level up to a paragraph or two. This technique is easily administered by selecting a passage that incorporates test specs and by recording the test-taker's output; the scoring is relatively easy because all of the test~taker's oral production is controlled. Because of theresults of research on the PhonePass test, reading aloud may actually be a surprisingly strong indicator of overall oral production ability.

For many decades, foreign language programs have used reading passages to analyze oral production. Prator's (1972) Manual ofAmerican English Pronunciation included a "diagnostic passage" of about 150 words that students could read aloud into a tape recOrder. Teachers listening to the recording would then rate students on a number of phonological factors (vowels, diphthongs, consonants, consonant clusters, stress, and intonation) by completing a two-page diagnostic checklist on which all errors or questionable items were noted. These checklists ostensibly offered direction to the teacher for emphases in the course to come.

An earlier form of the Test of Spoken English (fSE®, see below) incorporated
one read-aloud passage .of about 120 to 130 word Such a rating list does not indicate how to gauge intelligibility, which is mentioned in both lists. Such slippery terms remind us that oral production scoring, even with the controls that reading aloud offers, is still an inexact science. Underhill (1987, pp. 77-78) suggested some variations on the task of simply reading a short passage:
  •  reading a scripted dialogue, with someone else reading the other part
  •     reading sentences containing minimal pairs, for example:
             Try not to heat/hit the pan too much.
            The doctor gave me a bil/pill.
  • ·        reading information from a table or chart

      If reading aloud shows certain practical advantages (predictable output, practicality, reliability in scoring), th~re are several drawbacks to using this technique for ass~ssing oral production. Reading aloud is somewhat "!!!e:Y1h,!;n!~s,.in that we seldom read anything aloud to ·semeone else in the- real world, with--the .exception of a parent reading to a child, occasionally sharing a written story with someone, or giving a scripted oral presentation. Also, reading aloud calls on certain specialized oral abilities that may not indicate one's pragmatic ability to communicate orally ill face-ta-face contexts. You should therefore employ this technique with some caution, and certainly supplement it as an assessment task with other, more communicative procedures.

Sentence/Dialogue Completion Tasks and Oral Questionnaires

Another technique for targeting intensive aspects of language requires test-takers to read dialogue in which one speaker's lines have been omitted. Test-takers are first given time to read through the dialogue to get its gist and to think about appropriate lines to fill in. Then as the tape, teacher, or test administrator produces .

Picture-Cued Tasks

One of the more popular ways to elicit oral language performance at both intensive and extensive levels is a pictl1re-cued stimulus that requires a description from the testtaker. Pictures may be very simple, designed to elicit a word or a phrase; somewhat more elaborate and "busy"; or composed of a series that tells a story or incident. Here is an example of a picture-cued elicitation of the production of a simple minimal pair.

Translation (of T,imited Stretches of Discourse)

Translation is a part of our tradition in language teaching that we tend to discount or disdain, if only because our current pedagogical stance plays down its importance. Translation methods of teaching are certainly passe in an era of direct approaches to creating communicative classrooms. But we should remember that in countries where English is not the native or prevailing language, translation is a meaningful communicative device in contexts where the English user is. called on to be an interpreter. Also, translation is a well-proven communication strategy for learners of a second language.

Under certain constraints, then, it is not far-fetched to suggest translation as a device to check oral production. Instead of offering pictures or written stimuli, the test-taker is given a native language word, phrase, or sentence and is asked to translate it. Conditions may vary from expecting an instant translation of an orally elicited linguistic target toallowmg more thinking time before producing a translation of somewhat longer texts, which may option3.ny be offered to the test-taker in written form. (franslation of extensive, texts is discussed at the end of this chapter.) As an assessment procedure, the advantages of translation lie in its control of the output ofthe test-taker, whichof,:ot;U,"se means that scoring is more easily specified.


DESIGNING ASSESSMENT TASKS: RESPONSIVE SPEAKING

Assessment of responsive tasks involves brief interactions with an interlocutor, d~ffering from intensive tasks in the increased creativity given to the test-taker and   from interactive tasks by the somewhat limited length of utterances.

Questi!?n and Answer

Question-and-answer tasks can consist of one or two questions from an interviewer,or they can make up a portion of a whole battery of questions and prompts in an oral interview. They can vary from simple questions like "What is this called in English?" to complex questions like "What are the steps governments should take, if any, to stem the rate of deforestation in tropical countries?" The first question is intensive in its purpose; it is a display question intended to elicit a predetermined correct response. We have already looked at some of these types ot questions in the previous section. Questions at the responsive level tend to be genuine referential questions in which the test-taker is given more opportunity to·produce meaningful language in response.

Giving Instructions and Directions

We are all called on in our dally routines to read instructions on how to operate an appliance, how to put a bookshelf together, or how to create a delicious clam. Somewhat less frequent is the mandate to provide such instructions orally, but this speech act is still relatively common. Using such a stimulus in an assessment context provides an opportunity for the test-taker to engage in a relatively extended stretch of discourse, to be very clear and specific, and to use appropriate discourse markers and connectors. The technique is Simple: the administrator poses the problem, and the test-taker responds. Scoring is based primarily on comprehensibility and s~condari1y on other specified grammatical or discourse categOries. Here are some possibilities.

Paraphrasing

Another type of assessment task that can be categorized as responsive asks the testtaker to read or hear a limited number of sentences (perhaps two to five) and-produce a paraphrase of the sentence.

TEST OF SPOKEN ENGLISH (TSE@)

Somewhere straddling responsive, interactive, and extensive speaking tasks lies another popular commercial oral production assessment, theTest of Spoken English (TSE)'. The TSE is a 20-minute audiotaped test of oral language ability within anacademic or professional environment. TSE scores are used by many North American institutions of higher education to select international teaching assistants. The scores are also used for selecting and certifying health professionals such as physicians, nurses, pharmacists, physical therapists, and veterinarians.

The tasks on theTSE are designed to elicit oral production in various discourse categories rather than in selected phQ~ogical, ~a,ticalt or l~cal !argets. The follOwing content specifications for the TSE represent the discourse and pragmatic contexts assessed in each administration:

1. Describe something physical.
2. Narrate from presented material.
3. Summarize information of the speaker's own choice.
4. Give directions based on visual materials.
5. Give instructions.
6. Give an opinion.
7. Support an. opinion.
8. Compare/contrast.
9. Hypothesize.
10. Function "interactively."
11. Define.

Using these specifications, Lazaraton andWagner (1996) examined 15 different specific tasks in collecting background data from native and non-native speakers of  English.

a. giving  a personal deSCription
b. describing a daily routine
c.  suggesting a gift and supporting one's choice
d. recommending a place to visit and supporting one's choice
e. giving directions
f.  describing a favorite movie and supporting one's choice
g.  telling a story from pictures
h. hypothesizing about future action
i.  hypothesizing about a preventative action
j.  making a telephone call to the dry cleaner
k.describing an important news event

From their fmdings, the researchers were able to report on the validity of the tasks,
especially the match between the intended task functions and the actual outpu:: of
both native and non-native speakers.

DESIGNING ASSESSMENT TASKS: INTERACTIVE SPEAKING

The fmal two categories of oral production assessment (interactive and extensive speaking) include tasks that involve relatively long stretches of interactive discourse (interviews, role plays, discussions, games) and tasks. of equally long duration but that.

involve less interaction (speeches, telling longer stories, and extended explanations and translations).The obvious difference between the two sets of tasks is the degree of interaction with'an interlocutor. Also, interactive tasks are what some would describe as interpersonal, while the fmal category includes more transactional speech events.

Interview

When "oral production assessment" is mentioned, the first thing that comes to mind is an oral interview: a test administrator and a test-taker sit downJn a direct face-toface exchange and proceed through a protocol of questiof!s and directives. The interview, which may be tape-recorded for re-listening, is then scored on one or more parameters such as accuracy in pronunciation and/or grammar, vocabulary usage, fluency, sociolinguistic/pragmatic appropriateness, task accomplishment, and even comprehension. Interviews can vary in length from perhaps five to forty-five minutes, depending on their purpose and context. Placement interviews, designed to get a quick spoken sample from a student in order to verify placement into a course.

Every effective interview contains a number of nlandatory stages. 1W'0 decades ago, Michael Canale (1984) proposed a framework for oral proficiency testing that has withstgod the test of time. He . suggested that test-takers will perform at their best if they are led through four stages: 

  1.       Warm-up. In a minute or so of preliminary small talk, the interviewer directs mutual introductions, helps the test-taker become comfortable with the situation, apprises the test-taker of the format,and allays anxieties. No scoring of this phase takes place.
  2.       Level check. Through a series of preplanned questions, the interviewer stimulates the test-taker to respond using -expected or predicted forms and functions. If, for example, from previous test information, grades, or other data, the test..taker has been judged to be a "Level 2" (see below) speaker, the interviewer'S prompts will attempt to confirm this assumption. The responses may take very simple or very complex form, depending on the entry level of the learner. Questions are usually designed to elicit grammatical categories (such as past tense or subject-verb agreement), discourse structure (a sequence of events), vocabulary usage, and/or sociolinguistic factors (politeness conventions, formal/informal language.
  3.       Probe. Probe questions and prompts challenge test-takers to go to the heights of their ability, to extend beyond the limits ofthe interviewer'S expectation through increaSingly difficult questions. Probe questions may be complex in their framing and/or complex in their cognitive and linguistic demand. Through probe items, the interviewer discovers the ceiling or limitation of the test-taker's proficiency. This need not be a separate stage entirely, but might be a set of questions that are interspersed into the previous stage.
  4.        Wind-down. This fmal phase of the interview is simply a short period of time during which the interviewer encourages the test-taker to relax with some easy questions, sets the test-laker's mind at ease, and provides information about when and where to obtain the results of the interview. This part is not scored.

The success of an oral interview will depend on
  •  clearly specifying administrative procedures of the assessment (practicality),
  • focusing the questions and probes on the purpose of the assessment (validity),
  •  appropriately eliciting an optimal amount and quality of oral production from the test~taker (biased for best performance), and
  • creating a conSistent, workable scoring system (reliability).


This last issue is the thorniest. In oral production tasks that are open-ended and that involve a significant level of interaction, the interviewer is forced to make judgments that are susceptible to some unreliability. Through experience, training, and careful attention to the linguistic criteria being assessed, the ability to make su~h judgments accurately will be acquired. InTable 7.2, a set of deSCriptions is given for scoring open-ended oral interviews. These descriptions come from an earlier version of the Oral Proficiency Interview and are useful for classroom purposes.

The test administrator's .challenge is to assign a score, ranging from 1 to 5, for each of the six categories indicated above. It may look easy to do, but in reality the lines of distinction between levels is quite difficult to pinpoint. Some training or at least a good deal of interviewing experience is required to make accurate ·aSsessments of oral production in the six categories. Usually the six scores are then amalgamated into one holistic score, a process that might not be relegated to a simple mathematical average if you wish to put more weight on some categories than you do on others.

This five-point scale, once known as "FSI levels" (because they were frrst advocated by the Foreign Service Institute in Washington, D.C.), is still in popular use among U.S. government foreign service staff for designating profiCiency in a foreign language.

A variation on the usual one-on-one format with one intervie.wer and one testtaker is to place two test-takers at a time with the interviewer. An advantage of a two-on-one interview is the practicality of scheduling twice as rrrany candidates in the same time frame, but more significant is the opportunity for student-student interaction. By deftly posing questions, problems, and role plays, the interviewer can the output of the test-takers while lessening the need for his or her own output. A further benefit is the probable increase in authenticity when two testtakers can actually converse with each other. Disadvantages are equalizing the output between the two test-takers, discerning the interaction effect of unequal comprehension and production abilities, and scoring two people simultaneously.

Discussions and Conversations

As formal assessment devices, discussions and conversations with and among students are difficult to specify and even more difficult to score. But as informal techniques to assess learners, they offer a level of authenticity and spontaneity that other assessment techniques may not provide. Discussions may be especially appropriate tasks through which to elicit and observe such abilities as

§  topic nomination, maintenance, and termination;
§  attention getting, interrupting, floor holding, control;
§   clarifying, questioning, paraphrasing;
§   comprehension Signals (nodding, "uh-huh,""hmm," etc.);
§   negotiating meaning;
§  intonation patterns for pragmatic effect;
§  kinesics, . eye contact, proxemics, body language; and
§   politeness, formality, and other SOCiolinguistic factors.

AsseSSing the performance of participants through scores or checklists (in which appropriate or inappropriate manifestations of any category are noted)  should be carefully designed to suit the objectives of the observed discussion. Of course, discussion is an integrative task, and so it is also advisable to give some cognizance to comprehension performance in evaluating learners. 
Games
Among informal assessment devices are a variety of games that directly involve language production.

ORAL PROFICIENCY INTERVIEW (OPI)

The best-known oral interview format is one that has gone through a considerable metamorphosis over the last half-century, the Oral Proficiency Interview (OPI). Originally known as the Foreign Service Institute (FSI) test, the OPI is the result of a historical progression of revisions under the auspices of several agencies, including the Educational Testing Service and the American Council on Teaching Foreign Languages (ACTFL). The latter, a- professional society for research on foreign language instruction and assessment, has now become the prinCipal body for promoting the use of the OPI."The OP! is widely used across dozens of languages around the world. Only certified examiners are authorized to administer the OP!; certification workshops are available, at costs of around $700 for ACTFL members, through ACTFL at selected sites and conferences throughout the year.
DESIGNING ASSESSMENTS: EXTENSIVE SPEAKING

Extensive speaking tasks involve complex, relatively lengthy stretches of discourse. They are frequently variations on monologues, usually with minimal verbal interaction. Oral Presentations In !he academic and professional arenas, it would not be uncommon to be called on to present areport-,a-paper;a marketing plan,a-sales-idea, a design of a new product, or a method. A summary of oral assessment techniques would therefore be incomplete without some consideration of extensive speaking tasks. Once again the rules ! I for effective assessment must be invoked: (a) specify the criterion, (b) set appro priate tasks, (c) elicit optimal output, and (d) establish practical, reliable scoring procedures. And once again scoring is the key assessment challenge.

Picture-Cued Story-Telling

One of the most common techniques for eliciting oral production is through visual pictures, photographs, diagrams, and charts. We have already looked at this' elicitation device for intensive tasks, but at this level we consider a picture or a series of pictures as a stimulus for a longer story or deSCription .

It's always tempting to throw any picture sequence at test-takers and have them talk for a minute or so about them. But as is true of every assessment of speaking ability, the objective of eliciting narrative discourse needs to be clear. In the above example (with a little humor added!), are you testing for oral vocabulary (girl, alarm, coffee, telephone, wet, cat, etc.), for time relatives (before, after, when), for sentence connectors (then, and then, so), for past tense of irregular verbs (woke, drank, rang), and/or for fluency in general? Ifyou are eliCiting specific grammatical or discourse features, you might add to the directions something like "Tell the story that these pictures describe. Use the past tense of verbs." Your criteria for scoring need to be clear about what it is you are hoping to assess. Refer back to some of the guidelines suggested under the section on oral interviews, above, or to the OPI for  some general suggestions on scoring such a narrative.

Retelling a Story, News Event

In this type of task, test-takers hear or read a story or news event that they are asked to retell. This differs from the paraphrasing task discussed above (pages 161-162) in that it is a longer stretch of discourse and a different genre. The objectives in assigning such.a task vary from listening comprehension of the original to production of a number of oral discourse features (communicating sequences and relationships 01 events, stress and emphasis patterns, . "expression" in the case of a dramatic story), fluency, and interaction with the hearer. Scoring should of course meet the intended criteria.

Translation (of Extended Prose)

Translation of words, phrases, or short sentences was mentioned under the category of-intensive speaking. Here, longer texts are presented for the test-taker to read in the native language and then translate into English.Those texts could come in many forms: dialogue, directions for assembly of a product, a synopsis of a story or play or movie, directions on how to find something on a map, and other genres. The advantage of translation is in the control of the content, vocabulary, and, to some extent, the grammatical and discourse features. The disadvantage is that translation of longer texts is a highly specialized skill for which some individuals obtain post-baccalaureate degrees! To judge a nonspecialist's oral language ability on such a skill may be completely invalid, especially if the test-taker has not engaged in translation at this level. Criteria for scoring should therefore take into account not only the purpose in stimulating a translation but the possibility of errors that are unrelated to oral production ability.

EXERCISES

 [Note,: (I) Individual work; (G) Group or pair work; (C) Whole-class discussion.

  1.  (G) In the introduction to the chapter, the unique challenges of testing speaking were described (interaction effect, elicitation techniques, and scoring). In pairs, offer practical examples of one of the challenges, as aSSigned to your pair. Explain your examples to the class.
  2.        (C) Review the five basic types of speaking that were outlined at the beginning of the chapter. Offer examples of each and pay special attention to distinguishing between imitative and intensive, and between responsive and interactive.(G) Look at the list of micro- and macroskills of speaking on pages 142-143. In pairs, each assigned to a different skill (or two), brainstorm some tasks that  assess those skills. Present your findings to the rest of the class.
  3. (C) In Chapter 6, eight characteristics oflistening (page 122) that make listening "difficult"were listed. What makes speaking difficult? Devise a similar ijst that could fonn a set ofspecifications to pay special attention to in assessing speaking.
  4.    (G) Divide the five basic types of speaking among groups or pairs, one type  for each. Look at the sample assessment techniques provided and evaluate them according to the five pririciples (practicality, reliability, validity [especially face and content], authenticity, and washback). Present your critique to  the rest of the class.
  5.    (G) In· the same groups as in question #5 above, with the same type of speaking, design some other item types, different frorp the one(s) provided here, that assess the same type of speaking performance.
  6. (I) Visit the website listed for the PhonePass test. If you can afford it and you are a non-native speaker of English, take the test. Report back to the class on .'how valid, reliable, and authentic you felt the test was.
  7. (G) Several scoling scales are offered'in-this-chapter, ranging from simple (2-1-0) score categories to the more elaborate rubric used for the OP!. In groups, each assigned to a scoring scale, evaluate the strengths and weaknesses of each. Pay special attention to intra-rater and inter-rater reliability.
  8.    (C) If pOSSible, role-playa formal oral interview in your class, with one student (with beginning to intermediate profiCiency in a language) acting as the testtaker and another (with advanced proficiency) as the test administrator. Use the sample questions provided on pages 169-170 as a guide. This role play will require some preparation. The rest of the class will then evaluate the effectiveness of the oral interview. Finally, the test-taker and administrator can offer their perspectives on the experience.

FOR YOUR FURTHER READING

Underhill, Nic. (1987). Testing spoken language:A handbook of oral testing techniques. Cambridge: Cambridge University Press.
This practical tnanual .on assessing spoken language is still a widely used collection of techniques despite the fact that it was published in 1987. The chapter~re organized into types of tests, elicitation techniques, scoring systems, and a discussion of several types of validity along with reliability.
Brown,].D. (1998). New ways of classroom assessment. Alexandria,VA: Teachers of  English to Speakers of Other Languages.
One of the many volumes in TESOL's "New Ways" series, this one presents a collection of assessment techniques across a wide range of skill areas. The two sections on assessing oral skills offer 17 different techniq.





REFERENCES :



Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california. 





Tidak ada komentar:

Posting Komentar

ASSESSMENT FOR MEETING 15

            Assessing grammar,vocabulary Chapter one Differing notions of ‘grammar’ for assessment Introduction   Even...