SUMMARY (ASSESSING LISTENING,SPEAKING)
HLM : 116 - 184
ASSESSING
LISTENING
In earlier
chapters, a number of foundational principles of language assessment were
introduced. Concepts like practicality, reliability, validity, authenticity,
washback, direct and indirect "testitig,andformative-and sl.itpmative assessment·-are
bynow part of your vocabulary. You have become acquainted with some tools for
evaluating a "good" test, examined procedures for designing a
classroom test, and explored the complex process of creating different kinds of
test items. You have begun to absorb the intricate psychometric, educational,
and political issues that intertwine in the world of standardized and
standards-based testing.
Now our focus
will shift away from the standardized testing juggernaut to the level at which
you will usually work: the day-to-day classroom assessment of listening,
speaking, reading,
and writing. Since this is the level at which you will most frequently have the
opportunity to apply principles of assessment, the next four chapters ofthis
book will provide guidelines and hands-on practice in testing within a
curriculum of English as a second or foreign language.
But first, two important
caveats. The fact that the four ~nguage skills are discussed in four separate
chapters should in no way predispose you to think that those slillls are or
should be assessed in isolation. Every TESOL professional (see TBP, Chapter 15)
will tell you that the integration ofskills is ofparamount importance in
language learning. likewise, assessment is more authentic and provides more
washback when skills are integrated. Nevertheless, the skills are treated
independently here in order to identify prinCiples, test types, tasks, and
issues associated with each one.
Second, you may
already have scanned through this book to look for a chapter on
assessing grammar and vocabulary, or something in the way of a focus on form in
assessment. The treatment of form-focused assessment is not relegated to a
separate chapter here
for a very distinct reason: there is no such thing as a test of grammar
or vocabulary that does not invoke one or more of the separate skills of listening,
speaking, reading, or writing! It's not uncommon to fmd little "grammar
tests" and "vocabulary tests" in textbooks, and these may be
perfectly useful instruments. But responses on these quizzes are usually
written, with multiple-choice selection or ftll·in-the-blank items. In this
book, we treat the various linguistic forms (phonology, morphology, lexicon,
grammar, and discourse) within the context of skill areas. That way,we don't
perpetuate the myth that grammar and vocabulary and other ling1.listic forms
can somehow be disassociated from a mode of performance.
OBSERVING
THE PERFORMANCE OF THE FOUR SKIIS
Before focusing on listening itself, think about the two interacting
concepts of performance
and observation. All language users perform the acts of listening, speaking,
reading, and writing. They of course rely on their underlying competence in
order to accomplish these performances.When you propose to assess someone's
ability in one or a combination of the four skills, you assess that person's
competence, but you observe the person'sperformance. Sometimes the performance
does not indicate true
competence: a bad night's rest, illness, an emotional distraction, test
anxiety, a memory block, or other student-related reliability factors could
affect performance, thereby providing an unreliable measure of actual
competence.
So, one
important principle for assessing a learner's competence is to consider the
fallibility of the results of a single performance, such as that produced in a
test. As with any attempt at measurement, it is your obligation as a teacher to
triangulate your measurements: consider at least two (or more) performances
and/or contexts before drawing a conclusion. That could take the form of one or
more of the following designs:
- Several tests that are combined to form an assessment
- A single test with multiple test tasks to account for learning styles and performance variables.
- In-class and extra-class graded work
- Alternative forms of assessment (e.g., journal, portfolio, conference, obsen:ation, self-assessment, peeT~sessment).
Multiple
measures will always give you a more reliable and valid assessment than a
single measure.
A second principle
is one that we teachers often forget. We must rely as much as possible on observable
performance in our assessments ofstudents. Observable means being
able to see or hear the performance ofthe learner (the senses oftouch, taste,
and smell don't apply very
often to laI?-guage testing!). What, then, is obs~rvable among the four
skills of listening, speaking, reading, and writing? Table 6.1 ·Qffers an
answer.
Isn't it
interesting that in the case of the receptive skills, we can observe neither the
process of performing nor a
product?
I can hear your argument already: "But I can see that she's listening
because she's nodding her head and frowning and smiling and asking relevant
questions." Well, you're not observing the listening performance; you're
observing the result of the listening. You can no more observe listening (or
reading) than you can see the wind blowing.
THE
IMPORTANCE OF liSTENING
Listening has
often played second fiddle to its counterpart~ speaking. In the standardized testing
industry, a number of separate oral production tests are available (fest of Spoken English, Oral
ProfiCiency Inventory, and PhonePass, to name several that are described
Chapter 7 of this book), but it is rare to find just a listening test. One
reason for this emphasis is that listening is often implied as a component of speaking.
How could y~u speak a languag~ without also listening? In addition, the overtly
observable nature of speaking renders it more empirically measurable then listening.
But perhaps a deeper cause lies in universal biases toward speaking. A good
speaker is often (unwisely) valued more highly than a good listener. To
determine ifsomeone is a proficient user of a language, people customarily ask,
"Do you speak: Spanish?" People rarely ask, "Do you understand
and speak Spanish?"
Every teacher
oflanguage knows that one's oral production ability-other than monologues, speeches,
reading alo~d, and the like-is only as good as one's listening comprehension
ability. But of even further impact is the likelihood that input in the aural-oral
mode accounts for a·large proportion of successful language acquisition. In a
typical day, we do measurably more listening than speaking (with the exception of
one or two of your friends who may be nonstop chatterboxes!).Whether in the workplace,
educational, or home contexts, aural comprehension far outstrips oral production
in quantifiab~e terms of time, number of words, effort, and attention.
We therefore
ne.ed·-to pay close attention to listening as a mode of performance for
assessment in the classroom. In this chapter, we will begin with basic
prinCiples and types of listenitig, then move to a survey of tasks that can be
used to assess listening. (For a review of issues in teaching listening, you
may want to read Chapter 16 of TBE)
BASIC
TYPES OF IJSTENING
As with all
effective tests, designing appropriate assessment tasks in listening begins with the specification
of objectives, or criteria. Those objectives may be classified in terms -of
several types of listening performance. Think about what you do when you listen.
Literally in nanoseconds, the following processes flash through your brain:
- You recognize speech sounds and hold ~ temporary "imprint" of them in short-term memory.
- You simultaneously determine the type of speech event (monologue, interpersonal dialogue, transactional dialogue) that is being processed and attend to its context (who the speaker is, location, purpose) and the content of the message.
- You use (bottom-up) linguistic decoding skills and/or (top-down) background schemata to bring a plausible interpretation to the message, and assign a literal and intended meaning to the utterance.
- In most cases (except for repetition tasks, which involve shQrt-term memory only), you delete the exact linguistic form in which the message was originally received in favor of conceptually retaining important or relevant information in long-term memory.
- comprehending ofsurface structure elements such as phonemes,words, intonation, or a grammatical
- categoryunderstanding of pragmatic context
- determining meaning of auditory input
- developing the gist, a global or comprehensive understanding
From these stages we can derive four
commonly identified types of listening performance, each ofwhich comprises a
category within'whiCh t01consider assessment
tasks and procedures.
- Intensive. Listening for perception of the components (phonemes, words, intonation, discourse markers, etc.) of a larger stretch of language.
- Responsive. Listening to a relatively short stretch oflanguage (a greeting, question, command, comprehension check, etc.) in order to make an equally short response.
- Selective. Processing stretches of discourse such as short monologues for several minutes in order to "scan" for certain information.The purpose of such performance is not necessarily to look for global or general meanings, but to be able to comprehend designated information in a context of longer stretches of spoken language (such as classroom directions from a teacher, TV or radio news items, or stories). Assessm<:p.t tasks in selective listening could ask students, for example, to listen for names, numbers, a grammatical category, directions (in a map exercise), or certain facts and events.
- Extensive. Listening to· develop a top-down, global understanding of spoken language. Extensive performance ranges from listening to lengthy lectures to listening to a conversation and deriving a comprehensive message or purpose. Listening for the gist, for the main idea, and making inferences are all part of extensive listening.
xMICRO-
AND MACROSKII.lS OF LISTENING
A us ful way of synthesizing
the above two lists is to consider a finite number of micro- and macroskills
implied in the performance of listening comprehension. Richards'
(1983) list of microskills has proven useful in the domain of specifying objectives
for learning and may be even more useful in forcing test makers to carefully
identify specific assessment objectives. In the following box, the skills are
subdivided into what I prefer to think of as microskills (attending to the
smaller bits and chunks of language, in more of a bottom-up process) and
macroskills (focusing on the larger elements involved in a top-down approach to
a listening task). The microand macros kills provide 17 different objectives to
assess in listening.
Implied in the
taxonomy above is a notion of what makes many aspects of listening difficult,
or why listening is not simply a linear process of recording strings of
language as they are transmitted into our brains. Developing a sense ofwhich
aspects of listening performance are predictably difficult will help you to
challenge your students appropriately and to assign weights to items. Consider
the following list of what makes listening
difficult (adapted from Rich~ds, 1983; Dr, 1984; Dunkel, 1991):
- Clustering: attending to appropriate "chunks" of language-phrases, clauses, constituent.
- Redundancy: recognizing the kinds of repetitions, rephrasing, elaborations, and insertions that unrehearsed spoken language often contains, and benefiting from that recognition
- Reduced fonns: understanding the reduced fo~ms that may not have been a part,of an English learner's past learning experiences in classes where only formal "textbook" language has been presented.
- Perjonnance variables: being able to "weed out" heSitations, false starts, pauses, and corrections innat~ speech.
- Colloquial language: comprehending idioms, slang, reduced forms, shared cultural knowledge
- Rate ofdelivery: keeping up with the speed of delivery, processing automatically as the speaker continues
- Stress, rhythm, and intonation: correctly understanding prosodic elements of spoken language, which is almost always much more difficult than understanding the smaller phonological bits and pieces.
- Interaction: managing the interactive flow of language from listening to speaking to listening, etc.
DESIGNING
ASSESSMENT TASKS: INTENSIVE LISTENING
Once you have
determined objectives, your next step is to design the tasks,
including making decisions about how you
will elicit performance and how you will' expect the test-taker to respond. We
will look at tasks that range from intensive listening performance, such as
minimal phonemiC pair recognition, to extensive comprehension of language in
communicative contexts. The focus in this section is on the fllicroskills of
intensive listening.
DESIGNING
ASSESSMENT TASKS: RESPONSIVE liSTENING
A
question-and-answer format can provide some interactivity in these lower-end listening tasks. The
test-taker's response is the appropriate answer to a question. Appropriate
response to a question Test-takers hear: How much time did you take to do your
homework?
Test-takers read: (a) In about an hour.
(b) About an hour. (c) About $10. (d) Yes, I did.
The objective of
this item is recognition of the wh-question bow much and its appropriate response.
Distractors are chosen to repres’nt
common learner errors:
(a) responding to how much vs. how much
longer; (c) confusing how much in reference to time vs. the more frequent
reference to money; (d) confusing a wb-question with a yes/no question.
None of the
tasks so far discussed have to be framed in a multiple-choice format. They can
be offered in a more open-ended framework in which test-takers write or speak
the response',The above item would then look like this: Open-ended response to
a question
Test-takers hear: How much time did you
take to do your homework? Test-takers write or speak: If open-ended response
formats gain a small amount of authenticity and creativity,they of course
suffer some in their practicality, as teachers must then read students' responses
and judge their appropriateness, which takes time.
DESIGNING
ASSESSMENT TASKS: SELECTIVE IlSTENING
A third type of
listening performance is selective listening, in-which the test-taker listens to a limited
quantity of aural input and must discern within it some specific information. A number
of techniques have been used 'that require selective listening.
ListeningCloze
Listening cloze
tasks (sometit11es called cloze dict~tions or partial dictations) require the
test-taker to listen to a story. fllonologue,or conversation and simultaneously read the written text in which
selected words or phrases have been deleted. Cloze procedure is most commonly
associated with reading only In its generic form, the test consists of a
passage in which every nth word (typically every seventh word) is deleted and
the test-taker is asked to. supply an appropriate word. In a listening cloze
task, te~t-takers see a transcript of the passage that they are listening to and
flU· in the blanks with the words or phrases that they hear.
One pOlential
weakness of listening cloze techniques is that they may simply become reading
comprehen~ion tasks. Test-takers who are asked to listen to a story with
periodic deletions in the
written version may not need to listen at all, yet may still be able to respond with the appropriate
word or phrase. You can guard against this
eventuality if the blanks are items with high information load that cannot be easily
predicted simply by reading the passage. In the example below (adapted from Bailey,
1998, p. 16), suc~ a shortcoming was avoided by only the criterion of Numbers.
Other listening
cloze tasks may focus Qn a
gategory
such as verb tenses, articles, two-word verbs, prepositions, or transition
words/phrases. Notice two important structural differences between listening
cloze tasks and standard reading cloze. In a
listening cloze, deletions are governed by the objective of the test,
not by mathematical deletion of every nth word; and more than one word may be
deleted, as in the above example.
Listening cloze
tasks should normally use an exact word method of scoring, in which you accept
as PQnse only the, actual word or phrase that was spoken and consider other
appropriate words as incorrect. (See Chapter 8 for further discussion of these
two methods.) Such stringency is warranted; your objective is, after all, to
test listening comprehenSion, not grammatical or lexical expectancies.
Information
Transfer
Selective listening can also be
assessed through an infor:mation transfer technique in which aurally processed
information must be transferred to a visual representation, such as labeling a
diagram, ideniifying an element in a picture, completing a form, or showing
routes on a map.
At the lower end
of the scale of linguistic complexity, simple picture-cued items are sometimes
efficient rubrics for assessing certain selected information.
Sentence
Repetition
The task of
simply repeating a sentence or a partial sentence, or sentence repetition, is
also used as an assessment of listening comprehension. As in a dictation
(discussed below), the test-taker must retain a stretch of language long enough
to reproduce it. and then' must respond with an oral repetition of that
stimulus. Incorrect listening comprehension, whether at the phonemic or
discourse level, may be manifested in the correctness of the repetition. A
miscue in repetition is scored as a miscue in listening. In the case ofsomewhat
longer sentences, one could argue that the ability to recognize and retain
chunks of language as well as threads of meaning might be assessed through
repetition. In Chapter 7, we will look closely at PhonePass, a commercially
produced test that relies largely on sentence repetition to assess both oral
production and listening comprehension.
Sentence
repet~tion is far from a flawless listening assessment task. Buck (2001,p.79)
noted that such tasks "are not just tests of listening, but tests of
gc;.neral oral skills." Further, this task may test only recognition
ofsounds, and it can easily be contaminated by lack ofshort-term ~emory
ability, thus invalidating it as an assessment of comprehension alone. And the
teacher may never be able to distinguish a listening comprehension error from
an oral production ~rror. Therefore,sentence repetition tasks should be used
with caution.
DESIGNING
ASSESSMENT TASKS: EXTENSIVE LISTENING
Drawing a clear
distinction between any two of the categories of listening referred to here is
problematic, but perhaps the fuzziest division is between selective and extensive
listening. As we gradually move along the continuum from smaller to larger
stretches of language, and from micro- to macroskills of listening, the
probability of using more
extensiveJistening_tasks_jrrcl"eases.
Some important questions about designing assessments at this level emerge.
- Can listening performance be distinguished from cognitive processing factors such as memory, associations, storage, and recall?
- As assessment procedures become more communicative, does the task take into account test-takers' ability to use grammatical expectancies, lexical collocations, semantic interpretations, and pragmatic competence?
- Are test tasks themselves correspondingly content valid and authentic-that is, do they mirror real-world language and context?\
- As assessment tasks beco~e more and more open-ended, they more closely resemble pedagogical tasks, which leads one to ask what the difference is between assessment and teaching tasks. The answer is scoring: the former imply specified scoring procedures, while the latter do not.
We will try to
address these questions as ,ve look at a number of extensive or quasie}tensive
listening comprehension tasks.
Dictation
Dictation is a
widely researched genre of assessing listenit:lg comprehension. In a dictation, test-takers
hear a passage, typically of 50 to 100 words, recited three times: first, at
normal speed; then, with long pauses between phrases or natural word groups,
during which time test-takers write down what they have just heard; and finally,
at normal speed once more so they can check their work and proofread. Here is a
sample dictation at the intermediate level of English.
Dictations have
been used as assessment tools for decades. Some readers still cringe at the
thought of having to render a correctly spelled, verbatim version of a
paragraph or story recited by the teacher. Until research on integrative
testing was published (see Oller, 1971), dictations were thought to be not much
more than glorified spelling tests. However, the required integration of
listening and writing in a dictation, along with its presupposed knowledge of
grammatical and discourse expectancies, brought this technique back into
vogue..;. Hughes (1989), Cohen (1994), Bailey (1998), and Buck (2001) all
defend the plausibility of dictation as an integrative test that requires some
sophistication in the language in order to process and write down all segments
correctly. Thus, I include dictation here under the rubric of extensive tasks,
although I am more conlfortable with labeling it quasi extensive.
The difficulty
of a dictation task can be easily manipulated by the length of the word groups
(or bursts, as they are technically called), the length of the pauses, the speed
at which the text is read, and the complexity of the discourse, grammar, and vocabulary
used in the passage. Scoring
is another matter. Depending on your context and purpose in administering a
dictation, you will need to decide on.
scoring criteria for several possible kinds
of errors:
- spelling error only, ,but the word appears to have been heard correctly
- spelling 'and/or obvious misrepresentation of a word, illegible word
- grammatical error (For example, test-taker hears I can~t do it, writes I can do it.)
- skipped word or phrase
- permutation of words
- additional words not in
the original
- replacement of a word
with an appropriate synonym
Dictation seems
to provide a reasonably valid method for integrating listening and writing
skills and for tapping into the cohesive elements of language implied in short
passages. However, a word of caution lest you assume that dictation provides a
quick and easy method of assessing extensivelisterung comprehens~on. If the bursts
in a dictation are relatively long (more than five-word segments), this method places
a certain amount ofload on memory and processing ofmeaning (Buck, 2001, p. 78).
But only a moderate degree of cognitive processing is required, and claiming that
dictation fully assesses the ability to comprehend pragmatic or illocutionar elements'of language,
context, inference, or senlantics may be going too. far._Finally, one can
easily question the authenticity of dictation: it is rare in the real world for
people to write down more than a few chunks of information (addresses, phone numbers, grocery lists,
directions, for example) at a time.
Despite these
disadvantages, the practicality of the administration of dictations, a moderate degree of
reliability in a well-established scoring system, and a strong correspondence to other
language abilities speaks well for the inclusion of dictation among the possibilities for
assessing extensive (or quasi-extensive) listening comprehension.
Communicative
Stimulus-Response Tasks
Another-and more
authentic-example of extensive listening is found in a popular genre of
assessment. task in which the test-taker is presented with a stimulus monologue
or conversation and then is asked to respond to a set of compreh~slions.
sucntiSki--(as you saw in
Chapter 4 in the discussion of standardized testing) are
corrimonly used i.fl commercially produced profiCiency tests. The monologues, lectures.
and brief conversations used in such tasks are sometimes a little contrived, and
certainly the subsequent multiple-choice questions don't mirror communicative, real-life
situations. But with some care and creativity, one can create reasonably authentic
stimuli, and in some rare cases the response mode (as shown in one example
below) actually approaches complete authenticity. Here is a typical example of
such a task.
Authentic
Listening Tasks
Ideally, the
language assessment field would have a stockpile of listening test types that are cognitively
demanding. communicative, and authentic, not to mention interactive by means of
an integration with speaking. However, the nature of a test as a of performance
and a set of tasks with limited time frames implies an equally linlited
capacity to mirror all the real-world contexts of listening perfonnance.
"There is no such thing as a communicativet," stated Buck (200 1, p.
92). "Every test requires some comPofieiits-oicomIDiiiifcatlve language
ability, and no test covers them all.Similarly, with the notion of
authenticity, every task shares some characteristics with target-language
tasks, and no test is completely authentic."
Beyond the
rubrics of intensive, responsive, selective, and quasi-extensive communicative
contexts described above, can we assess aural comprehension in a truly communicative context?
Can we, at this end of the range of listening tasks, ascertain from test-takers
that they have processed the main idea(s) of a lecture, the gist of a story,
the pragmatics of a conversation, or the unspoken inferential data present in most
authentic aural input? Can we assess a test-taker's comprehension of humor, idiom,
and metaphor? The answer is a cautious yes, but not without some concessions to
practicality.· And the answer is a more certain yes if we take the liberty of stretching
the concept of assessment to extend beyond tests and into a broader framework
of tna!iy Here are some
possibilities.
- Note-taking. In the academic world, classroom lectures by professors are common features of a non-native English-user's experience. One form of a midterm examination at the American Language Institute at San Francisco State University (Kahn, 2002) uses a IS-minute lecture as a stimulus. One among several response formats includes note-taldng by the test-takers. These notes are evaluated by the teacher on a 30-point system, as follows:
- Editing. Another authentic task provides both a written and a spoken stimulus, and requires the test-taker to listen for discrepancies. Scoring achieves relatively high reliability as there are usually a small number of specific differences that must be identified. Here is the way the task proceed.
3. Interpretive tasks. One of the intensive listening
tasks described above was
paraphrasing
a story or conversation. An interpretive task extends the stimulus material to
a longer stretch of discourse and forces the test-taker to infer a response.
Potential stimuli include
- · song lyrics,
- [recited] poetry, .
- ·
radio/television news
reports, and
- ·
an oral account of an
experience.
Test-takers are then directed to
interpret the stimulus by answering a few questions
(in open-ended form). Questions might
be:
·
"Why was the
Singer feeling sad?"
·
"What events might
have led up to the reciting of this poem?"
·
"What do you think the political
activists might do next, and why?"
·
"What do you think
the storyteller felt about the mystorious disappearance of her necklace?"
This kind of
task moves us away from what might traditionally be considered a test
toward an informal assessment, or
possibly even a pedagogical technique or activity. But the task conforms to
certain time limitations, and the questions can be quite specific, even though
they ask the test-taker to use inference.\Ylhile reliable scori.tlg may be an
issue (there may be more than one correct interpretation), the authenticity of
the interaction in this task and potential washback tothe student surely give
it some prominence among communicative assessment procedures.
4. Retelling. In a related task, test-takers
listen to a story or news event and simply retell it, or summarize it, either
orally (on an audiotape) or in writing. In so doing, test-takers must identify
the gist, main idea, purpose, supporting points, and/or conclusion to show full
comprehension. Scorillg is partially predetermined by specifying a minimu
number of elements that must appear in the retelling. Again reliability may
suffer, and the time and effort needed to read and evaluate the response lowers
practicality. Validity, cognitive processing, communicative ability, and
authenticity are all well incorporated into the task.
A ftfth category
of listening comprehension was hinted at earlier in the chapter: interactive
listening. Because such interaction presupposes a process of speaking in
concert with listening, the interactive nature of listening will be addressed
in the next chapter. Don't forget that a significant proportion of realworld
listening performance is interactive. With the exception of media input, speeches,
lectures, and eavesdropping, many of our listening efforts are directed toward
a two-way process of speaking and listening in face-ta-face conversations.
EXERCISES
[Note: (I)
Individual work; (G) Group or pair work; (C) Whole-class discussion.]
- (C) In Table 6.1 on page 118, it is noted that one cannot actually observe listening and reading performance. Do you agree?-:Anddo you agree that there isn't even a product to observe for speaking, listening, and reading? How, then, can one infer the competence of a test-taker to speak, listen, and read a language?
- (C) Given that we spend much more time listening than we do speaking, why are there many more tests of speaki1l:g than listeninG.
- (G) Look at the list of micro- and macroskills of listening on pages 121-122. In pairs, each assigned to a different skill (or two), brainstorm some tasks that assess those skills. Present your fmdings to the rest ot-the class.
- (G) Eight characteristics of listening that make listening "difficult" are listed on page 122. In pairs, each asSigned to an assessment task itemized in this chapter, decide which of the eight factors, in order of significance; contribute to the potential difficulty of the items. Report back to the class.
- (G) Divide the basic types of listening among groups or pairs, one type for each. Look at the sample assessment teclmiques provided and evaluate them according the five principles (practicality, reliability, validity [face and content], authenticity, and washback). Present your critique to the rest of the class.
- (G) In the same groups as in #5 above and with the same type of listening, design some other item types, different from the one(s) provided here, that assess the same type of listening performance.
- (G) With a linguistic objective assigned to each pair or group, construct a listening cloze test for two-word verbs, verb tenses, prepositions, transition words, articles, and/or other grammatical categorieS,
- (I/C) On page 131, you are reminded that dictations are considered by some assessment specialists to be integrative (requiring the integration of listening, writing, reading [proofreading], along with attendant grammatical and discourse abilities). Is this a valid claim? Justify your response.
FOR
YOUR FURTIlER READING
Buck, Gary. (2001). Assesstng listening.
Cambridge: Cambridge University Press.
One of a series
of very useful ref~rence books on assessing specific skill areas published by Cambridge
University Press, this. one gives an overview of research and pedagogy on
listening comprehension and demonstrates many different assessment procedures in common
use.
Richards, Jack C. (1983). Listening
comprehension: Approach, design, procedure.
TESOL Quarterly,
17, 219-239.
Even though
Richards published this article in 1983, it still provides a standard backdrop for teaching
listening skills. While formal assessment is not directly addressed, informal
assessmeht is implied in its pedagogical focus on practical classroom
techniques.
Mendelsohn, David J. (1998). Teaching
listening. Annual Review of Applied
Linguistics, 18,
81-101.
Mendelsohn's
overview of research ort teaching 1i~tening proVides an excellent foundation for
understanding assessment tasks. Me focuses on a strategy based approach to
teaching listening and adds an annotated bibliography of professional resource
books.
ASSESSIIG
SPEAKING
From a pragmatic
view of language performance, listening and speaking are almost always closely
interrelated. While it is possible to isolate some listening performance types
(see Chap"ter 6),'it is very difficult to isolate oral-production tasks
that do not directly involve the interaction of aural comprehension. Only in
limited contexts of speaking (monologues, speeches, or telling a story and
reading aloud) can we assess oral language without the aural participation of
an interlocutor.
While speaking
is a productive skill that can be directly and empirically observed, those
observations are invariably colored by the accuracy and effectiveness Of a
test-ta}{er'$ list~I1iI1g skill, which necessarily
compromisestlieh"rellability
and
validity of an oral production test. How do you know for certain that a
speaking score is exclusively a measure of oral production without the
potentially frequent clarifications of an interlocutor? TItis interaction of
speaking and listening challenges the designer of an oral production test to
tease apart, as much as possible, the factors accounted for by aural intake.
Another challenge is the
design of elicitation techniques. Because most speaking is the product of
creative construction oflinguistic strings, the speaker makes choices of
leXicblf,srructure, and discourse:-If-your-goal is to-have test-takers demonstrate
certain spoken grammatical categories, for example, the stimulus you design
must elicit those grammatical categories in ways that prohibit the test-taker from
avoiding or paraphrasing and thereby dodging production of the target form.
All of these
issues will be addressed in this chapter as we review types of spoken language
and micro- and macros kills of speaking, then outline numerous tas~s for
assessing speaking.
BASIC
TYPES OF SPEAKING
In Chapter 6, we
cited four categories of listening performance assessment tasks. A similar
taxonomy emerges for oral production.
- bnitative. At one end of a continuum of types of speaking performance is the ability to simply parrot back (imitate) a word or phrase or possibly a sentence. While this is a purely phonetic level of oral production, a number of prosodiC, lexical, and grammatical properties of language may be included in the criterion performance.We are interested only in what is traditionally labeled "pronunciation"; no inferences are made about the test-taker's ability to understand or convey meaning or to participate in an interactive conversation. The only role of listening here is in the short-term storage of a ptonlpt, just long enough to, allow the speaker to retain the short stretch of language that must be imitated.
- Intensive. A second type of speaking frequently employed in assessment contexts is the production of short stretches of oral language designed to demonstrate competence in a narrow band of grammatical, phrasal, lexical, or phonological relationships (such as prosodic elements-intonation, stress, rhythm, juncture). The speaker must be aware of semantic properties in order to be able to respond, but interaction with an interlocutor or test administrator is minimal at best. Examples of intensive assessment tasks include directed response tasks, reading aloud, sentence and dialogue completion; limited picture-cued tasks ill:~luding simple sequences; and translation up to the simple Sentence level.
- Responsive. ,Responsive assessment tasks include interaction and test com- v prehension but at the somewhat limited level of very short conversations, standard greetings and small talk, simple requests and comments, and the liken
- Interactive. The difference between responsive and interactive" speaking is in the length and complexity of the interaction, which sometimes includes mUltiple exchanges and/or multiple participants. Interaction can take the two forms of transactional language, which has the purpose of exchanging specific information, or interpersonal exchanges, which have the purpose of maintaining social relationships. (In tfie three dialogues cited above, A and B were transactional, and C was interpersonal.) In interpersonal exchanges, oral production can become pragmatically complex with the need to speak in a casual register and use colloquial language, ellipSis, slang, humor, and other sociolinguistic conventions.
- Extensive (monologue). Extensive oral production tasks include speeches, oral presentations, and story-telling, during which the opportunity for oral interaction from listeners is either highly limited (perhaps to nonverbal responses) or ruled out altogether. Language style is frequently more deliberative (planning is involved) and" formal for extensive tasks, but we cannot rule out certain informal monologues" such as casually delivered spe"ech (for exatPple, my vacation in the mountains, a recipe for outstanding pasta primavera, recounting the plot of a novel or movie).
5.
MICRO-
AND MACROSKUJS OF SPEAKING
In Chapter 6, a
list of listening micro- and macroskills enumerated the various components of
listening that make up criteria for assessment. A similar list of speaking
skills can be drawn up for the same purpose: to serve as a taxonomy of skills
from which you 'will select one or several that will become the objective(s) of
an assessment task. The microskills refer to producing the smaller chunks of
language such as phonemes, mofQ!!emes, words, collocations, and phrasal units.
The macroskills imply thespeaker's focus on die-larger eiements: flu~ng,
dis-course, function, style, cohesion, nonverbal c?mmunication, and strategic.
? Ptions.
The niicro-and macros kills total roughly 16 different objectives to assess in
speaking.
There is such an
array of oral production tasks that a complete treatment is almost impossible
within the confines of one chapter in this book. Below is a consideration of
the most common techniques with brief allusions to related tasks. As already
noted in the introduction to this chapter, consider three important issues as
you set out to design tasks:
1.No
speaking task is capable of isolating the single skill of'oral production.
Concurrent involvement of the additional performance of aural comprehension,
and possibly reading, is
usually necessary.
2.EliCiting
the specific criterion you have designated for a task can be tricky because
beyond the word level, spoken language offers a number of productive options to
test-takers.
3.Because
of the above two characteristics of oral production assessment, it is important
to carefully specify scoring procedures for a response so that ultimately you
achieve as high a reliability index as possibie.
DESIGNING
ASSESSMENT TASKS: IMITATIVE SPEAKING
You may be
surprised to see the inclusion ofsimple phonological imitation in a
consideration of assessment of oral production. After all, endless repeating of
words, phrases, and sentences was the province of the long-since-discarded
Audiolingual Method, and in an era of communicative language teaching, many
believe that nonmeaningful imitation ofsounds is fruitless. Such opinions-have
faded in recentyears as we discovered that an overemphasis on fluency can
sometimes lead to the decline of accuracy in speech. And so we have been paying
more attention to pronunciation, especially' suprasegmentals, in an attempt to
help learners be more comprehensible.
An occasional
phonologically focused repetition task is warranted as long as repetition tasks are
not allowed to occupy a dominant role in an overall oral production assessment, and as long
as you artfully avoid a negative washback effect. Such tasks range from word
level to sentence level, usually with each item focusing on. a specific
phonological criterion. In a simple repetition task, test-takers repeat the
stimulus, whether it is a pair ofwords, a sentence, or perhaps a question (to
test for intonation production).
PHONEPASS®
TEST
An example of a
popular test that uses imitative (as well as intensive) production tasks is
PhonePass, a widely used, commercially available speaking test in many
countries. Among a number of speaJdng tasks on the test, repetition of
sentences (of 8 to 12 words) occupies a prominent role. It is remarkable that
research on the PhonePass test has supported the construct validity of its
repetition tasks not just for a testtaker's phonological ability but also for
discourse and overall oral production ability (fownshend et al., 1998;
Bernstein et aI., 2000; Cascallar & Bernstein, 2000).
The PhonePass
test elicits computer-assisted oral production over a telephor~,e. Test-takers.
read aloud, repeat sentences, say words, and answer questions. With a downloadable
test sheet as a reference, test-takers are directed to telephone a designated
number and listen for
directions. The test has five sections.
DESIGNING
ASSESSMENT TASKS: INTENSIVE SPEAKING
At the intensive
level, test-takers are prompted to produce short stretches of discourse (no
more than a sentence) through which they demonstrate linguistic ability at a
specified level of language. Many tasks are "cued" tasks in that they
lead the testtaker into a narrow band of possibilities.
Parts C and D of
the PhonePass test fulfill the criteria of intensive tasks as they elicit
certain expected forms of language. Antonyms like high and low, happy and sad
are prompted
so that the, automated scoring mechanism anticipates only one word. The
either/or task of Part D fulfills the same criterion. Intensive tasks may also be
described as limited response tasks (Madsen, 1983), or mechanical tasks (Underhill,
1987), or what classroom pedagogy would label as controlled responses.
Directed
Response Tasks
In this type of
task, the test administrator elicits a particular grammatical form or a transformation of a
sentence. Such tasks are clearly mechanical and not communicative, but they do
require minimal processing ofmeaning in order to produce the correct
grammatical output.
Read-Aloud
Tasks
Intensive
reading-aloud tasks include reading beyond the sentence level up to a paragraph or two. This
technique is easily administered by selecting a passage that incorporates test specs
and by recording the test-taker's output; the scoring is relatively easy because all of the
test~taker's oral production is controlled. Because of theresults of research
on the PhonePass test, reading aloud may actually be a surprisingly strong
indicator of overall oral production ability.
For many
decades, foreign language programs have used reading passages to analyze oral
production. Prator's (1972) Manual ofAmerican English Pronunciation included a
"diagnostic passage" of about 150 words that students could read
aloud into a tape recOrder. Teachers listening to the recording would then rate
students on a number of phonological factors (vowels, diphthongs, consonants,
consonant clusters, stress, and intonation) by completing a two-page diagnostic
checklist on which all errors or questionable items were noted. These
checklists ostensibly offered direction to the teacher for emphases in the
course to come.
An earlier form
of the Test of Spoken English (fSE®, see below) incorporated
one read-aloud passage .of about 120 to
130 word Such
a rating list does not indicate how to gauge intelligibility, which is
mentioned in both lists. Such slippery terms remind us that oral production
scoring, even with the controls that reading aloud offers, is still an inexact
science. Underhill (1987, pp. 77-78) suggested some variations on the task of
simply reading a short passage:
- reading a scripted dialogue, with someone else reading the other part
- reading sentences containing minimal pairs, for example:
Try not to heat/hit the
pan too much.
The doctor gave me a
bil/pill.
- · reading information
from a table or chart
If reading aloud
shows certain practical advantages (predictable output, practicality,
reliability in scoring), th~re are several drawbacks to using this technique
for ass~ssing oral production. Reading aloud is somewhat "!!!e:Y1h,!;n!~s,.in
that we seldom read anything aloud to ·semeone else in the- real world,
with--the .exception of a parent reading to a child, occasionally sharing a written
story with someone, or giving a scripted oral presentation. Also, reading aloud
calls on certain specialized oral abilities that may not indicate one's
pragmatic ability to communicate orally ill face-ta-face contexts. You should
therefore employ this technique with some caution, and certainly supplement it
as an assessment task with other, more communicative procedures.
Sentence/Dialogue
Completion Tasks and Oral Questionnaires
Another
technique for targeting intensive aspects of language requires test-takers to read dialogue in which
one speaker's lines have been omitted. Test-takers are first given time to read
through the dialogue to get its gist and to think about appropriate lines to
fill in. Then as the tape, teacher, or test administrator produces .
Picture-Cued
Tasks
One of the more
popular ways to elicit oral language performance at both intensive and
extensive levels is a pictl1re-cued stimulus that requires a description from
the testtaker. Pictures
may be very simple, designed to elicit a word or a phrase; somewhat more
elaborate and "busy"; or composed of a series that tells a story or
incident. Here is an example of a picture-cued elicitation of the production of
a simple minimal pair.
Translation
(of T,imited Stretches of Discourse)
Translation is a
part of our tradition in language teaching that we tend to discount or disdain,
if only because our current pedagogical stance plays down its importance.
Translation methods of teaching are certainly passe in an era of direct approaches
to creating communicative classrooms. But we should remember that in countries
where English is not the native or prevailing language, translation is a meaningful
communicative device in contexts where the English user is. called on to be an
interpreter. Also, translation is a well-proven communication strategy for learners
of a second language.
Under certain
constraints, then, it is not far-fetched to suggest translation as a device to
check oral production. Instead of offering pictures or written stimuli, the test-taker
is given a native language word, phrase, or sentence and is asked to translate
it. Conditions may vary from expecting an instant translation of an orally
elicited linguistic target toallowmg more thinking time before producing a
translation of somewhat longer texts, which may option3.ny be offered to the
test-taker in written form. (franslation of extensive, texts is discussed at the
end of this chapter.) As an assessment procedure, the advantages of translation
lie in its control of the output ofthe test-taker, whichof,:ot;U,"se means
that scoring is more easily specified.
DESIGNING
ASSESSMENT TASKS: RESPONSIVE SPEAKING
Assessment of
responsive tasks involves brief interactions with an interlocutor, d~ffering
from intensive tasks in the increased creativity given to the test-taker and from
interactive tasks by the somewhat limited length of utterances.
Questi!?n
and Answer
Question-and-answer
tasks can consist of one or two questions from an interviewer,or they can make
up a portion of a whole battery of questions and prompts in an oral interview.
They can vary from simple questions like "What is this called in English?"
to complex questions like "What are the steps governments should take, if any,
to stem the rate of deforestation in tropical countries?" The first
question is intensive in its purpose; it is a display question intended to
elicit a predetermined correct response. We have already looked at some of these
types ot questions in the previous section. Questions at the responsive level tend
to be genuine referential questions in which the test-taker is given more opportunity
to·produce meaningful language in response.
Giving
Instructions and Directions
We are all
called on in our dally routines to read instructions on how to operate an appliance, how to put a
bookshelf together, or how to create a delicious clam. Somewhat less frequent
is the mandate to provide such instructions orally, but this speech act is
still relatively common. Using such a stimulus in an assessment context
provides an opportunity for the test-taker to engage in a relatively extended
stretch of discourse, to be very clear and specific, and to use appropriate
discourse markers and connectors. The technique is Simple: the administrator poses
the problem, and the test-taker responds. Scoring is based primarily on
comprehensibility and s~condari1y on other specified grammatical or discourse
categOries. Here are some possibilities.
Paraphrasing
Another type of
assessment task that can be categorized as responsive asks the testtaker to
read or hear a limited number of sentences (perhaps two to five) and-produce a
paraphrase of the sentence.
TEST
OF SPOKEN ENGLISH (TSE@)
Somewhere
straddling responsive, interactive, and extensive speaking tasks lies another popular
commercial oral production assessment, theTest of Spoken English (TSE)'. The
TSE is a 20-minute audiotaped test of oral language ability within anacademic
or professional environment. TSE scores are used by many North American
institutions of higher education to select international teaching assistants. The
scores are also used for selecting and certifying health professionals such as physicians,
nurses, pharmacists, physical therapists, and veterinarians.
The tasks on
theTSE are designed to elicit oral production in various discourse categories rather than
in selected phQ~ogical, ~a,ticalt or l~cal !argets. The follOwing content
specifications for the TSE represent the discourse and pragmatic contexts
assessed in each administration:
1. Describe something physical.
2. Narrate from presented material.
3. Summarize information of the
speaker's own choice.
4. Give directions based on visual
materials.
5. Give instructions.
6. Give an opinion.
7. Support an. opinion.
8. Compare/contrast.
9. Hypothesize.
10. Function "interactively."
11. Define.
Using these
specifications, Lazaraton andWagner (1996) examined 15 different specific tasks
in collecting background data from native and non-native speakers of English.
a. giving
a personal deSCription
b. describing
a daily routine
c. suggesting
a gift and supporting one's choice
d. recommending
a place to visit and supporting one's choice
e. giving
directions
f. describing
a favorite movie and supporting one's choice
g. telling
a story from pictures
h. hypothesizing
about future action
i. hypothesizing
about a preventative action
j.
making a telephone call to the dry cleaner
k.describing
an important news event
From their
fmdings, the researchers were able to report on the validity of the tasks,
especially the match between the
intended task functions and the actual outpu:: of
both native and non-native speakers.
DESIGNING
ASSESSMENT TASKS: INTERACTIVE SPEAKING
The fmal two
categories of oral production assessment (interactive and extensive speaking) include tasks
that involve relatively long stretches of interactive discourse (interviews, role
plays, discussions, games) and tasks. of equally long duration but that.
involve less interaction (speeches,
telling longer stories, and extended explanations and translations).The obvious
difference between the two sets of tasks is the degree of interaction with'an
interlocutor. Also, interactive tasks are what some would describe as
interpersonal, while the fmal category includes more transactional speech
events.
Interview
When "oral
production assessment" is mentioned, the first thing that comes to mind is
an oral interview: a test administrator and a test-taker sit downJn a direct
face-toface exchange and proceed through a protocol of questiof!s and
directives. The interview, which may be tape-recorded for re-listening, is then
scored on one or more parameters such as accuracy in pronunciation and/or
grammar, vocabulary usage, fluency,
sociolinguistic/pragmatic
appropriateness, task accomplishment, and even comprehension. Interviews can vary in
length from perhaps five to forty-five minutes, depending on their purpose and
context. Placement interviews, designed to get a quick spoken sample from a
student in order to verify placement into a course.
Every effective
interview contains a number of nlandatory stages. 1W'0 decades ago, Michael
Canale (1984) proposed a framework for oral proficiency testing that has withstgod
the test of time. He . suggested that test-takers will perform at their best if
they are led through
four stages:
- Warm-up. In a minute or so of preliminary small talk, the interviewer directs mutual introductions, helps the test-taker become comfortable with the situation, apprises the test-taker of the format,and allays anxieties. No scoring of this phase takes place.
- Level check. Through a series of preplanned questions, the interviewer stimulates the test-taker to respond using -expected or predicted forms and functions. If, for example, from previous test information, grades, or other data, the test..taker has been judged to be a "Level 2" (see below) speaker, the interviewer'S prompts will attempt to confirm this assumption. The responses may take very simple or very complex form, depending on the entry level of the learner. Questions are usually designed to elicit grammatical categories (such as past tense or subject-verb agreement), discourse structure (a sequence of events), vocabulary usage, and/or sociolinguistic factors (politeness conventions, formal/informal language.
- Probe. Probe questions and prompts challenge test-takers to go to the heights of their ability, to extend beyond the limits ofthe interviewer'S expectation through increaSingly difficult questions. Probe questions may be complex in their framing and/or complex in their cognitive and linguistic demand. Through probe items, the interviewer discovers the ceiling or limitation of the test-taker's proficiency. This need not be a separate stage entirely, but might be a set of questions that are interspersed into the previous stage.
- Wind-down. This fmal phase of the interview is simply a short period of time during which the interviewer encourages the test-taker to relax with some easy questions, sets the test-laker's mind at ease, and provides information about when and where to obtain the results of the interview. This part is not scored.
The success of
an oral interview will depend on
- clearly specifying administrative procedures of the assessment (practicality),
- focusing the questions and probes on the purpose of the assessment (validity),
- appropriately eliciting an optimal amount and quality of oral production from the test~taker (biased for best performance), and
- creating a conSistent, workable scoring system (reliability).
This last issue
is the thorniest. In oral production tasks that are open-ended and that involve a significant
level of interaction, the interviewer is forced to make judgments that are
susceptible to some unreliability. Through experience, training, and careful
attention to the linguistic criteria being assessed, the ability to make su~h judgments
accurately will be acquired. InTable 7.2, a set of deSCriptions is given for scoring
open-ended oral interviews. These descriptions come from an earlier version of
the Oral Proficiency Interview and are useful for classroom purposes.
The test
administrator's .challenge is to assign a score, ranging from 1 to 5, for each
of the six categories indicated above. It may look easy to do, but in reality
the lines of distinction between levels is quite difficult to pinpoint. Some
training or at least a good deal of interviewing experience is required to make
accurate ·aSsessments of oral production in the six categories. Usually the six
scores are then amalgamated into one holistic score, a process that might not
be relegated to a simple mathematical average if you wish to put more weight on
some categories than you do on others.
This five-point
scale, once known as "FSI levels" (because they were frrst advocated
by the Foreign Service Institute in Washington, D.C.), is still in popular use
among U.S. government foreign service staff for designating profiCiency in a
foreign language.
A variation on
the usual one-on-one format with one intervie.wer and one testtaker is to place
two test-takers at a time with the interviewer. An advantage of a two-on-one
interview is the practicality of scheduling twice as rrrany candidates in the same
time frame, but more significant is the opportunity for student-student interaction.
By deftly posing questions, problems, and role plays, the interviewer can the
output of the test-takers while lessening the need for his or her own output. A
further benefit is the probable increase in authenticity when two testtakers
can actually converse with each other. Disadvantages are equalizing the output
between the two test-takers, discerning the interaction effect of unequal comprehension
and production abilities, and scoring two people simultaneously.
Discussions
and Conversations
As formal
assessment devices, discussions and conversations with and among students are
difficult to specify and even more difficult to score. But as informal
techniques to assess learners, they offer a level of authenticity and
spontaneity that other assessment techniques may not provide. Discussions may
be especially appropriate tasks through which to elicit and observe such
abilities as
§ topic
nomination, maintenance, and termination;
§ attention
getting, interrupting, floor holding, control;
§ clarifying, questioning, paraphrasing;
§ comprehension Signals (nodding,
"uh-huh,""hmm," etc.);
§ negotiating meaning;
§ intonation
patterns for pragmatic effect;
§ kinesics,
. eye contact, proxemics, body language; and
§ politeness, formality, and other
SOCiolinguistic factors.
AsseSSing the
performance of participants through scores or checklists (in which appropriate
or inappropriate manifestations of any category are noted) should
be carefully designed to suit the objectives of the observed discussion. Of course,
discussion is an integrative task, and so it is also advisable to give some
cognizance to comprehension performance in evaluating learners.
Games
Among informal
assessment devices are a variety of games that directly involve language
production.
ORAL
PROFICIENCY INTERVIEW (OPI)
The best-known
oral interview format is one that has gone through a considerable metamorphosis
over the last half-century, the Oral Proficiency Interview (OPI). Originally
known as the Foreign Service Institute (FSI) test, the OPI is the result of a
historical progression of revisions under the auspices of several agencies,
including the Educational Testing Service and the American Council on Teaching
Foreign Languages (ACTFL). The latter, a- professional society for research on foreign
language instruction and assessment, has now become the prinCipal body for
promoting the use of the OPI."The OP! is widely used across dozens of
languages around the world. Only certified examiners are authorized to
administer the OP!; certification workshops are available, at costs of around $700
for ACTFL members, through ACTFL at selected sites and conferences throughout
the year.
DESIGNING
ASSESSMENTS: EXTENSIVE SPEAKING
Extensive
speaking tasks involve complex, relatively lengthy stretches of discourse. They are frequently
variations on monologues, usually with minimal verbal interaction. Oral
Presentations In !he academic and professional arenas, it would not be uncommon
to be called on to present areport-,a-paper;a marketing plan,a-sales-idea, a
design of a new product, or a method. A summary of oral assessment techniques
would therefore be incomplete without some consideration of extensive speaking
tasks. Once again the rules ! I
for
effective assessment must be invoked: (a) specify the criterion, (b) set appro priate tasks, (c)
elicit optimal output, and (d) establish practical, reliable scoring procedures.
And once again scoring is the key assessment challenge.
Picture-Cued
Story-Telling
One of the most
common techniques for eliciting oral production is through visual pictures,
photographs, diagrams, and charts. We have already looked at this' elicitation device
for intensive tasks, but at this level we consider a picture or a series of
pictures as a stimulus for a longer story or deSCription .
It's always
tempting to throw any picture sequence at test-takers and have them talk for a
minute or so about them. But as is true of every assessment of speaking ability,
the objective of eliciting narrative discourse needs to be clear. In the above example
(with a little humor added!), are you testing for oral vocabulary (girl, alarm,
coffee, telephone, wet, cat, etc.), for time relatives (before, after, when),
for sentence connectors (then, and then, so), for past tense of irregular verbs
(woke, drank, rang), and/or for fluency in general? Ifyou are eliCiting
specific grammatical or discourse features, you might add to the directions something
like "Tell the story that these pictures describe. Use the past tense of
verbs." Your criteria for scoring need to be clear about what it is you
are hoping to assess. Refer back to some of the guidelines suggested under the
section on oral interviews, above, or to the OPI for some
general suggestions on scoring such a narrative.
Retelling
a Story, News Event
In this type of
task, test-takers hear or read a story or news event that they are asked to
retell. This differs from the paraphrasing task discussed above (pages 161-162)
in that it is a longer stretch of discourse and a different genre. The
objectives in assigning such.a task vary from listening comprehension of the
original to production of a number of oral discourse features (communicating
sequences and relationships 01 events, stress and emphasis patterns, .
"expression" in the case of a dramatic story), fluency, and
interaction with the hearer. Scoring should of course meet the intended
criteria.
Translation
(of Extended Prose)
Translation of
words, phrases, or short sentences was mentioned under the category
of-intensive speaking. Here, longer texts are presented for the test-taker to
read in the native language and then translate into English.Those texts could come
in many forms: dialogue, directions for assembly of a product, a synopsis of a
story or play or movie, directions on how to find something on a map, and other
genres. The advantage of translation is in the control of the content, vocabulary,
and, to some extent, the grammatical and discourse features. The disadvantage
is that translation of longer texts is a highly specialized skill for which
some individuals obtain post-baccalaureate degrees! To judge a nonspecialist's
oral language ability on such a skill may be completely invalid, especially if
the test-taker has not engaged in translation at this level. Criteria for scoring
should therefore take into account not only the purpose in stimulating a
translation but the possibility of errors that are unrelated to oral production
ability.
EXERCISES
[Note,: (I) Individual work; (G) Group or pair
work; (C) Whole-class discussion.
- (G) In the introduction to the chapter, the unique challenges of testing speaking were described (interaction effect, elicitation techniques, and scoring). In pairs, offer practical examples of one of the challenges, as aSSigned to your pair. Explain your examples to the class.
- (C) Review the five basic types of speaking that were outlined at the beginning of the chapter. Offer examples of each and pay special attention to distinguishing between imitative and intensive, and between responsive and interactive.(G) Look at the list of micro- and macroskills of speaking on pages 142-143. In pairs, each assigned to a different skill (or two), brainstorm some tasks that assess those skills. Present your findings to the rest of the class.
- (C) In Chapter 6, eight characteristics oflistening (page 122) that make listening "difficult"were listed. What makes speaking difficult? Devise a similar ijst that could fonn a set ofspecifications to pay special attention to in assessing speaking.
- (G) Divide the five basic types of speaking among groups or pairs, one type for each. Look at the sample assessment techniques provided and evaluate them according to the five pririciples (practicality, reliability, validity [especially face and content], authenticity, and washback). Present your critique to the rest of the class.
- (G) In· the same groups as in question #5 above, with the same type of speaking, design some other item types, different frorp the one(s) provided here, that assess the same type of speaking performance.
- (I) Visit the website listed for the PhonePass test. If you can afford it and you are a non-native speaker of English, take the test. Report back to the class on .'how valid, reliable, and authentic you felt the test was.
- (G) Several scoling scales are offered'in-this-chapter, ranging from simple (2-1-0) score categories to the more elaborate rubric used for the OP!. In groups, each assigned to a scoring scale, evaluate the strengths and weaknesses of each. Pay special attention to intra-rater and inter-rater reliability.
- (C) If pOSSible, role-playa formal oral interview in your class, with one student (with beginning to intermediate profiCiency in a language) acting as the testtaker and another (with advanced proficiency) as the test administrator. Use the sample questions provided on pages 169-170 as a guide. This role play will require some preparation. The rest of the class will then evaluate the effectiveness of the oral interview. Finally, the test-taker and administrator can offer their perspectives on the experience.
FOR
YOUR FURTHER READING
Underhill, Nic. (1987). Testing spoken
language:A handbook of oral testing techniques. Cambridge: Cambridge
University Press.
This practical
tnanual .on assessing spoken language is still a widely used collection of
techniques despite the fact that it was published in 1987. The chapter~re
organized into types of tests, elicitation techniques, scoring systems, and a
discussion of several types of
validity
along with reliability.
Brown,].D. (1998). New ways of classroom
assessment. Alexandria,VA: Teachers of English
to Speakers of Other Languages.
One of the many
volumes in TESOL's "New Ways" series, this one presents a collection
of assessment techniques across a wide range of skill areas. The two sections
on assessing oral skills offer 17 different techniq.
REFERENCES :
Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california.
Tidak ada komentar:
Posting Komentar