Assessing grammar,vocabulary
Chapter one
Differing notions of ‘grammar’ for assessment
Introduction
Even with the
sudden increase of research since the mid-1980s on the teaching and learning of
grammar, there still remains a surprising lack of consensus on (1) what
constitutes grammatical knowledge, (2) what type of assessment tasks might best
allow teachers and testers to infer that grammatical knowledge has been
acquired and (3) how to design tasks that elicit grammatical knowledge from
students for some specific assessment purpose, while at the same time providing
reliable and valid measures of performance.
What is meant by ‘grammar’ in theories of language?
Grammar and linguistics
Such linguistic
grammars are typically derived from data taken from native speakers and
minimally constructed to describe well-formed utterances within an individual
framework. These grammars strive for internal consistency and are mainly
accessible to those who have been trained in that particular
paradigm. In the syntactocentric view of language, formal
grammar is defined as a systematic way of accounting for and predicting an
‘ideal’ speaker’s or hearer’s knowledge of the language. This is done by a set
of rules or ‘principles’ that can be
used to generate all well-formed or grammatical utterances
in the language. This approach typically examines sounds that are combined to
form words, words that are put together to form phrases, phrases combined to
form clauses, and clauses assembled to form sentences.
In other words, this approach is predominantly concerned with the structure of
clauses and sentences, leaving the literal meaning and contextual use of these
forms to other approaches (i.e., to the fields of semantics and pragmatics). To
illustrate, consider the following sentence:
(1.1) Reggio and Messina were taken to the vet’s
this morning.
(1.1) Reggio and Messina were taken to the vet’s
this morning [by someone].
(1.2) [someone] took Reggio and Messina to the vet’s
this morning.
Form- based perspectives of language
traditional grammar
drew on data from literary texts to provide rich and lengthy descriptions of
linguistic form. Unlike some other syntactocentric
theories, traditional grammar also revealed the linguistic meanings of these
forms and provided information on their usage in a sentence (Celce-Murcia and
Larsen-Freeman, 1999). Traditional grammar supplied
an extensive set of prescriptive rules along with the exceptions. A typical
rule in a traditional English grammar might be:
The first-person singular of the present tense verb
‘to be’ is ‘I am’. ‘Am’ is used with ‘I’ in all cases, except in first-person
singular negative tag and yes/no questions, which are contracted. In this case,
the verb ‘are’ is used instead of ‘am’. For example, ‘I’m in a real bind,
aren’t I?’ or ‘Aren’t I trying my best?’
Figure 1.1
shows how a structural grammar might analyze statements and yes/no questions in
English.
Statements
Subject+Verb +Direct object+prepositional phrase
Steve
reads novels during the summer.
Yes/No questions
Auxiliary +Subject +Verb +Direct object
+Prepositional phrase
Does
Steve read novels
during the summer?
Figure 1.1 Structural analysis of statements and
yes/no questions in English
structural grammars are not based on a set of
prescriptive rules. Rather, they seek to describe the language as it appears
with a strict focus on grammatical form. Although descriptive linguistics has
provided numerous insights into the structure of languages,
it downplayed the semantic aspects of grammar, and provided little information
on how linguistic forms are used in context. Nonetheless,
many L2 educators continue to consider this theory a valuable resource for use
in syllabus design, grammar teaching and assessment.
Unlike the traditional or structural grammars that
aim to describe one particular language, transformational generative
grammar endeavored to provide a ‘universal’ description of language behavior
revealing the internal linguistic system for which all humans are predisposed
(Radford, 1988). Transformational-generative grammar claims that the underlying
properties of any individual language system can be uncovered by means of a
detailed, sentence-level analysis.
Nonetheless, both semantics and pragmatics, together
with phonology, morphology and syntax, are critical for assessing the
communicative success of an utterance within a given context.
To illustrate these shortcomings, consider the
following two syntactically identical
sentences.
(1.3) It is raining.
(1.4) It is working
In short, as a model for communicative teaching and
testing, the syntactocentric perspective has much to contribute; however, used
alone, it may not be appropriate for all situations, and must, therefore, be
adopted judiciously. Another example of the
theoretical limitations of applying a purely syntactocentric
approach to L2 educational contexts is seen in the following two pairs of
utterances.
Context: A French person, who speaks only French, is
having a discussion with two Americans, who both speak English and French
fluently. During the discussion, one American (Joe) lapses into English. The
other American (Sue) says:
(1.5) Sue: Would you please speak French? [request
and perhaps criticism]
(1.6) Joe: Oh, no problem. [acknowledgment and
agreement to comply]
Later, noticing that Joe has not stopped speaking
English, Sue repeats:
(1.7) Sue: Would you please speak French? [request
and criticism/chastisement]
(1.8) Joe: Sorry, I forgot. [apology and excuse]
To highlight further a need to account for meaning
on a lexicogrammatical level, consider the different interpretations of the
modal auxiliary ‘can’ in the following sentences:
Can you speak Kurdish? (ability or potential)
Can I have some milk, please? (request)
Can I go to the movies tonight, please? (request for
permission)
Can I buy you a beer? (offer)
Can we talk at 10? (suggestion)
Can they still be at work? (speculation)
Can it get any warmer? (theoretical possibility)
Other linguistic theories, however, are better
equipped to examine how speakers and writers actually exploit linguistic forms
during language use. For example, if we wish to explain how seemingly similar
structures like I like to read and I like reading connote different meanings, we
might turn to those theories that study grammatical form and use interfaces.
Katz
and Fodor (1963) looked at the connections between lexical forms and
grammatical forms by examining the features of words that encode grammar. They
found that in addition to encoding semantic features and restrictions, a word
also contains a number of syntactic features including the part of speech
(noun, verb, adjective), countability (singular, plural), gender (masculine,
feminine), and it can mark prepositional cooccurrence restrictions such as when
the word think is followed by a preposition (about, of, over) or is followed by
a that-clause.
One
type of information relates to the frequency and distribution of grammatical
forms. For example, Grabowski and Mindt’s (1995) study of 4,240 regular verb
types found in the Brown Corpus (Francis and Kucˇera, 1964) and
Lancaster-Oslo-Bergen Corpus (Johansson et al., 1978) discovered that regular
verbs accounted for only 42.3% of the total English verb tokens, with irregular
verbs making up the rest. Moreover, of these irregular verbs, 60% were
accounted for by be, have, or do, and 23.6% by say, make, go, take, come, see,
know, get, give, find,think,tell,become,show,leave,feel,put. In sum, these 20
verbs constituted an amazing 83.6% of the irregular verbs in the corpora.
corpus
linguistics has provided information on the different semantic functions of
lexical items. For example, a corpus linguist could examine the distribution and
frequency of occurrence of the word black and discover that it relates to
color, race, profit, cleanliness, amount of light and so forth
Communication-based
perspectives of language
A
pedagogical grammarrepresents an eclectic, but principled description of the
target-language forms, created for the express purpose of helping teachers
understand the linguistic resources of communication. These grammars provide
information about how language is organized and offer relatively accessible ways
of describing complex, linguistic phenomena for pedagogical purposes. The more
L2 teachers understand how the grammatical system works, the better they will
be able to tailor this information to their specific instructional contexts.
However,
in the tradition of pedagogical grammars, they also invoked other linguistic
theories and methods of analysis to explain the workings of grammatical form,
meaning and use when a specific grammar point was not amenable to a
transformational-generative analysis. For example, to explain the form and
meanings of prepositions, they drew upon case grammar (Fillmore, 1968) and to
describe the English tense-aspect system at the semantic level, they referred
to Bull’s (1960) framework relating tense to time. Celce-Murcia and
Larsen-Freeman’s (1999) book and other useful pedagogical English grammars
(e.g., Swan, 1995; Azar, 1998) provide teachers and testers alike with
pedagogically oriented grammars that are an invaluable resource for organizing
grammar content for instruction and assessment.
Chapter
two
Research
on L2 grammar teaching, learning and assessment
Introduction
This
has considerably broadened our notion of grammar and has led to a deeper
understanding of the role that grammar plays in conveying meaning in
communication.
Research
on L2 teaching and learning
To determine if students had actually learned under
the different conditions, teachers have used diverse forms of assessment and
drawn their own conclusions about their students. In so doing, these teachers
have acquired a considerable amount of anecdotal evidence on the strengths and
weaknesses of using different practices to implement L2 grammar instruction.
These experiences have led most teachers nowadays to ascribe to an eclectic
approach to grammar instruction, whereby they draw upon a variety of different
instructional techniques, depending on the individual needs, goals and learning
styles of their students.
In
recent years, some of these same questions have been addressed by second
language acquisition (SLA) researchers in a variety of empirically based
studies. These studies have principally focused on a description of how a
learner’s interlanguage (Selinker, 1972), or how a learner’s L2, develops over
time and on the effects that L2 instruction may have on this progression. In
most of these studies, researchers have investigated the effects of learning
grammatical forms by means of one or more assessment tasks. Based on the
conclusions drawn from these assessments, SLA researchers have gained a much
better understanding of how grammar instruction impacts both language learning
in general and grammar learning in particular. However, in far too many SLA
studies, the ability under investigation has been poorly defined or defined with
no relation to a model of L2 grammatical ability. Also, the empirical evidence
to support the learning claims have sometimes lacked credibility or
generalizability, and the scoring of the tasks or the reliability of the
measuring instruments have often not been reported.
The
SLA research looking at the role of grammar instruction in SLA might be
categorized into three strands. One set of studies has looked at the
relationship between the acquisition of L2 grammatical knowledge and different
language-teaching methods. These are referred to as the comparative methods
studies. A second set of studies has examined the acquisition of L2 grammatical
knowledge through what Long and Robinson (1998) call a ‘non-interventionist’
approach to instruction. These studies have examined the degree to which
grammatical ability could be acquired incidentally(while doing something else)
or implicitly (without awareness), and not through explicit(with awareness)
grammar instruction. A third set of studies has investigated the relationship
between explicit grammar instruction and the acquisition of L2 grammatical
ability. These are referred to as the interventionist studies, and are a topic
of particular interest to language teachers and testers.
Comparative
methods studies
The
comparative methods studies sought to compare the effects of different
language-teaching methods on the acquisition of an L2. More generally, these
studies were in reaction to form-focused instruction (referred to as ‘focus on
forms’ by Long, 1991), which used a traditional structural syllabus of
grammatical forms as the organizing principle for L2 instruction. According to
Ellis (1997), form-focused instruction contrasts with meaning-focused
instruction in that meaning-focused instruction emphasizes the communication of
messages (i.e., the act of making a suggestion and the content of such a
suggestion) while formfocused instruction stresses the learning of linguistic
forms. These can be further contrasted with form-and-meaning focused
instruction (referred to by Long (1991) as ‘focus-on-form’), where grammar
instruction occurs in a meaning-based environment and where learners strive to
communicate meaning while paying attention to form.
At
the end of the first and second years, students were tested in reading, writing,
listening and speaking. It must be noted, however, that these skill-based tests
were essentially form-focused grammar tests designed to measure knowledge of
linguistic forms while performing one of the language skills. For example, the
following speaking task provided students with a spoken present tense sentence,
and students were asked to say the same sentence in the past.
A:
Er spielt mit seinem Freund. (He plays with his friend.)
B:
Er spielte mit seinem Freund. (He played with his friend.)
Non-interventionist
studies
Many
interlanguage studies also showed that learners acquiring any
individualgrammaticalfeaturesuchasnegatives,interrogatives,relative clauses,
word order, or pronouns appeared to pass through a relatively fixed
developmental sequence toward mastering that form (Ellis, 1994). For example,
ESL learners learning the interrogatives would first use word(s) plus rising
intonation (You going?).
Empirical
studies in support of non-intervention
This
study sought to demonstrate that the development of grammatical ability could
be achieved through a task-based, rather than a form-focused, approach to language
teaching, provided that the tasks required learners to engage in meaningful
communication.
Possible
implications of fixed developmental order to language assessment
The
notion that structures appear to be acquired in a fixed developmental order and
in a fixed developmental sequence might conceivably have some relevance to the assessment
of grammatically ability. .In other words, one task could potentially tap into development
a level one, while another taps into developmental level two, and so forth.
In
a study on the effects of form-focused instruction and corrective feedback on
the acquisition of questions, Spada and Lightbown (1993) did just that.
Problems
with the use of development sequences as a basis for assessment
First,
the number of grammatical sequences that show a fixed order of acquisition is
very limited, far too limited for all but the most restricted types of grammar
tests. For example, what is the order for acquiring the modals, the
conditionals, or the infinitive or gerund complements? Second, much of the
research on acquisitional sequences is based on data from naturalistic
settings, where students are provided with considerable exposure to the
language.
Interventionist
studies
In
fact, several (e.g., Schmidt, 1983; Swain, 1991) have maintained that although
some L2 learners are successful in acquiring selected linguistic features
without explicit grammar instruction, the majority fail to do so. Testimony to
this is the large number of non-native speakers who emigrate to countries
around the world, live there all their lives and fail to learn the target
language, or fail to learn it well enough to realize their personal, social and
long-term career goals. In these situations, language teachers affirm that formal
grammar instruction of some sort can be of benefit. Furthermore, most language
teachers would contend that explicit grammar instruction, including systematic
error correction and other instructional techniques, contributes immensely to
their students’ linguistic development.
Empirical
studies in support of intervention
Similar
results were found by Doughty (1991), who compared the effectiveness of
naturalistic exposure to the target language with different types of instruction
in the acquisition of relative clauses. Using intermediate-level ESL students,
she asked one group, the control group, to read passages on the computer that
contained relative clauses. A second group, the meaning-oriented group, was
asked to read the same passages, except these students were also provided with
highlighted or capitalized lexical and semantic rephrasings of the relative
clauses, so the forms would potentially become salient and ‘noticed’. A third
group, the rule-oriented group, read the same passages, except they were also
given explicit rule statements below each relative clause so that the rules
would become salient. Knowledge of the relatives was measured by written
grammaticality-judgment, sentence-combination and gap-filling tasks, and by
sentence-level oral tasks based on pictures. Although no attempt was made to
measure literal, intended, or pragmatic meaning independent of the relative
clause forms, Doughty found that on the post-tests, the rule and
meaning-oriented groups outperformed the control group in their ability to use
relative clauses. However, the meaning-oriented group performed better than the
other two groups on the overall comprehension of the text. In short, this study
showed that naturalistic exposure alone was less effective than
form-and-meaning-based instruction in promoting the acquisition of relative
clause forms.
Research
on instructional techniques and their effects on acquisition
Feedback-based
techniques involve ways of providing negative evidence of grammar performance.
For example, ‘recast’ is a feedback based technique, where an utterance
containing an error is repeated without the error. Another is referred to as
‘garden path’ since learners are explicitly shown the linguistic rule and
allowed to generalize with other examples; however, when the generalization
does not hold (negative evidence), further instruction is provided. Finally,
metalinguistic feedback involves the use of linguistic terminology to promote
‘noticing’.
Grammar
processing and second language development
It
is important for language teachers and testers to understand these processes,
especially for classroom assessments. In the grammar-learning process, explicit
grammatical knowledge refers to a conscious knowledge of grammatical forms and
their meanings. Explicit knowledge is usually accessed slowly, even when it is
almost fully automatized (Ellis, 2001b).
Implicit
grammatical knowledge refers to ‘the knowledge of a language that is typically
manifest in some form of naturally occurring language behavior such as
conversation’ (Ellis, 2001b, p. 252). In terms of processing time, it is
unconscious and is accessed quickly.
In
fact, ‘focused instructional treatments of whatever sort far surpass non- or
minimally-focused exposure to the L2’ (Norris and Ortega, 2000, p. 463), and
this result holds in both the short and the long term (Doughty and Williams,
1998).
Implications
for assessing grammar
The
studies investigating the effects of teaching and learning on grammatical performance present a number of
challenges for language assessment.
The
information from these assessment should show how well students could apply the
forms in contexts where fluent and spontaneous language use is not required and
where time could be taken to figure out the answers. To obtain information on
the students’ implicit knowledge of grammatical forms, testers would need to
create tasks designed to elicit the fluent and spontaneous use of grammatical
forms in situations where automatic language use was required. In other words,
to infer that students could understand and produce grammar in spontaneous
speech, testers
wouldneedtopresentstudentswithtasksthatelicitcomprehensionorfull production in
real time (e.g., listening and speaking). Although this idea is interesting,
the introduction of speed into an assessment should be done with caution since
it is often difficult to determine the impact of speed on the test taker.
At
the same time, the research in SLA on the effectiveness of instructionaltreatmentshighlightsthecriticalrolethatgrammaticalassessment plays
in how language educators decide if learners are able to recognize and produce
the target-language structures. Assessment is used not only to determine the
state of a learner’s interlanguage, but also to ascertain the impact of
instructional treatments. It is thus surprising how little attention has been
devoted to ensuring that outcome instruments provide valid and reliable
measures of grammatical ability in SLA research. Because of this lack of rigor,
readers are often left questioning
theviabilityoftheresearch.Thus,SLAresearchersneedtoinformreaders about how the
tests used in their research were conceptualized, developed and scored.
Chapter
three
The
role of grammar in models of communicative language ability
Introduction
Although,
over the years, grammar instruction has changed considerably in communicative
language classrooms and research on how best to teach and learn it has proliferated,
this has had surprisingly little impact on how grammatical ability is assessed
in second and foreign language educational contexts. Far too many language
educators still use only multiple-choice tests of grammar and vocabulary in
assessing grammatical ability, or they use grammaticality judgments – if, in
fact, grammatical ability is assessed at all! Also, most language educators
remain wedded to a definition of grammatical knowledge that is limited to
sentence-level morphosyntactic form, even though in their classrooms, meaning
and grammar in discourse contexts are emphasized.
The
role of grammar in models of communicative competence
To
illustrate, imagine we wanted to determine a student’s grammatical knowledge of
the simple present, the simple past and the present perfect tenses as used in
conversational narratives.
In
sum, many different models of communicative competence have emerged over the
years. The more recent depictions have presented much broader
conceptualizations of communicative language ability; however, definitions of
grammatical knowledge have remained more or less the same – morphosyntax. Also,
within these expanded models, more detailed specifications are needed for how
grammatical form might interact with grammatical meaning to communicate literal
and intended meanings, and how form and meaning relate to the ability to convey
pragmatic meanings.
Rea-Dickins’definition
of grammar
Rea-Dickins
(1991) further stated that the goal of communicative grammar tests is to
provide an ‘opportunity for the test-taker to create his or her own message and
to produce grammatical responses as appropriate to a given context’ (p. 125).
This underscores the notion that pragmatic appropriateness or accept ability
can add a crucial dimension to communication, and must not be ignored If the message is not understood as intended,
the message can be repaired or mis understandings can persist. Nonetheless,
Rea-Dickins’ emphasis on grammar as pragmatics correctly reminds us of the
close relationship among grammar, semantics and pragmatics.
Larsen-Freeman’s
definition of grammar
LarsenFreeman’s
(1991, 1997) framework for the teaching of grammar in communicative language
teaching contexts. characterized grammatical knowledge along three dimensions:
linguistic form, semantic meaning and pragmatic use. Form is defined as both
morphology, or how words are formed, and syntactic patterns, or how words are
strung together. This dimension is primarily concerned with linguistic
accuracy. The meaning dimension describes the inherent or literal message conveyed
by a lexical item or a lexico-grammatical feature. This dimension is mainly
concerned with the meaningfulnessof an utterance. The use dimension refers to
the lexico-grammatical choices a learner makes to communicate appropriately
within a specific context. Pragmatic use describes whenand whyone linguistic
feature is used in a given context instead of another, especially when the two
choices convey a similar literal meaning. In this respect, pragmatic use is
said to embody presuppositions about situational context, linguistic context,
discourse context, and sociocultural context. This dimension is mainly
concerned with making the right choice of forms in order to convey an
appropriate message for the context. According to Larsen-Freeman (1991), these three
dimensions may be viewed as independent or interconnected. For example, a
linguistic form such as the articles in English displays a syntactic, semantic
and pragmatic dimension, even though, perhaps in the classroom, it might be
necessary to focus more on the pragmatic aspect, which can pose the greatest
challenge to learners.
What
is meant by ‘grammar’for assessment purposes?
In
one testing situation the assessment goal might be to obtain information on
students’ knowledge of linguistic forms in minimally contextualized sentences,
while in another, it might be to determine how well learners can use linguistic
forms to express a wide range of communicative meanings. Regardless of the
assessment purpose, if we wish to make inferences about grammatical ability on
the basis of a grammar test or some other form of assessment, it is important
to know what we mean by ‘grammar’ when attempting to specify components of
grammatical knowledge for measurement purposes. language knowledge consists of
grammatical knowledge and pragmatic knowledge.
Grammatical
knowledge embodies two highly related components: grammatical form and
grammatical meaning. Grammatical form includes a host of forms, for example, on
the phonological, lexical, morphosyntactic, cohesive, information management,
and interactional levels. Knowledge of grammatical form, therefore, refers to
the knowledge of one or more of these linguistic forms. Grammatical meaning is
sometimes used to refer to the literal meaning expressed by sounds, words, phrases
and sentences, where the meaning of an utterance is derived from its component
parts or the ways in which these parts are ordered in syntactic structure. Some
linguists have referred to this as semantic meaning, utterance meaning or the
compositionality of an utterance (Jaszczolt, 2002). Others (e.g., Grice, 1957;
Levinson, 1983) have referred to it as literal meaning, sentence meaning or
conventional meaning.
Although literal meaning allows us to identify
what is said by a speaker, Jaszczolt (2002) notes that some utterances may not
be sufficiently informative for the speaker’s meaning to be fully conveyed (p.
54). Grammatical meaning refers to instances of language use in which what is said
is what is meant literally and is closely related to what the speaker intends
to communicate. First, the notion of‘ conveying literal meaning’ is important
since in many cases, the primary assessment goal is to determine if learners
are able to use forms to get their basic point across accurately and
meaningfully.
CHAPTER
FOUR
Towards
a definition of grammatical ability
Introduction
discussed
the role of grammar in models of communicative competence and showed how a more
detailed depiction of grammar was needed in order to assess how learners use
grammatical forms as a resource for conveying a variety of meanings.
What
is meant by grammatical ability?
Defining
grammatical constructs
In
other words, the type, range and scope of grammatical features required to
communicate accurately and meaningfully will vary from one situation to another.
For example, the type of grammatical knowledge needed to write a formal
academic essay would be very different from that needed to make a train
reservation. Given the many possible ways of interpreting what it means to
‘know’ grammar, it is important that we define what we mean by ‘grammatical
knowledge’ for any given testing situation.
The
many possible ways of interpreting what it means to ‘know grammar’ or to have
‘grammatical ability’ highlight the importance in language assessment of
defining key terms. Some of the same terms used by different testers reflect a
wide range of theoretical positions in the field of applied linguistics. These
include knowledge, competence, ability, proficiency and performance, to name a
few. These concepts are abstract, not directly observable in tests and open to
multiple definitions and interpretations.
Definition
of key terms
Grammatical knowledge
Language
knowledge is then a mental representation of informational structures related
to language. The exact components of language knowledge, like any other
construct, need to be defined. In this book, grammar refers to a system of
language whereas grammatical knowledge is defined as a set of internalized
informational structures related to the theoretical model of grammar proposed
in Figure 3.2 (p. 62). In this model, grammar is defined in terms of grammatical
form and meaning, which are available to be accessed in language use.
Grammatical ability
Grammatical
ability is, then, the combination of grammatical knowledge and strategic
competence; it is specifically defined as the capacity to realize grammatical
knowledge accurately and meaningfully in testing or other language-use
situations.
Grammatical performance
grammatical
performance is defined as the observable manifestation of grammatical ability in
language use. In grammatical performance, the underlying grammatical ability of
a test-taker may be masked by interactions with other attributes of the
examinee or the test task.
Metalinguistic knowledge
Finally,
metalanguage is the language used to describe a language. It generally consists
of technical linguistic or grammatical terms (e.g., noun, verb). Metalinguistic
knowledge, therefore, refers to informational structures related to linguistic
terminology.
What
is ‘grammatical ability’for assessment purposes?
Assessment
of grammatical ability in this book is based on several specific definitions.
First, grammar encompasses grammatical form and meaning, whereas pragmatics is
a separate, but related, component of language. . Finally, in cases where
grammatical ability is assessed by means of an interactive test task involving
two or more interlocutors, the way grammatical ability is realized will be
significantly impacted by both the contextual and the interpretative demands of
the interaction.
Knowledge
of phonological or graphological form and meaning
Knowledge
of graphological form enables us to understand and produce features of the
writing system as they are used to convey meaning in testing or language-use
situations. Graphological form includes sound–spelling correspondences
(bear/bare), and other orthographical conventions.
Knowledge
of lexical form and meaning
Knowledge
of lexical form enables us to understand and produce those features of words
that encode grammar rather than those that reveal meaning. This includes words
that mark gender (e.g., waitress), countability (e.g., people) or part of
speech (e.g., relate, relation). For example, when the word think in English is
followed by the preposition about before a noun, this is considered the
grammatical dimension of lexis, representing a co-occurrence restriction with
prepositions.
Knowledge
of lexical meaning allows us to interpret and use words based on their literal
meanings. Lexical meaning here does not encompass the suggested or implied
meanings of words based on contextual, sociocultural, psychological or
rhetorical associations. For example, the literal meaning of a rose is a kind
of flower, whereas a rose can also be used in a non-literal sense to imply a
number of sociocultural meanings depending on the context.
Finally,
the choice of some lexical forms is based purely on usage norms, preferences or
expectations, and not solely on grammatical grounds.
Knowledge
of morphosyntactic form and meaning
This
includes the articles, prepositions, pronouns, affixes (e.g., -est), syntactic
structures, word order, simple, compound and complex sentences, mood, voice and
modality. A learner who knows the morphosyntactic form of the English
conditionals would know that: (1) an if-clause sets up a condition and a result
clause expresses the outcome; (2) both clauses can be in the sentence-initial
position in English; (3) if can be deleted under certain conditions as long as
the subject and operator are inverted; and (4) certain tense restrictions are
imposed on if and result clauses
Morphosyntactic
forms carry morphosyntactic meaningswhich allow us to interpret and express
meanings from inflections such as aspect and time, meanings from derivations
such as negation and agency, and meanings from syntax such as those used to
express attitudes (e.g., subjunctive mood) or show focus, emphasis or contrast
(e.g., voice and word order). For example, a student who knows the
morphosyntactic meaning of the English conditionals would know how to express a
factual conditional relationship (If this happens, that happens), a predictive
conditional relationship (If this happens, that will happen), or a hypothetical
conditional relationship (If this happened, that would happen).
Knowledge
of cohesive form and meaning
Knowledge
of cohesive form enables us to use the phonological, lexical and
morphosyntactic features of the language in order to interpret and express
cohesion on both the sentence and the discourse levels. Cohesive form is
directly related to cohesive meaning through cohesive devices (e.g., she, this,
here)which create links between cohesive forms and their referential meanings
within the linguistic environment or the surrounding co-text. Cohesive
form on a phonological level (common literary terms called assonance and
alliteration) can be seen in an excerpt from a poem by Paul Verlaine (1866) in
Poèmes saturniens: Les
sanglots longs de l’automne (The long sobs of the autumn) blessent mon cœur
(wound my heart) d’une langueur monotone (with a monotonous languor)
Knowledge
of information management form and meaning
Knowledge
of information management for mallows us to use linguistic forms as are source for
interpreting and expressing the information structure of discourse. Some
resources that help manage the presentation of information include, for example,
prosody, wordorder, tense-aspectand parallel structures. These forms are used to
create information management meaning.
For
example, consider how word order can emphasize the new information variation in
the following sentences.
1.
Liz gave Steve the wine.
2.
Liz gave the wine to Steve.
Knowledge
of interactional form and meaning
Knowledge
of interactional form enables us to understand and use linguistic forms as a
resource for understanding and managing talk-ininteraction. These forms include
discourse markers and communication management strategies. Discourse markers
consist of a set of adverbs, conjunctions and lexicalized expressions used to
signal certain language functions. For example, well can signal
disagreement, ya know or ahhuh can signal shared knowledge, and by the way can
signal topic diversion. Finally,
from a pragmatic perspective, interactional forms and meanings embody a number
of implied meanings. Consider the following examples.
Example
1 A: Sorry. I didn’t have money to buy the flowers.
B: Hello . . .? Today’s her birthday. You
could’a told me.
Example
2 A: Wow, those kids’re really a handful!
B: Thank you.
Knowledge
of interactional form and meaning
Knowledge
of interactional form enables us to understand and use linguistic forms as a
resource for understanding and managing talk-ininteraction. These forms include
discourse markers and communication management strategies. Discourse markers
consist of a set of adverbs, conjunctions and lexicalized expressions used to
signal certain language functions. For example, well . . . can signal
disagreement, ya know or ahhuh can signal shared knowledge, and by the way can
signal topic diversion.
Similar
to cohesive forms and information management forms, interactional forms use
phonological, lexical and morphosyntactic resources to encode interactional
meaning. For example, in saying *What means that?, the learner knows how to
repair a conversation by asking for clarification, but does not know the form of
the request.
Finally,
from a pragmatic perspective, interactional forms and meanings embody a number
of implied meanings. Consider the following examples.
Example
1 A: Sorry. I didn’t have money to buy the flowers. B: Hello . . .? Today’s her
birthday. You could’a told me.
Example
2 A: Wow, those kids’re really a handful! B: Thank you.
CHAPTER
FIVE
Designing
test tasks to measure L2 grammatical ability
Introduction
However,
some of the most important factors that affect grammar-test scores, aside from
grammatical ability, are the characteristics of the test itself. In fact,
anyone who has ever taken a grammar test, or any test for that matter, knows
that the types of questions on the test can severely impact performance. For
example, some test-takers perform better on multiple-choice tasks than on oral
interview tasks; others do better on essays than on cloze tasks; and still
others score better if asked to write a letter than if asked to interpret a
graph. Each of these tasks has a set of unique characteristics, called
test-task characteristics. These characteristics can potentially interact with
the characteristics of the examinee (e.g., his or her grammatical knowledge,
personal attributes, topical knowledge, affective schemata) to influence test
performance. Given the potential impact of test-task characteristics on
performance, it is important for test developers to understand the individual
characteristics of the tasks they use and to follow systematic procedures for
designing and developing tasks that will elicit the best possible
manifestations of grammatical ability.
How
does test development begin?
A
TLU task is one of many languageuse tasks that test-takers might encounter in
the target language use domain. It is to this domain that language testers
would like to make inferences about language ability, or more specifically,
about grammatical ability.
In
this example, the TLU situation is language instruction at flight school, and
the assessment purpose is to measure the student’s ability to use grammar as a
resource for communication in this setting.
What
do we mean by ‘task’?
Traditionally,
‘task’ has referred to any activity that requires students to do something for
the intent purpose of learningthe target language. A task then is any activity
(i.e., short answers, role-plays) as long as it involves a linguistic or
nonlinguistic (circle the answer) response to input.
The
first involves task-naturalness,a condition where ‘a grammatical construction
may arise naturally during the performance of a particular task, but the task
can often be performed perfectly well, even quite easily, without it’ (p. 132).
For example, in a task designed to elicit past modals in the context of a
murder mystery, we expect forms like: the butler could have done it or the maid
might have killed her, but we might get forms like: Maybe the butler did itor I
suspect the maid killed her. The second condition is task-utility, where ‘it is
possible to complete the task [meaningfully] without the structure, but with
the structure the task becomes easier’ (ibid.). For example, in a comparison
task, I once had a student say: *Shiraz is beautiful city,but Esfahan is
very,very,very, beautiful city in Iran.Had he known the comparatives or the
superlatives, his message could have been communicated much more easily. The
final and most interesting condition for grammar assessment entails
taskessentialness. This is where the task cannot be completed unless the
grammatical form is used. For example, in a task intended to distinguish
stative from dynamic adjectives, the student would need to know the difference
between I’m really boredand I’m really boringin order to complete the task.
Obviously task essentialness is the most difficult, yet the most desirable
condition to meet in the construction of grammar tasks. In real-life domains,
language is used as a resource for transaction and negotiated interaction; in
language-instruction domains, language is used in the context of language
learning, whether that involves interaction or not. Let us now examine the
individual characteristics of tasks.
What
are the characteristics of grammatical test tasks?
As
all language teachers know, the kinds of tasks we use in tests and their
quality can greatly influence how students will perform. In other words,
specifically designed tasks will work to produce the types of variability in
test scores that can be attributed to the underlying constructs given the
contexts in which they were measured (Tarone, 1998). To understand the
characteristics of test tasks better, we turn to Bachman and Palmer’s (1996)
framework for analyzing target language use tasks and test tasks.
The
Bachman and Palmer framework
These
five aspects describe characteristics of (1) the setting, (2) the test rubrics,
(3) the input, (4) the expected response and (5) the relationship between the
input and response. This framework can be used to (1) describe the TLU tasks as
a basis for designing test tasks; (2) specify the test tasks; and (3) compare.
Describing
grammar test tasks
Traditionally,
there have been many attempts at categorizing the types of tasks found on
tests. Some have classified tasks according to scoring procedure. For example,
objective test tasks (e.g., true–false tasks) are those in which no expert
judgment is required to evaluate performance with regard to the criteria for
correctness. Subjective test tasks (e.g., essays) are those that require expert
judgment to interpret and evaluate performance with regard to the criteria for
correctness.
Selected-response
task types
Selected-response
tasks present input in the form of an item, and testtakers are expected to
select the response. Other than that, all other task characteristics can vary.
For example, the form of the input can be language, non-language or both, and
the length of the input can vary from a word to larger pieces of discourse.
Finally, selected-response tasks can vary in terms of reactivity, scope and
directness.
-
The
multiple-choice (MC) task
-
Multiple-choice
error identification task
-
The
matching task
-
The
discrimination task
-
The
noticing task
Limited-production
task types
Limited-production
tasks are intended to assess one or more areas of grammatical knowledge
depending on the construct definition. Unlike selected-response items, which
usually have only one possible answer, the range of possible answers for
limited-production tasks can, at times, be large – even when the response involves
a single word.
-
The
gap-filling task
-
The
short-answer task
-
The
dialogue (or discourse) completion task (DCT)
-
Extended-production
tasks
-
The
information-gap task (info-gap)
-
Story-telling
and reporting tasks
-
The
role-play and simulation tasks
CHAPTER
SIX
Developing
tests to measure L2 grammatical ability
Introduction
Building
on the procedures for designing grammar-test tasks, this chapter addresses the
process of grammar-test construction, that is the principles underlying the
design, development and scoring of grammatical assessments.
What
makes a grammar test ‘useful’?
Score-based
inferences from grammar tests can be used to make a variety of decisions. For
example, classroom teachers use these scores as a basis for making inferences
about learning or achievement. These inferences can then serve to provide
feedback for learning and instruction, assign grades, promote students to the
next level, or even award a certificate.
The
information derived from language tests, of which grammar tests are a subset,
can be used to provide test-takers and other test-users with formative and
summative evaluations.
Score-based
inferences from grammar tests can also be used to make, or contribute to,
decisions about program placement. Many language testers (e.g., Harris, 1969;
Lado, 1961) have addressed this question over the years. Most recently, Bachman
and Palmer (1996) have proposed a framework of test usefulness by which all
tests and test tasks can be judged, and which can inform test design,
development and analysis. They consider a test ‘useful’ for any particular
testing situation to the extent that it possesses a balance of the following
six complementary qualities: reliability, construct validity, authenticity,
interactiveness, impact and practicality. They further maintain that for a test
to be ‘useful’, it needs to be developed with a specific purpose in mind, for a
specific audience, and with reference to a specific target language use (TLU)
domain. Given the importance of these qualities for grammar assessment, I will
describe them in some detail.
The
quality of reliability
This
consistency of measurement is referred to as test reliability, and it ranges on
a scale from zero (no consistency) to one (perfect consistency).
Another
way is to adopt objective scoring procedures. Objective scoring techniques
involve no expert decision-making in the scoring process such as in the scoring
of selected-response items. In cases where right/wrong scoring is not
appropriate, the scoring process can be ‘objectified’ by training raters to
score consistently according to an agreed-upon scoring rubric, and by having
more than one independent rater judging performance. Finally, reliability can
be raised by increasing the number of tasks on a test, the number of
test-takers or the number of judges.
The
quality of construct validity
Construct
validity also has to do with the domain of generalization to which our score
interpretations generalize’ (p. 21). In other words, construct validity not
only refers to the meaningfulness and appropriateness of the interpretations we
make based on test scores, but it also pertains to the degree to which the
score-based interpretations can be extrapolated beyond the testing situation to
a particular TLU domain (Messick 1993). Construct validity of score-based
interpretations needs to be supported through the collection and analysis of
data grounded in research and theory. In sum, construct validity is clearly one
of the most important qualities a test can possess.
The
quality of authenticity
A
third quality of test usefulness is authenticity, a notion much discussed in
language testing since the late 1970s, when communicative approaches to
language teaching were first taking root. Building on these discussions, Bachman
and Palmer (1996) refer to ‘authenticity’ as the degree of correspondence
between the test-task characteristics and the TLU task characteristics.
Finally,
authenticity, in my view, is also enhanced when the linguistic characteristics
of the test input appear ‘natural’. In other words, to the greatest extent
possible, the written or spoken input should resemble naturalistic discourse,
conforming to the norms, preferences and expectations of naturally occurring
talk or text. Similarly, tasks should be devised to elicit natural-sounding
responses.
In
sum, test authenticity resides in the relationship between the characteristics
of the TLU domain and characteristics of the test tasks, and although a test
task may be highly authentic, this does not necessarily mean it will engage the
test-taker’s grammatical ability.
The
quality of interactiveness
A
fourth quality of test usefulness outlined by Bachman and Palmer (1996) is
interactiveness. This quality refers to the degree to which the aspects of the
test-taker’s language ability we want to measure (e.g., grammatical knowledge,
language knowledge) are engaged by the testtask characteristics (e.g, the input
response, and relationship between the input and response) based on the test
constructs.
Consider,
for example, the chemistry lab report task whose input requires examinees to
invoke strategies to use their grammatical knowledge, the focus of measurement,
to express their ideas about the lab procedure (topical knowledge). This task
is likely to be more interactive than a task that is unsuccessful in engaging
aspects of the test-taker’s language ability to such a degree. The engagement
of these construct-relevant characteristics with task characteristics is the
essence of actual language use. Note again that for grammar assessment, what is
important is that the task succeeds in engaging the examinee’s grammatical
ability as intended by the test design. A task may be interactive because it
engages the examinee’s topical knowledge and positive affective schemata;
however, if the purpose of the test is to measure grammatical ability and the
task does not engage the ability of interest, this is all construct irrelevant.
If the construct is defined in such a way that it includes both grammatical
knowledge and topical knowledge (i.e., language for specific purposes), then the
task should be designed to engage these two constructs and little else.
The
quality of impact
Bachman
and Palmer (1996) refer to the degree to which testing and test score decisions
influence all aspects of society and the individuals within that society as test
impact. In terms of impact, most educators would agree that tests should
promote positive test-taker experiences leading to positive attitudes (e.g., a
feeling of accomplishment) and actions (e.g., studying hard).
A
special case of test impact is washback, which is the degree to which testing has
an influence on learning and instruction. Washback can be observed in grammar
assessment through the actions and attitudes that test-takers display as a
result of their perceptions of the test and its influence over them. For
example, examinees who are able to use corrective feedback from assessments to
clarify or extend their knowledge of grammar, or improve their ability to write
lab reports, would most likely perceive these tests as being ‘useful’.
The
quality of practicality
Test
practicality is not a quality of a test itself, but is a function of the extent
to which we are able to balance the costs associated with designing,
developing, administering, and scoring a test in light of the available
resources (Bachman, personal communication, 2002).
In
sum, the characteristics of test usefulness, proposed by Bachman and Palmer
(1996), are critical qualities to keep in mind in the development of a grammar
test.
Overview
of grammar-test construction
As
a result, there is no one ‘right’ way to develop a test; nor are there any
recipes for ‘good’ tests that could generalize to all situations. There are,
however, several frameworks of test development that have been proposed (e.g.,
Alderson, Clapham and Wall, 1995; Bachman and Palmer, 1996; Brown, 1996; Davidson
and Lynch, 2002) which serve to guide the test-development process so that the
qualities of test usefulness will not be ignored.
Test
development is often presented as a linear process consisting of a number of
stages and steps. In reality, the process is anything but linear. Instead, it
should be viewed as iterative and recursive, where knowledge and experience
gained at one stage of the process will require the reassessment of a previous
stage, followed by a series of readjustments.
Stage
1:Design
According
to Bachman and Palmer (1996, p. 88), this document should contain the following
components:
1.
a description of the purpose(s) of the test,
2.
a description of the TLU domains and task types,
3.
a description of the test-takers,
4.
a definition of the construct(s) to be measured,
5.
a plan for evaluating test usefulness, and
6.
a plan for dealing with resources.
Stage
2:Operationalization
The
outcome of the operationalization phase is both a blue print for the entire test
including scoring materials and a draft version of the actual test. According to
Bachmanand Palmer (1996), the blueprint contains two parts: a description of the
overall structure of the test and asset of test-task specifications for each task.
The blueprint serves as the basis for item writing and scoring.
-
Specifying
the scoring method
-
Scoring
selected-response tasks
-
Scoring
limited-production tasks
-
Scoring
extended-production tasks
-
Using
scoring rubrics
-
Grading
Stage
3:Test administration and analysis
The
actual administration of the test should transpire in a setting that is
physically comfortable and free from distraction, and a supportive testing
environment should be established. Instructions should be clear and the
administration orderly. Test administration provides an excellent opportunity
for collecting information about the test-takers’ initial reaction to the test
tasks and information about certain test procedures such as the allotment of
time.
Test
analyses provide different types of information to evaluate the characteristics
of test usefulness. This information serves as a basis for revising the test
before it goes operational, at which time further data are collected and
analyses performed in an iterative and recursive manner.
CHAPTER
SEVEN
Illustrative
tests of grammatical ability
Introduction
Some
of these tests contain separate sections that are exclusively devoted to the
assessment of grammatical ability, while others measure grammatical knowledge
along with other components of language ability in the context of language use
– that is while test-takers are listening, speaking, reading or writing. The
purpose of examining these tests is to illustrate how a few large-s calegrammar
test shave been designed and operationalized in light of their purpose(s), intended
use(s) and the construct(s) they are trying to measure.
The
First Certificate in English Language Test (FCE)
-
Purpose
-
Construct
definition and operationalization
-
Measuring
grammatical ability through language use
-
The
FCE and the qualities of test usefulness
-
Summary
The
Comprehensive English Language Test (CELT)
-
Purpose
-
Construct
definition and operationalization
-
Measuring
grammatical ability through language use
-
The
CELT and the qualities of test usefulness
-
Summary
The
Community English Program (CEP) Placement Test
-
Purpose
-
Construct
definition and operationalization
-
Measuring
grammatical ability through language use
-
The
CEP Placement Test and the qualities of test usefulness
-
Summary
CHAPTER
EIGHT
Learning-oriented
assessments of grammatical ability
Introduction
In
the context of learning grammar, learning-oriented assessment of grammar
reflects a growing belief among educational assessment experts (e.g., Stiggins,
1987; Gipps, 1994; Pellegrinio, Baxter and Glaser, 1999; Rea-Dickins and
Gardner, 2000) that if assessment, curriculum and instruction were more
integrally connected, student learning would improve (National Research
Council, 2001b). This approach attempts to provide teachers and learners with
summative and/or formative information on the test-takers’ grammatical ability.
Summative information from assessment allows teachers to assign grades based on
specific assessment criteria, report student progress at a single moment or over
time, and reward and motivate student learning. Formative information from
assessment provides teachers and learners with concrete information on what
aspects of the grammar students have and have not mastered and involves them in
the regulation and assessment of their own learning, so that further learning
can take place independently or in collaboration with teachers and other
students.
A
learning-oriented approach to grammar assessment addresses the following
questions.
•
How do I know if my students have learned and internalized the grammar points
covered in the course?
•
How do I know if my students can use these grammar points to communicate
spontaneously in real-life situations?
•
How do I know if the test tasks make it essential for my students to use the
target grammar points?
•
How can I use grammar assessment results to provide feedback for guiding
learning?
•
How will the results from this grammar test provide information to me on what
to (re)teach?
•
How can I design interesting and cognitively engaging grammar tasks so my
students will enjoy learning grammar?
What
is learning-oriented assessment of grammar?
The
terms alternative assessment, authentic assessment and performance assessment
have all been associated with calls for reform to both large-scale and
classroom assessment contexts. Alternative assessment emphasizes an alternative
to and rejection of selected-response, timed and one-shot approaches to
assessment, whether they occur in large-scale or classroom assessment contexts.
Alternative assessment encourages assessments in which students are asked to
perform, create, produce or do meaningful tasks that both tap into higher-level
thinking (e.g., problem-solving) and have real-world implications (Herman et
al., 1992). Alternative assessments are scored by humans, not machines.
Similar
to alternative assessment, authentic assessment stresses measurement practices
which engage students’ knowledge and skills in ways similar to those one can
observe while performing some real-life or ‘authentic’ task (O’Malley and
Valdez-Pierce, 1996). It also encourages tasks that require students to perform
some complex, extended production activity, and emphasizes the need for
assessment to be strictly aligned with classroom goals, curricula and
instruction. Self assessment is considered a key component of this approach.
Performance assessment refers to the evaluation of outcomes relevant to a domain
of interest (e.g., grammatical ability), which are derived from the observation
of students performing complex tasks that invoke real world applications
(Norris et al., 1998). As with most performance data, assessments are scored by
human judges (Stiggins, 1987; Herman et al., 1992;Brown,1998)accordingtoascoringrubricthatdescribeswhattesttakersneedtodoinordertodemonstrateknowledgeorabilityatagiven
performance level. Bachman (2002) characterized language performance assessment
as typically: (1) involving more complex constructs than those measured in
selected-response tasks; (2) utilizing more complex and authentic tasks; and
(3) fostering greater interactions between the characteristics of the
test-takers and the characteristics of the assessment tasks than in other types
of assessments. Performance assessment encourages self-assessment by making
explicit the performance criteria in a scoring rubric. In this way, students
can then use the
criteriatoevaluatetheirperformanceandcontributeproactivelytotheir own learning.
Finally,
learning-oriented assessment is designed to be an integral part of instruction,
occurring formally or informally at any stage of the learning process. Learning-oriented
assessment data can also be collected at one point in time or over a period of
time. Unlike large-scale assessments, learning-oriented assessment is
fundamentally iterative and recursive in that feedback from one assessment is
intended to provide information for subsequent learning and assessment, until a
criterion level of mastery has been achieved. Finally, these assessments are
scored by machines or humans, depending on the nature of the task and the
scoring procedures, as described in earlier chapters.
Implementing
learning-oriented assessment of grammar
Considerations
from grammar-testing theory
-
Implications
for test design
-
Implications
for operationalization
-
Planning
for further learning
Considerations
from L2 learning theory
-
SLA
processes – briefly revisited
-
Assessing
for intake
-
Assessing
to push restructuring
-
Assessing
for output processing
Illustrative
example of learning-oriented assessment
Background
The
goal of the achievement tests is ‘to measure the students’ knowledge of grammar,
vocabulary, pronunciation, reading and writing, as taught in each unit’
(Purpura et al., 2001, p. iii). The test results are intended to indicate
mastery of the learning points in the unit being tested and to determine if
students are ready for the next unitor level of the program.
Finally,
the writing section aimed to measure the test-takers’ ability to write a
recommendation paragraph using the present perfect tense. This task hoped to
elicit the learners’ implicit knowledge of grammatical form and meaning.
Test-takers have to read a situation and brainstorm information. They then have
to use this information to write a recommendation paragraph to the principal,
justifying their choice for the award. They are reminded to check their work
for organization and for the use of the present perfect tense. The
brainstorming task was intended to be scored with a three-point holistic rubric
defined in terms of topical control, task fulfillment and information
relevance/validity.
Making
assessment learning-oriented
From
a learning perspective, the achievement test was based on the premise that
students had had plenty of opportunities in class to demonstrate their
understanding of the present perfect tense and to receive feedback. Thus, it
was presumed that assessment was taking place at some point beyond intake. For
this reason, no comprehension tasks were included in the test. It was also
presumed that most students were well on their way toward incorporating the
target grammar into their interlanguage and that it would make sense to have
information on the degree to which students had learned the present perfect
tense and the degree to which this knowledge was implicit. For this reason,
both simple and complex tasks were used in the test.
In
sum, the On Target achievement test attempted to take into consideration
elements from both grammar-testing theory and L2 learning theory in achieving a
learning-oriented assessment mandate.
CHAPTER
NINE
Challenges
and new directions in assessing grammatical ability
Introduction
Research
and theory related to the teaching and learning of grammar have made significant
advances over the years. In applied linguistics, our understanding of language
has been vastly broadened with the work of corpus-based and communication-based
approaches to language study, and this research has made path ways into recent
pedagogical grammars. Also, our conceptualization of language proficiency has
shifted from an emphasis on linguistic form to one on communicative language
ability and communicative language use, which has, in turn, led to a demphasis
on grammatical accuracy and a greater concern for communicative effectiveness.
The
state of grammar assessment
In
a few cases, grammatical ability has been tested in both ways – as ‘a “body” of
knowledge and “a means to an end” with attention to . . . conveying appropriate
meanings in messages rather than an exclusive emphasis on accuracy of form and
structure’ (Rea-Dickins, 2001, p. 28). This has led to examinations in which
grammatical ability is measured by one or more separate-and-explicit,
selected-response or limitedproduction tasks of grammatical knowledge, as well
as one or more extended-production tasks designed to measure, amongst other
things, the test-takers’ implicit knowledge of grammar while speaking or
writing.
Challenge
1: Defining grammatical ability
While
the current research on learner-oriented corpora has shown great promise, many
more insights on learner errors and interlanguage development could be obtained
if other components of grammatical form (e.g., information management forms and
interactional forms) and if grammatical meaning were also tagged at both the
sentence and the discourse levels. For example, in a talk on the use of corpora
for defining learning problems of Korean ESL students at the University of
Illinois, Choi (2003) identified the following errors as passive errors:
1:
*The color of her face was changed from a pale white to a bright red.
2:
*It is ridiculous the women in developing countries are suffered.
While
it is true that the students may have overused the passive in these sentences,
it is clear that they have a full understanding of passive form, but not of
passive meaning, so that it can be used correctly. In sentence 1, the student
has failed to learn that ‘change’ requires the active voice since it is an
agentless ‘change-of-state’ or ergative verb, and in sentence 2, ‘suffer’
denotes a physical state and is intransitive, thereby making passivization
unlikely. As a result, these sentences might be tagged for meaning and not
form. This information could ultimately provide a more comprehensive
understanding of learner errors than a depiction based solely on form. It would
also root learner errors stemming from performance data to a broader model of
language proficiency.
Challenge
2: Scoring grammatical ability
Another
challenge relates to the scoring of grammatical ability in complex performance
tasks. In instances where the assessment goals call for the use of complex
performance tasks, we need to be sure to use welldeveloped scoring rubrics and
rating scales to guide raters to focus their judgments only on the constructs
relevant to the assessment goal. McNamara (1996) stresses that the scales in
such tasks represent, explicitly or implicitly, the theoretical basis upon
which the performance is judged. Therefore, clearly defined constructs of
grammatical ability and how they are operationalized in rating scales are
critical.
Challenge
3: Assessing meanings
The
‘communicative’ in communicative language teaching, communicative language
testing, communicative language ability, or communicative competence refers to
the conveyance of ideas, information, feelings, attitudes and other intangible
meanings (e.g., social status) through language. Therefore, while the
grammatical resources used to communicate these meanings precisely are
important, the notion of meaning conveyance in the communicative curriculum is
critical. Therefore, in order to test something as intangible as meaning in
second or foreign language use, we need to define what it is we are testing.
Challenge
4: Reconsidering grammar-test tasks
The
fourth challenge relates to the design of test tasks that are capable of both
measuring grammatical ability and providing authentic and engaging measures of
grammatical performance. Since the early 1960s, language educators have
associated grammar tests with discrete-point, multiple-choice tests of
grammatical form. These and other ‘traditional’ test tasks (e.g.,
grammaticality judgments) have been severely criticized for lacking in
authenticity, for not engaging test-takers in language use, and for promoting
behaviors that are not readily consistent with communicative language teaching.While
there is a place for discrete-point tasks in grammar assessment, language
educators have long used a wide range of simple and complex tasks in which to
assess test-takers’ explicit and implicit knowledge of grammar. In fact, in a
small-scale study designed to discover teacher practices in testing grammar in
primary, secondary and adult-school contexts, Rea-Dickins (2001) noted that 61
of the 70 teachers reported testing grammar explicitly, while 27 reported
assessing it indirectly through the language skills. Furthermore, 67 out of the
70 teachers reported testing grammar, and only one actually stated that it
should not be tested. In short, grammar testing in classrooms is alive and
well. One
discursive practice that naturally elicits the past and past continuous tenses
is the ‘eyewitness report’ between an eyewitness to some sudden event or close
call and a reporter. For example: Reporter:
‘What were you doing when the electricity went out?’ Interviewee:
‘I was having a heaping plate of pasta.’
Challenge
5: Assessing the development of grammatical ability
The
fifth challenge revolves around the argument, made by some researchers, that
grammatical assessments should be constructed, scored and interpreted with
developmental proficiency levels in mind. This notion stems from the work of
several SLA researchers (e.g. Clahsen, 1985; Pienemann and Johnson, 1987;
Ellis, 2001b) who maintain that the principal finding from years of SLA research
is that structures appear to be acquired in a fixed order and a fixed
developmental sequence. Furthermore, instruction on forms in non-contiguous
stages appears to be ineffective. As a result, the acquisitional development of
learners, they argue, should be a major consideration in the L2 grammar
testing. In terms of test construction, Clahsen (1985) claimed that grammar
tests should be based on samples of spontaneous L2 speech with a focus on
syntax and morphology, and that the structures to be measured should be
selected and graded in terms of order of acquisition in natural L2 development.
Furthermore, Ellis (2001b) argued that grammar scores should be calculated to
provide a measure of both grammatical accuracy and the underlying acquisitional
development of L2 learners. In the former, the target-like accuracy of a
grammatical form can be derived from a total correct score or percentage. In the
latter, the developmental proficiency can be derived from scores linked to
different stages of the interlanguage continuum. In this view, it was argued,
students and teachers can be provided with information that reflects both
target-like and developmental criteria with regard to knowledge of specific
grammatical forms. If these claims are accepted, the ensuing challenge to
language testers and SLA researchers is to adapt current test design and
scoring procedures to incorporate findings from this research.
Final
remarks
This
research has also highlighted the important role that meaning plays in learning
grammatical forms. In the same way, most language teachers and SLA researchers
around the world have never really given up grammar testing. Admittedly, some
have been perplexed as to how grammar assessment could be compatible with a
communicative language teaching agenda, and many have relied on assessment
methods that do not necessarily meet the current standards of test construction
and validation. With the exception of ReaDickins and a few others, language
testers have been of little help. In fact, a number of influential language
proficiency exams have abandoned the explicit measurement of grammatical
knowledge and/or have blurred the boundaries between communicative effectiveness
and communicative precision (i.e., accuracy).
CHAPTER
ONE
The
place of vocabulary in language assessment
Introduction
At ®rst glance, it may seem that
assessing the vocabulary knowledge
of
second language learners is both necessary and reasonably straightforward. It is
necessary in the sense that words are the basic building blocks of
language, the units of meaning from which larger structures such as
sentences, paragraphs and whole texts are formed.For native speakers, although
the most rapid growth occurs in childhood, vocabulary knowledge continues to
develop naturally in adult
life
in response to new experiences, inventions, concepts, social trends and
opportunities for learning. For learners, on the other hand, acquisition of
vocabulary is typically a more conscious and demanding process. Even at an
advanced level, learners are aware of limitations in their knowledge of second
language (or L2) words. They
experience
lexical gaps, that is words they read which they simply do not understand, or
concepts that they cannot express as adequately as they could in their
®rst language (or L1). Many learners see second language acquisition as
essentially a matter of learning vocabulary, so they devote a great
deal of time to memorising lists of L2 words and rely on their bilingual
dictionary as a basic communicative resource. Moreover, after a
lengthy period of being preoccupied with the development of grammatical
competence, language teachers and applied linguistic researchers now generally
recognise the importance of
vocabulary
learning and are exploring ways of promoting it moreeffectively. Thus, from
various points of view, vocabulary can be seen as a priority area in
language teaching, requiring tests to monitor thelearners' progress in
vocabulary learning and to assess how adequate their vocabulary
knowledge is to meet their communication needs.
Vocabulary
assessment seems straightforward in the sense that word lists are readily
available to provide a basis for selecting a set of words to be tested. In
addition, there is a range of well-known item types thatare convenient to use
for vocabulary testing. Here are some examples:
Multiple-choice (Choose the correct
answer)
The principal was irate when she
heard what the students had
done.
a. surprised
b. interested
c. proud
d. angry
Completion (Write in the missing word)
At last the climbers reached the s--------- of the mountain.
Translation (Give the L1 equivalent of
the underlined word)
They worked at the mill.
Matching (Match each word with its
meaning)
1 accurate a. not changing
2 transparent b. not friendly
3 constant c. related to seeing things
4 visual d. greater in size
5 hostile e. careful and exact
f. allowing light to go through
g. in the city
These test items
are easy to write and to score, and they makeef®cient use of testing time.
Multiple-choice items in particular have been commonly used in standardised
tests. A professionally produced,multiple-choice
vocabulary test is highly reliable and distinguishes learners effectively
according to their level of vocabulary knowledge.Furthermore, it will usually
be strongly related to measures of the learners' reading comprehension ability.
Handbooks on language testing
published in the 1960s and 1970s (for example Lado, 1961;Harris, 1969; Heaton,
1975) devote a considerable amount of space to vocabulary testing,
with a lot of advice on how to write good items and avoid various
pitfalls.
Tests containing
items such as those illustrated above continue to be written and used by
language teachers to assess students' progress in vocabulary learning
and to diagnose areas of weakness in their knowledge of
target-language words, i.e. the language which they are learning. Similarly,
scholars with a specialist interest in the learning and teaching of
vocabulary (see, for example, McKeown and Curtis,1987; Nation, 1990; Coady and
Huckin, 1997; Schmitt and McCarthy,1997) generally take it for granted that it
is meaningful to treat words
as
independent units and to devise tests that measure whether ± and how well ± learners
know the meanings of particular words.
Recent
trends in language testing
However, scholars in the ®eld of
language testing have a rather different perspective on vocabulary-test items of
the conventional kind.Such items ®t neatly into what language testers call the
discretepoint approach to testing. This involves designing tests to assess whether learners have
knowledge of particular structural elements of the language: word
meanings, word forms, sentence patterns, sound contrasts and so on. In
the last thirty years of the twentieth century, language testers
progressively moved away from this approach, to the extent that such
tests are now quite out of step with current thinking about how to
design language tests, especially for pro®ciency assessment.
A number of
criticisms can be made of discrete-point vocabulary tests.
-
It is dif®cult to make
any general statement about a learner's vocabulary on the basis of scores in
such a test. If someone gets 20
items
correct out of 30, what does that say about the adequacy of the learner's
vocabulary knowledge?
-
Being pro®cient in a second language is not
just a matter of knowing
a lot of words ± or grammar rules, for that matter ± but being able to exploit
that knowledge effectively for various communicative purposes. Learners can
build up an impressive knowledge
of
vocabulary (as re¯ected in high test scores) and yet be incapable of understanding a
radio news broadcast or asking for assistance atan enquiry counter.
-
Learners need to show
that they can use words appropriately in their own speech and writing, rather
than just demonstrating that
they
understand what a word can mean. To put it another way, the standard discrete-point
items test receptive but not productive competence.
-
In normal language use,
words do not occur by themselves or in isolated sentences but as integrated
elements of whole texts and
discourse.
They belong in speci®c conversations, jokes, stories, letters, textbooks,
legal proceedings, newspaper advertisements and so on. And the way that
we interpret a word is signi®cantly in¯uenced by the context in which it
occurs.
-
In communication
situations, it is quite possible to compensate for lack of knowledge of
particular words. We all know learners who are remarkably
adept at getting their message across by making the best use of limited
lexical resources. Readers do not have to understand every word in order to
extract meaning from a text satisfactorily. Some words can be ignored, while
the meaning of others can
be
guessed by using contextual clues, background knowledge of the subject matter and so
on. Listeners can use similar strategies, as well as seeking
clari®cation, asking for a repetition and checking that they have interpreted
the message correctly.
The widespread acceptance of the
validity of these criticisms has led
to
the adoption- particularly in the major English-speaking countries of the communicative
approach to language testing. Today's language pro®ciency tests do not set out
to determine whether learners
know
the meaning of magazine or put on or approximate; whether they can get the
sequence of tenses right in conditional sentences; or whether they can
distinguish ship and sheep. Instead, the tests are based on tasks
simulating communication activities that the learners are likely to be
engaged in outside of the classroom. Learners may be asked to write a letter
of complaint to a hotel manager, to show that they understand the
main ideas of a university lecture or to discuss in an interview how they
hope to achieve their career ambitions. Presumably good vocabulary knowledge
and skills will help test-takers to
perform
these tasks better than if they lack such competence, but neither vocabulary nor
any other structural component of the language is the primary focus of the
assessment. The test-takers are
judged
on how adequately they meet the overall language demands of the task.
Recent books on
language testing by leading scholars such asBachman and Palmer (1996) and
McNamara (1996) demonstrate how
the
task has become the basic element in contemporary test design. This is consistent with
broader trends in Western education systems away from formal
standardised tests made up of multiple items to measure students' knowledge
of a content area, towards what is
variously
known as alternative, performance-based or standardsbased assessment (see, for
example, Baker, O'Neil and Linn, 1993; Taylor, 1994; O'Malley and Valdez
Pierce, 1996), which includes
judging
students' ability to perform more open-ended, holistic and real-world' tasks
within their normal learning environment. Is there a place, then, for vocabulary
assessment within task-based
language
testing? To look for an answer to this question, we can turn to Bachman and Palmer's
(1996) book Language Testing in Practice,which is a comprehensive and
in¯uential volume on language-test
design
and development. Following Bachman's (1990) earlier work, the authors see the
purpose of language testing as being to allow us to make inferences about
learners' language ability, which consists of two components. One is
language knowledge and the other is
strategic
competence. That is to say, learners need to know a lot about the vocabulary,
grammar, sound system and spelling of the target language, but they also need to
be able to draw on that knowledge effectively for communicative purposes under
normal time constraints. As I noted above, one of the main criticisms of
discrete-point vocabulary
items is that they focus entirely on the knowledge component of language
ability.
Within the
Bachman and Palmer framework, language knowledge is classi®ed into numerous
areas, as presented in Table 1.1. The table shows that language
knowledge covers more areas than I indicated in the previous paragraph,
but at the same time knowledge of vocabulary appears to be just a
minor component of the overall system, a subsub-category of organisational
knowledge. It is classi®ed as part of Grammatical knowledge, which suggests a
very narrow view of vocabulary as a stock of meaningful word forms that ®t into
slots in sentence
frames. I will have a great deal more to say about the nature of vocabulary in
Chapter 2, but for now let me point out that vocabulary knowledge is a
signi®cant element in several other categories of The table. The most obvious area is
Sociolinguistic knowledge, which
includes
`natural or idiomatic expressions', `cultural references' and `®gures of speech'.
Most people would regard these as belonging to the vocabulary of the
language. In addition, the sociolinguistic
gambar
tests focus on just one of the areas of
language knowledge, such as
vocabulary.
They give as an example a test for primary school children learning English as a
foreign language in an Asian country. In the context of a teaching
unit on `Going to the zoo', the students are tested on their
knowledge of the names of zoo animals (Bachman and Palmer, 1996: 354±365).
The authors argue that, even at this elementary level of language learning,
vocabulary testing should relate to
some
meaningful use of language outside the classroom.
However, their
main concern is with the development of test tasks that not only draw on
various areas of language knowledge but also require learners to
show that they can activate
that knowledge effectively in communication. An illustration of the latter kind
of task
is found
in an academic writing test for non-native speakers of English entering a writing
programme in an English-medium university (Bachman and Palmer, 1996: 253±284). The
test-takers are required to
write
a proposal for improving the institution's admissions procedures. Rather than
the single global scale that is often employed to rate performance on
such a task, Bachman and Palmer advocate the use of several analytic
scales, which provide separate ratings for different components of
the language ability to be tested. In the case of the academic writing
test, they developed ®ve scales, for knowledge of syntax, vocabulary,
rhetorical organisation, cohesion and register. Thus, vocabulary is
certainly being assessed here, but not separately; it is part of a larger
procedure for measuring the students' academicwriting ability.
Three
dimensions of vocabulary assessment
Up to this point, I have outlined two
contrasting perspectives on the
role
of vocabulary in language assessment. One point of view is that it is perfectly sensible
to write tests that measure whether learners know the meaning and usage
of a set of words, taken as independent semantic units. The other view is that
vocabulary must always be
assessed
in the context of a language-use task, where it interacts in a natural way with other
components of language knowledge. To some extent, the two views are complementary
in that they relate to different purposes of assessment. Conventional
vocabulary tests are most
likely
to be used by classroom teachers for assessing progress invocabulary learning
and diagnosing areas of weakness. Other users of these tests are
researchers in second language acquisition with a special interest in how
learners develop their knowledge of, and ability to use, target-language words.
On the other hand, researchers
in
language testing and those who undertake large testing projects tend to be more
concerned with the design of tests that assess learners' achievement or
pro®ciency on a broader scale. For such purposes, vocabulary knowledge has a
lower pro®le, except to the extent
that
it contributes to, or detracts from, the performance of communicative tasks.
As with most
dichotomies, the distinction I have made between the two perspectives on
vocabulary assessment oversimpli®es the matter.There is a whole range of
reasons for assessing vocabulary knowledge and use, with a
corresponding variety of testing procedures. In order to map out the scope of
the subject, I propose three dimensions, as presented in Figure
1.1.
The
dimensions represent ways in which we can expand our conventional ideas about
what a vocabulary test is in order to include a wider range of lexical
assessment procedures. I introduce the dimensions here, then illustrate and
discuss them at various points in the
following chapters. Let us look at each
one in turn.
Discrete
- embedded
The ®rst
dimension focuses on the construct which underlies the assessment instrument.
In language testing, the term construct refers to the mental attribute
or ability that a test is designed to measure. In the case of a
traditional vocabulary test, the construct can usually be labelled as `vocabulary
knowledge' of some kind. The practical signi®-cance of de®ning the construct is
that it allows us to clarify the
meaning
of the test results. Normally we want to interpret the scores on a vocabulary test as
a measure of some aspect of the learners' vocabulary knowledge, such as their
progress in learning words from
the
last several units in the course book, their ability to supply derived forms of base words
(like scientist and scienti®c, from science), or their skill at inferring the
meaning of unknown words in a reading passage. Thus, a discrete test
takes vocabulary knowledge as a distinct construct, separated from other
components of language competence.
Whether
it is valid to do so is a matter for debate and an issue that
gambar
return to in
Chapter 4. However, most existing vocabulary tests are designed on the
assumption that it is meaningful to treat them as an independent construct
for assessment purposes and can thus be classi®ed as discrete measures in the
sense that I am de®ning it here.
In contrast, an
embedded vocabulary measure is one that contributes to the assessment of a
larger construct. I have already given an example of such a measure, when I
referred to Bachman and Palmer's
task
of writing a proposal for the improvement of university admissions procedures.
In this case, the construct can be labelled academic writing
ability', and the vocabulary scale is one of ®ve ratings which form a
composite measure of the construct. Another example of an embedded
measure is found in reading tasks consisting of a written text
followed by a set of comprehension questions. It is common practice to
include in such tests a number of items assessing the learners'
understanding of particular words or phrases in text.
Usually the vocabulary item scores are not separately counted; they simply form part
of the measure of the learners' `readingcomprehension ability'. In that sense,
vocabulary assessment is more
embedded
here than in the academic-writing test, where the vocabulary rating may well be
included in a pro®le report of each learner's writing ability.
It is important
to understand that the discrete±embedded distinction does not refer primarily
to the way that vocabulary is presented to the test-takers. Many discrete
vocabulary tests do require the learners to respond to words which are
presented in isolation or in a short
sentence,
but this is not what makes the test discrete. Rather, it is the fact that the test is
focusing purely on the construct of vocabulary knowledge. A test can
present words in quite a large amount of context and still be a discrete measure
in my sense. For instance, I
can
take a suitable reading passage, select a number of content words or phrases in it and
write a multiple-choice item for each one, designed to assess whether learners
can understand what the vocabulary
item
means as it is used in the text. This may appear to be very much the same kind of test
as the one I described in the last paragraph to illustrate what an
embedded measure is, but the crucial difference is that in this case all
the items are based on vocabulary
in
the passage and
I interpret the test score as measuring how well the learners can understand what those
words and phrases mean. I do not see it as assessing their reading
comprehension ability or any other broader construct. Thus, to
determine whether a particular vocabulary measure is discrete or embedded, you
need to consider its purpose
and
the way the results are to be interpreted.
Selective
-
comprehensive
The second
dimension concerns the range of vocabulary to be included in the assessment. A
conventional vocabulary test is based on a
set of target words selected by the test-writer, and the test-takers are
assessed according to how well they demonstrate
their knowledge of the meaning or use of
those words. This is what I call a selective vocabulary measure. The
target words may either be selected as individual words and then incorporated
into separate test items, or
alternatively
the test-writer ®rst chooses a suitable text and then uses certain words from it
as the basis for the vocabulary assessment.On the other hand, a comprehensive
measure takes account of all
the
vocabulary content of a spoken or written text. For example, let us take a speaking test
in which the learners are rated on various criteria, including
their range of expression. In this case, the raters are not listening for
particular words or expressions but in principle are forming a judgement
of the quality of the test-takers' overall vocabulary use. Similarly,
as we shall see in Chapter 7, some researchers have investigated productive
vocabulary use by setting
learners
a written composition task and then counting the number of different words or the
number of `sophisticated', low-frequency words used.
Comprehensive
measures can also be applied to the input material for reading or
listening tests. It is common practice for test-writers to use a readability
measure as one way of judging the suitability of a text for the assessment
of a particular group of test-takers. Readability formulas almost always
include a vocabulary component, typically in the form of a
calculation of the percentage of `long' words in the text. It is well established
in English that there is an inverse relationship between the length of a
word and its frequency of occurrence in the language, which means
that a text with a high proportion of long words is likely to
challenge the learners both linguistically and conceptually. Although of course
other factors in¯uence readability andlistenability as well, the use of a
readability formula in this way
illustrates
a vocabulary-assessment measure that is both comprehensive and embedded.
Context-independent
- context-dependent
The role of context, which is an old
issue in vocabulary testing, is the
basis
for the third dimension. Traditionally contextualisation has meant that a word is
presented to test-takers in a sentence rather than as an isolated
element. From a contemporary perspective, it is necessary to broaden
the notion of context to include whole texts and, more generally,
discourse. In addition, we need to recognise that contextualisation is
more than just a matter of the way in which vocabulary is
presented. The key question is to what extent the testtakers are being assessed
on the basis of their ability to engage with the context provided in
the test. In other words, do they have to make use of contextual
information in order to give the appropriate response to the test task, or can
they just respond as if the words were in isolation?
We can
illustrate the distinction by looking at a vocabulary item embedded in a
reading-comprehension test.
Humans have an
innate ability to recognise the taste of salt because it provides us
with sodium, an element which is essential to life. Although too
much salt in our diet may be unhealthy, we must consume a
certain amount of it to maintain our wellbeing.
What is the meaning of consume in
this text?
a.use up completely
b. eat or drink
c.spend wastefully
d.destroy
The point about
this test item is that all four options are possible meanings of the word
consume. Thus, the test-takers need some understanding of the context in order to
be con®dent that they have
chosen
the correct option, rather than simply relying on the fact that they have learned `eat
and drink' as the meaning of consume. To that extent, the item is
context dependent. I will explore this matter further in discussing
the vocabulary items in the Test of English as a Foreign Language
(TOEFL) in Chapter 5.
The issue of
context dependence also arises with cloze tests, in which words are
systematically deleted from a text and the testtakers' task is to write a
suitable word in each blank space. As we shall see in Chapter 4,
language testing researchers have debated whether cloze-test items can mostly
be answered correctly just by looking at the immediate context of the blank (the
phrase or clause in which it
occurs),
or whether it is necessary to draw on information from the wider context of the
passage in many cases. Some researchers have made detailed analyses
of the contextual information required to respond to individual cloze items, while
others have sought to show
more
globally that cloze-test items are or are not context dependent in a broad sense. Thus,
the degree of context dependence can be approached either as a characteristic of
individual test items or as a
property
of the test as a whole.
Generally
speaking, vocabulary measures embedded in writing and speaking tasks are
context dependent in that the learners are assessed on the appropriateness
of their vocabulary use in relation to the task. Judgements about
appropriateness take us beyond the text to consider the wider social context.
For instance, take a pro®ciency test in which the test-takers are doctors and
the test task is a role play
simulating
a consultation with a patient. If vocabulary use is one of the criteria used in
rating the doctors' performance, they need to demonstrate an ability
to meet the lexical requirements of the situation; for example: understanding
the colloquial expressions that patients use for common symptoms and ailments,
explaining medical concepts
in lay terms, avoiding medical jargon, offering reassurance to someone who is upset
or anxious, giving advice in a suitable tone and so on. Vocabulary
use in the task is thus in¯uenced by the doctor's status as a highly educated
professional, the expected role
relationship
in a consultation and the affective dimension of the situation. This is a
much broader view of context than we are used to thinking of in relation
to vocabulary testing, but a necessary one nonetheless if we are
to assess vocabulary in contemporary performance tests.
An
overview of the book
The three dimensions are not intended to
form a comprehensive model
of vocabulary assessment. Rather, they provide a basis for locating the variety of
assessment procedures currently in use within a common framework and,
in particular, they offer points of contact between tests which
treat words as discrete units and ones that assess vocabulary more integratively
in a task-based testing context. At
various
points through the book I refer to the dimensions and exemplify them. Since a
large proportion of work on vocabulary assessment to date has involved
instruments which are relatively discrete, selective and context independent in nature,
this approach may seem to
be
predominant in several of the following chapters. However, my aim is to present a
balanced view of the subject, and I discuss measures that are more embedded,
comprehensive and context dependent wherever the opportunity arises, and
especially in the last two
chapters
of the book.
Chapter 2 takes
up the question of what we mean by vocabulary. We tend to think of it
as consisting of individual words, as in the headwords of a
dictionary; however, even the de®nition of a `word' isby no means
straightforward. It is also necessary to consider lexical units that are larger
than single words, such as compound nouns,phrasal verbs, idioms and ®xed
expressions of various kinds. For
assessment
purposes, vocabulary is not just a set of linguistic units but also an attribute
of individual language learners, in the form of vocabulary knowledge
and the ability to access that knowledge for communicative purposes.
To explore
further the nature of vocabulary ability, in Chapter 3 I review the main lines
of enquiry by researchers on second language vocabulary acquisition.
Apart from the extensive work on methods of conscious vocabulary
learning, researchers are investigating how acquisition of word knowledge
occurs in a more incidental fashion
through
reading and listening activities. Other areas of interest are the ability of learners to
guess the meaning of unknown words which they encounter in their
reading, and the strategies they use to overcome gaps in their vocabulary
knowledge when engaged in speaking and writing tasks.
In Chapter 4 I
consider research in language testing that either has involved the
investigation of vocabulary tests or has a bearing on vocabulary assessment.
One issue in this area is whether the notion of a `pure' vocabulary
test is at all tenable. I trace the move away from discrete-point
vocabulary tests and look in some detail at the extent to which the cloze
procedure and its variants can be regarded as measures of vocabulary.
Much recent work on vocabulary testing has focused on estimating
how many words learners know (or their vocabulary size). A complementary
perspective is provided by other
studies
that seek to assess the quality (or `depth') of their vocabulary knowledge.
Chapter 5 presents case studies of four
vocabulary tests:
-
Nation's Vocabulary Levels Test;
-
Meara and Jones's
Eurocentres Vocabulary Size Test;
-
Paribakht and Wesche's
Vocabulary Knowledge Scale; and
-
the vocabulary items in
the Test of English as a Foreign Language (TOEFL).
In addition to being in¯uential
instruments in their own right, these tests exemplify several of the main
currents in vocabulary testing
discussed
in the previous chapter.
Practical issues
in the design of vocabulary tests are discussed in Chapter 6, which
focuses on relatively discrete and selective tests. The chapter includes
discussion of two speci®c examples of test design from my own experience.
One looks at some typical items for classroom progress tests, and the other is
an account of my efforts to
develop
a workable test to measure depth of vocabulary knowledge.
Chapter 7
focuses on comprehensive measures of vocabulary which can be used in
task-based language testing, particularly for embedded assessment. The largest
section of the chapter covers procedures that have been applied to
the assessment of learners' writing. These include `objective'
counts of the relative proportions of different types of word in a
composition, as well as `subjective' rating scales. I also consider the application
of comprehensive measures, such as readability formulas, to the analysis of
input material for tests involving
reading
and listening tasks.
Finally, in
Chapter 8, I look at current and future directions in work on vocabulary
assessment. This includes discussion of ways in which computer-based corpus
research can contribute to the development of vocabulary measures.
A second major theme is the need to
broaden
our view of the nature of vocabulary. More consideration should be given to the
role of multi-word lexical items in language use. Another priority
is to gain a better understanding of the vocabulary of speech, as distinct from
written language. There should also be more focus on the social dimension of
vocabulary use.
References
Purpura, james. 2004. ASSESSING GRAMMAR. United Kingdom: University Press Cambridge.
Read, John. 2000. ASSESSING VOCABULARY. United Kingdom: University Press Cambridge.