Rabu, 29 April 2020

ASSESSMENT 9 -10

Summary  BEYOND TESTS:
 ALTERNATIVES IN ASSESSMENT
          
              Halaman 251-280

 In the public eye, tests have acquired an aura of infallibility in our culture of mass producing everything, including the education of school children. Everyone wants a test for everything, especially if the test is cheap, quickly administered, and scored instantaneously. But we saw in Chapter 4 that while the standardized test industry has become a powerful juggernaut of influence on decisions about people's lives, it also has come under severe criticism from the public (Kohn, 2000). A more bal- anced viewpoint is offered by Bailey (1998, p. 204): "One of the disturbing things about tests is the extent to which many people accept the results uncritically, while others believe that all testing is invidious. But tests are simply measurement tools: It is the use to which we put their results that can be appropriate or inappropriate.

 It is clear by now that tests are one of a number of possible types of assessment. In Chapter 1, an important distinction was made between testing and assessing. Tests are formal procedures, usually administered within strict time limitations, to sample the performance of a test-taker in a specified domain. Assessment connotes a much broader concept in that most of the time when tcachers are teaching, they are also assessing. Assessment includes all occasions from informal impromptu obscrvations and comments up to and including tests,


 Early in the decade of the 1990s, in a culture of rebellion against the notion that all people and all skills could be mcasured by traditional tests, a novel concept emerged that began to be labeled "alternative" assessment. As teachers and students were becoming aware of the shortcomings of standardized tests, "an alternative to standardized testing and all the problems found with such testing" (Huerta-Macías, 1995, p. 8) was proposed. That proposal was to assemble additional measures of students-portfolios, journals, obscrvations, selfassessments, peerassessments, and the like-in an effort to triangulate data about students. For some, such alternatives held "ethical potential" (Lynch, 2001, p. 228) in their promotion of fairness and the balance of power relationships in the classroom.

 So they proposed to refer to "alternatives" in asses ment instead. Their term is a perfect fit within a model that considers tests as subset of assessment. Throughout this book, you have been reminded that all tes are assessments but, more important, that not all assessments are tests.
 The defining characteristics of the various alternatives in assessment that has been commonly used across the profession were aptly summed up by Brown ar Hudson (1998, pp. 654-655). Alternatives in assessments.

 1. require students to perform, create, produce, or do something;
2. use real-world contexts or simulations;
 3. are nonintrusive in that they extend the day-to-day classroom activities;
 4. allow students to be assessed on what they normally do in class every day;
 5. use tasks that represent meaningful instructional activities;
 6. focus on processes as well as products;
 7. tap into higher-level thinking and problem-solving skills:
 8. provide information about both the strengths and weaknesses of students:
 9. are multiculturally sensitive when properly administered;
10. ensure that people, not machines, do the scoring, using human judgment;
11. encourage open disclosure of standards and rating criteria; and
 12. call upon teachers to perform new instructional and assessment roles.

THE  DILEMMA OF MAXIMIZING BOTH PRACTICALITY ND WASHBACK.

The principal purpose of this chapter is to examine some of the alternatives assessment that are markedly different from formal tests. Tests, especially large-sca standardized tests, tend to be one-shot performances that are timed, multiple-choica decontextualized, norm-referenced, and that foster extrinsic motivation. On a other hand, tasks like portfolios, journals, and selfassessment are.

 • open-ended in their time orientation and format,
 • contextualized to a curriculum,
• referenced to the criteria (objectives) of that curriculum, and
 • likely to build intrinsic motivation .

 One way of looking at this contrast poses a challenge to you as a teacher: test designer. Formal standardized tests are almost by definition highly practical able instruments. They are designed to minimize time and money on the part of s designer and test-taker, and to be painstakingly accurate in their scori Alternatives such as portfolios, or conferencing with students on drafts of writing work, or observations of learners over time all require considerable time and effort on the part of the teacher and the student. Even more time must be spent if the teacher hopes to offer a reliable evaluation within students across time, as well as across students (taking care not to favor one student or group of students). But the alternative techniques also offer markedly greater washback, are superior formative measures, and, because of their authenticity, usually carry greater face validity.

This relationship can be depicted in a hypothetical graph that shows practicality/reliability on one axis and washback/authenticity on the other, as shown in Figure 10.1. Notice the implied negative correlation: as a technique increases in its washback and authenticity, its practicality and reliability tend to be lower. Conversely, the greater the practicality and reliability, the less likely you are to achieve beneficial washback and authenticity. I have placed three types of assess ment on the regression line to illustrate.

The figure appears to imply the inevitability of the relationship: large-scale multiple-choice tests cannot offer much washback or authenticity, nor can portfo- lios and such alternatives achieve much practicality or reliability. This need not be the case! The challenge that faces conscientious teachers and assessors in our pro- fession is to change the directionality of the line: to "flatten" that downward slope to some degree, or perhaps to push the various assessments on the chart leftward and upward. Surely we should not sit idly by, accepting the presumably inescapable conclusion that all standardized tests will be devoid of washback and authenticity. With some creativity and effort, we can transform otherwise inau- thentic and negative-washback-producing tests into more pedagogically fulfilling


Assessment learning experiences. A number of approaches to accomplishing this end are pos sible, many of which have already been implicitly presented in this book:
 • building as much authenticity as possible into multiple-choice task types and items
 • designing classroom tests that have both objectivescoring sections and open-ended response sections, varying the performance tasks
• turning multiple-choice test results into diagnostic feedback on areas of needed improvement
• maximizing the preparation period before a test to elicit performance rele- vant to the ultimate criteria of the test
• teaching test-taking strategics
 • helping students to see beyond the test: don't "teach to the test"
• triangulating information on a student before making a final assessment competence,

PERFORMANCE-BASED ASSESSMENT
           
Before proceeding to a direct consideration of types of aliternatives in assessment word about performance-based assessment is in order. There has been a gre deal of press in recent years about performance-based assessment, sometime merely called performance assessment (Shohamy, 1995; Norris et al., 1998). Is the different from what is being called "alternative assessment"?

 The push toward more performance-based assessment is part of the same g eral educational reform movement that has raised strong objections to using s dardized test scores as the only measures of student competencies (see. example, Valdez Pierce & O'Malley, 1992; Shepard & Bliem, 1993). The argument you can guess, was that standardized tests do not elicit actual performance on part of test-takers. If a child were asked, for example, to write a description ofe as seen from space, to work cooperatively with peers to design a three-dimensi model of the solar system, to explain the project to the rest of the class, and to a notes on a videotape about space travel, traditional standardized testing would involved in none of those performances. Performance-based assessment, however, would require the performance of the above-named actions, or samples thereof, which would be systematically evaluated through direct observation by a teacher and/or possibly by self and peers.
Performance-based assessment implies productive, observable skills, such as speaking and writing, of content-valid tasks. Such performance usually, but not always, brings with it an air of authenticity-real-world tasks that students have had time to develop. It often implies an integration of language skills, perhaps all four skills in the case of project work. Because the tasks that students perform are con- sistent with course goals and curriculum, students and teachers are likely to be more motivated to perform them, as opposed to a set of multiple-choice questions about facts and figures regarding the solar system.

O'Malley and Valdez Pierce (1996) considered performance-based assessment to be a subset of authentic assessment. In other words, not all authentic assessment is performance-based. One could infer that reading, listening, and thinking have many authentic manifestations, but since they are not directly observable in and of themselves, they are not performance-based. According to O'Malley and Valdez Pierce (p. 5), the following are characteristics of performance assessment:

 1. Students make a constructed response.
2. They engage in bigherorder thinking, with open-ended tasks.
 3. Tasks are meaningful, engaging, and autbentic.
 4. Tasks call for the integration of language skills.
5. Both process and product are assessed.
 6. Depth of a student's mastery is emphasized over breadth.

Performance-based assessment needs to be approached with caution. It is tempting for teachers to assume that if a student is doing something, then the process has fulfilled its own goal and the evaluator needs only to make a mark in the grade book that says "accomplished" next to a particular competency. In reality, per- formances as assessment procedures need to be treated with the same rigor as tra- ditional tests. This implies that teachers should

state the overall goal of the performance.
• specify the objectives (criteria) of the performance in detail,
• prepare students for performance in stepwise progressions,
 • use a reliable evaluation form, checklist, or rating sheet,
 • treat performances as opportunitics for giving feedback and provide that feedback systematically, and
 • if possible, utilize self- and peerassessments judiciously.

 To sum up, performance assessment is not completely synonymous with the con- cept of alternative assessment. Rather, it is best understood as one of the primary traits of the many available alternatives to assessment.

PORTFOLIOS

 One of the most popular alternatives in assessment, especially within a framework of communicative language teaching, is portfolio development. According to Genesce and Upshur (1996), a portfolio is "a purposeful collection of students' work that demonstrates... their efforts, progress, and achievements in given areas (p.99). Portfolios include materials such as

 essays and compositions in draft and final forms; reports, project outlines;
 • poetry and creative prose;
 • artwork, photos, newspaper or magazine clippings;
• audio and/or video recordings of presentations, demonstrations, etc.;
• journals, diaries, and other personal reflections ;
• tests, test scores, and written homework exercises;
 • notes on lectures; and
 • self- and peerassessments-comments, evaluations, and checklists.

 Until recently, portưolios were thought to be applicable only to younger children who assemble a portfolio of artwork and written work for presentation to a teacher and/or a parent. Now learners of all ages and in all fields of study are benefiting from the tangible, handson nature of portfolio development.

 Gottlieb (1995) suggested  a developmental scheme for considering the nature and purpose of portfolios, using the acronym CRADLE to designate six possible attributes of a portfolio:
Collecting
Reflecting
Assessing
Documenting
Linking
Evaluating
As Collections, portfolios are an expression of students' lives and identities. The appropriate freedom of students to choose what to include should be respected, but at the same time the purposes of the portfolio need to be clearly specified. Reflective practice through journals and self-assessment checklists is an important ingredient of a successful portfolio. Teacher and student both need to take the role of Assessment seriously as they evaluate quality and development over time. We need to recognize that a portfolio is an important Document in demonstrating stu- dent achievement, and not just an insignificant adjunct to tests and grades and other more traditional evaluation. A portfolio can serve as an important Link between stu- dent and teacher, parent, community, and peers; it is a tangible product, created with pride, that identifies a student's uniqueness. Finally, Evaluation of portfolios requires a time-consuming but fulfilling process of generating accountability.

The advantages of engaging students in portfolio development have been extolled in a number of sources (Genesee & Upshur, 1996,O'Malley & Valdez Pierce, 1996; Brown & Hudson, 1998; Weigle, 2002). A synthesis of those characteristics gives us a number of potential benefits. Portfolios

 • foster intrinsic motivation, responsibility, and ownership,
• promote student-teacher interaction with the teacher as facilitator,
 • individualize learning and celebrate the uniqueness of each student,
• provide tangible evidence of a student's work,
• facilitate critical thinking, selfassessment, and revision processes,
 • offer opportunities for collaborative work with peers, and
 • permit assessment of multiple dimensions of language learning.

At the same time, care must be taken lest portfolios become a haphazard pile of "junk" the purpose of which is a mystery to both teacher and student. Portfolios can fail if objectives are not clear, if guidelines are not given to students, if system- atic periodic review and feedback are not present, and so on. Sometimes the thought of asking students to develop a portfolio is a daunting challenge, especially for new teachers and for those who have never created a portfolio on their own.

Successful portfolio development will depend on following a number of steps and guidelines.

1.                          State objectives clearly. Pick one or more of the CRADLE attributes named above and specify them as objectives of developing a portfolio. Show how those purposes are connected to, integrated with, and/or a reinforcement of your already stated curricular goals. A portfolio attains maximum authenticity and washback when it is an integral part of a curriculum, not just an optional box of materials. Show students how their portfolios will include materials from the course they are taking and how that collection will enhance curricular goals.
2.                          Give guidelines on wbat materiais to include. Once the objectives have been determined, name the types of work that should be included. There is some disagreement among "experts" about how much negotiation should take place be- tween student and teacher over those materials. Hamp-Lyons and Condon (2000) suggested advantages for student control of portfolio contents, but teacher guid- ance will keep students on target with curricular objectives. It is helpful to give clear directions on how to get started since many students will never have com- piled a portfolio and may be mystified about what to do. A sample portfolio from a previous student can help to stimulate some thoughts on what to include.
3.                          Communicate assessment criteria to students. This is both the most im- portant aspect of portfolio development and the most complex. Two sources- self-assessment and teacher assessment-must be incorporated in order for students to receive the maximum benefit. Self-assessment should be as clear and simple as possible. O'Malley and Valdez Pierce (1996) suggested the following half- page selfevaluation of a writing sample (with spaces for students to write) for ele mentary school English language students.

Portfolio self-assessment questions (O'Maley & Valdez Pierce, 1996, p. 42)
1. Look at your writing sample.
a. What does the sample show that you can do?
b. Write about what you did well.
2. Think about realistic goals. Write one thing you need to do better. Be specific.
Genesee and Upshur (1996) recommended using a questionnaire format for selfassessment, with questions like the following for a project:

The teacher's assessment might mirror self-assessments, with similar questions designed to highlight the formative nature of the assessment. Conferences are important checkpoints for both student and teacher. In the case of requested written responses from students, help your students to process your feedback and show them how to respond to your responses. Above all, maintain reliability assessing portfolios so that all students receive equal attention and are assessed by the same criteria.
 An option that works for some contexts is to include peerassessment or smal group conferences to comment on one another's portfolios. Where the classroom community is relatively closely knit and supportive and where students are willing to expose themselves by revealing their portfolios, valuable feedback can be achieved from peer reviews. Such sessions should have clear objectives lest thes erode into aimless chatter. Checklists and questions may serve to preclude such= eventuality.
One could argue that it is inappropriate to reduce the personalized and cre- ative process of compiling a portfolio to a number or letter grade and that it is more appropriate to offer a qualitative evaluation for a work that is so open-ended. Such evaluations might include a final appraisal of the work by the student, with questions such as those listed above for selfassessment of a project, and a narrative evaluation of perceived strengths and weakness by the teacher. Those final evalu- ations should emphasize strengths but also point the way toward future learning challenges.
It is clear that portfolios get a relatively low practicality rating because of the time it takes for teachers to respond and conference with their students. Nevertheless, following the guidelines suggested above for specifying the criteria for evaluating portfolios can raise the reliability to a respectable level, and without ques- tion the washback effect, the authenticity, and the face validity of portfolios remain exceedingly high.
 In the above discussion, I have tried to subject portfolios to the same specifi- cations that apply to more formal tests: it should be made clear what the objectives are, what tasks are expected of the student, and how the learner's product will be evaluated . Strict attention to these demands is warranted for successful portfolio development to take place.

JOURNALS
Fifty years ago, journals had no place in the second language classroom. When lan guage production was believed to be best taught under controlled conditions, the concept of "free" writing was confined almost exclusively to producing essays or assigned topics, Today, journals occupy a prominent role in a pedagogical mode that stresses the importance of self- reflection in the process of students taking con trol of their own destiny.
 A journal is a log (or "account") of one's thoughts, feelings, reactions, assess ments, ideas, or progress toward goals, usually written with little attention to struc ture, form, or correctness. Learners can articulate their thoughts without the threa of those thoughts being judged later (usually by the teacher). Sometimes journals are rambling sets of verbiage that represent a stream of consciousness with no par ticular point, purpose, or audience. Fortunately, models of joumal use in educationa practice have sought to tighten up this style of journal in order to give them some focus (Staton et al. 1987). The result is the emergence of a number of overlapping categories or purposes in journal writing, such as the following:
language-learning logs grammar journals
• responses to readings
• strategies-based learning logs
 • selfassessment reflections
 • diaries of attitudes, feelings, and other affective factors
 • acculturation logs
Most classroom-oriented journals are what have now come to be known as dia logue journals. They imply an interaction between a reader (the teacher) and the student through dialogues or responses. For the best resuits, those responses shouid be dispersed across a course at regular intervals, perhaps weekly or biweekly. One of the principal objectives in a student's dialogue journal is to carry on a convers tion with the teacher. Through dialogue journals, teachers can become better acquainted with their students, in terms of both their learning progress and ther affective states, and thus become better equipped to meet students' individual needs.
 The following journal entry from an advanced student from China, and the teacher's response, is an illustration of the kind of dialogue that can take place.

 Dialogue journal sample

Journal entry by Ming Ling, China:

 Yesterday at about eight o'clock I was sitting in front of my table holding a fork and eating tasteless noodles which I usually really like to eat but I lost my taste yesterday because I didn' t feel well. I had a headache and a fever. My head seemed to be broken. I sometimes felt cold, sometimes hot. I didn't feel comfortable standing up and I didn't feel comfortable sitting down. I hated everything around me. It seemed to me that I got a great pressure from the atmosphere and I could not breath I was so sleepy since I had taken some medicine which functioned ay an antibiotic.
 The room was so quiet. I was there by myself and felt very solitary. This dinner reminded me of my mother. Whenever I was sick in China, my mother always took care of me and cooked rice gruel, which has to cook more than three hours and is very delicious, I think. I would be better very soon under the care of my mother. But yesterday, I had to cook by myself even though I was sick, The more I thought, the les I wanted to eat, Haif an hour passed. The noodles were cold, but I was still sitting there and thinking about my mother. Finally I threw out the noodles and went to bed.
 Teacher's response:

This is a powerful piece of writing because you really communicate what you were feeling. You used vivid details, like "cating tasteless noodles," "my head seemed to be broken" and "rice gruel, which has to cook more than three hours and is very delicious." These make it easy for the reader to picture exactly what you were going through. The other strong point about this piece is that you bring the reader full circie by beginning and ending with "the noodles."

Being alone when you are sick is difficult. Now, I know why you were so quiet in class.
 If you want to do another entry related to this one, you could have a dialogue with your "sick" self. What would your "healthy" self say to the "sick" self? Is there some advice that could be exchanged about how to prevent illness or how to take care of yourself better when you do get sick? Start the dialogue with your "sick" self speaking first.
With the widespread availability of Internet communications, journals and other student-teacher dialogues have taken on a new dimension. With such inno- vations as "collaboratories" (where students in a class are regularly carrying on email discussions with each other and the teacher), on-line education, and distance learning, journals-out of several genres of possible writing-have gained additional prominence.
Journals obviousty serve important pedagogical purposes: practice in the mechanics of writing, using writing as a "thinking" process, individualization, and communication with the teacher. At the same time, the assessment qualities of journal writing have assumed an important role in the teaching-learning process. Because most journals are-or should be-a dialogue between student and teacher, they afford a unique opportunity for a teacher to offer various kinds of feedback.
On the other side of the issue, it is argued that journals are too free a form to be assessed accurately. With so much potential variability, it is difficult to set up cri- teria for evaluation. For some English language learners, the concept of free and unfettered writing is anathema. Certain critics have expressed ethical concerns: stu- dents may be asked to reveal an inner self, which is virtually unheard of in their own culture. Without a doubt, the assessing of journal entrics through responding is not an exact science.
It is important to turn the advantages and potential drawbacks of journals into positive general steps and guidelines for using journals as assessment instruments. The following steps are not coincidentally parallel to those cited above for portfolio development:
 1. Sensitively tntroduce students to the concept of journal writing. For many students, especially those from educational systems that play down the notion of teacher-student dialogue and collaboration, journal writing will be difficult at first. University-level students, who have passed through a dozen years of product writ- ing, will have particular difficulty with the concept of writing without fear of a teacher's scrutinizing every grammatical or spelling error. With modeling, assur- ance, and purpose, however, students can make a remarkable transition into the po- tentially liberating process of journal writing. Students who are shown examples of journal entries and are given specific topics and schedules for writing will become comfortable with the process.
 2. State the objective(s) of the journal. Integrate journal writing into the ob- jectives of the curriculum in some way, especially if journal entries become topics of class discussion. The list of types of journals at the beginning of this section may coincide with the following examples of some purposes of journals:
Language-learning logs. In English language teaching, learning logs have the advantage of sensitizing students to the importance of setting their own goals and then self-monitoring their achievement. McNamara (1998) suggested restricting the number of skills, strategies, or language categories that students comment on; oth erwise students can become overwhelmed with the process. A weekly schedule a limited number of strategies usually accomplishes the purpose of keeping stu dents on task.
Grammar journals. Some journals are focused only on grammar acquisition These types of journals are especially appropriate for courses and workshops tha focus on grammar. "Error logs" can be instructive processes of consciousness raising for students: their successes in noticing and treating errors spur them to maintain the process of awareness of error.
Responses to readings. Journals may have the specified purpose of simpie responses to readings (and/or to other material such as lectures, presentations, film and videos). Entries may serve as precursors to freewrites and help learners tos out thoughts and opinions on paper. Teacher responses aid in the further develop ment of those ideas.
Strategies-based learning logs. Closely allied to language-learning logs are spe cialized journals that focus only on strategies that learners are seeking to become aware of and to use in their acquisition process. In H. D. Brown's (2002) Strategie. for Success: A Practical Guide to Learning Englisb, a systematic strategies-basec journal-writing approach is taken where, in each of 12 chapters, learners become aware of a strategy, use it in their language performance, and reflect on that proces in a journal.
Selfassessment reflections. Journals can be a stimulus for self-assessment in : more open-ended way than through using checklists and questionnaires. With the possibility of a few stimulus questions, students journals can extend beyond the scope of simple one-word or one-sentence responses.
Diaries of attitudes, feelings, and other affective factors. The affective states o learners are an important element of self-understanding. Teachers can thereby become better equipped to effectively facilitate learners' individual journeys toward their goals Acculturation logs. A variation on the above affectively based journals is one that focuses exclusively on the sometimes difficult and painful process of accultur ation in a non-native country. Because culture and language are so strongly linked awareness of the symptoms of acculturation stages can provide keys to eventual lan guage success.
 3. Give guidelines on what kinds of topics to include. Once the purpose or type of journal is clear, students will benefit from models or suggestions on what kinds of topics to incorporate into their journals.
 4. Carefully specify the criteria for assessing or grading journals. Students need to understand the freewriting involved in journals, but at the same time, they need to know assessment criteria. Once you have clarified that journals will not be evaluated for grammatical correctness and rhetorical conventions, state how they will be evaluated. Usually the purpose of the journal will dictate the major assess ment criterion. Effort as exhibited in the thoroughness of students' entries will no doubt be important. Also, the extent to which entries reflect the processing of course content might be considered. Maintain reliability by adhering conscien tiously to the criteria that you have set up.
 5. Provide optimal feedback in your responses. McNamara (1998, p. 39) rec ommended three different kinds of feedback to journals:
 1. cheerleading feedback, in which you celebrate successes with the students or encourage them to persevere through difficulties,
2. instructional feedback, in which you suggest strategies or materials, suggest ways to fine-tune strategy use, or instruct students in their writing, and
 3. reality-check feedback, in which you help the students set more realistic expectations for their language abilities.
The ultimate purpose of responding to student journal entries is well captured in McNamara's threefold classification of feedback. Responding to journals is a very personalized matter, but closely attending to the objectives for writing the journal and its specific directions for an entry will focus those responses appropriately.
Peer responses to journals may be appropriate if journal comments are rela- tively "cognitive," as opposed to very personal. Personal comments could make stu- dents feel threatened by other pairs of eyes on their inner thoughts and feelings.
6. Designate appropriate time frames and schedules for review. Journals. like portfolios, need to be esteemed by students as integral parts of a course.There- fore, it is essential to budget enough time within a curriculum for both writing jour- nals and for your written responses. Set schedules for submitting journal entries periodically; return them in short order.

 CONFERENCES AND INTERVIEWS

 For a number of years, conferences have been a routine part of language classrooms, especially of courses in writing. In Chapter 9, reference was made to conferencing as a standard part of the process approach to teaching writing, in which the teacher, in a conversation about a draft, facilitates the improvement of the written work. Such interaction has the advantage of one-on-one interaction between teacher and student, and the teacher's being able to direct feedback toward a student's specific needs. Conferences are not limited to drafts of written work. Including portfolios and journals discussed above, the list of possible functions and subject matter for con- ferencing is substantial:

commenting on drafts of essays and reports
 • reviewing portfolios
• responding to journals
 • advising on a student's plan for an oral presentation
 • assessing a proposal for a project
• giving feedback on the results of performance on a test
• carifying understanding of a reading
• exploring strategies-based options for enhancement or compensation
 • focusing on aspects of oral production
• checking a student's self-assessment of a performance
• setting personal goals for the near future
• assessing general progress in a course

Conferences must assume that the teacher plays the role of a facilitator and guide, not of an administrator, of a formal assessment. In this intrinsically motivating atmosphere, students need to understand that the teacher is an ally who is encour- aging self-reflection and improvement. So that the student will be as candid as pos- sible in self-assessing, the teacher should not consider a conference as something to be scored or graded. Conferences are by nature formative, not summative, and their primary purpose is to offer positive washback.

 Discussions of alternatives in assessment usually encompass one specialized kind of conference: an interview. This term is intended to denote a context in which a teacher interviews a student for a designated assessment purpose. (We are not talking about a student conducting an interview of others in order to gather information on a topic.) Interviews may have one or more of several possible goals, in which the teacher.

• assesses the student's oral production,
• ascertains a student's needs before designing a course or curriculum,
• seeks to discover a student's learning styles and preferences,
• asks a student to assess his or her own performance, and
• requests an evaluation of a course.

Because interviews have multiple objectives, as noted above, it is difficult to generalize principles for conducting them, but the following guidelines may help to frame the questions efficiently:

1. Offer an initial atmosphere of warmth and anxiety-lowering (warm-up).
2. Begin with relatively simple questions.
 3. Continue with level-check and probe questions, but adapt to the interviewee as needed.
4. Frame questions simply and directly.
5. Focus on only one factor for each question. Do not combine several objec- tives in the same question.
 6. Be prepared to repeat or reframe questions that are not understood,
7. Wind down with friendly and reassuring closing comments.

How do conferences and interviews score in terms of principles of assessment Their practicality, as is true for many of the alternatives to assessment, is low because they are time-consuming. Reliability will vary between conferences and interviews In the case of conferences, it may not be important to have rater reliability because the whole purpose is to offer individualized attention, which will vary greatly from student to student. For interviews, a relativcly high level of reliability should be maintained with careful attention to objectives and procedures. Face validity for both can be maintained at a high level due to their individualized nature. As long as the subject matter of the conference/interview is clearly focused on the course and course objectives, content validity should also be upheld. Washback potential and authenticity are high for conferences, but possibly only moderate for interviews unless the results of the interview are clearly folded into subsequent learning.

OBSERVATIONS

All teachers, whether they are aware of it or not, observe their students in the class room almost constantly. Virtually every question, every response, and almost every nonverbal behavior is, at some level of perception, noticed. All those intuitive per- ceptions are stored as little bits and pieces of information about students that can orm a composite impression of a student's ability. Without ever administering a test or a quiz, teachers know a lot about their students. In fact, experienced teachers are so good at this almost subliminal process of assessment that their estimates of a stu- dent's competence are often highly correlated with actual independently adminis- ered test scores. (Sece Acton, 1979, for an example.)

 How do all these chunks of information become stored in a teacher's brain cells? Usually not through rating sheets and checklists and carefully completed observation charts. Still, teachers' intuitions about students' performance are not nfallible, and certainly both the reliability and face validity of their feedback to stu- dents can be increased with the help of empirical means of obscrving their lan- guage performance. The value of systematic observation of students has been extolled for decades (Flanders, 1970; Moskowitz, 1971; Spada & Frölich, 1995), and ts utilization greatly enhances a teacher's intuitive impressions by offering tangible corroboration of conclusions. Occasionally, intuitive information is disconfirmed by observation data.

We will not be concerned in this section with the kind of observation that rates formal presentation or any other prepared, prearranged performance in which the student is fully aware of some evaluative measure being applied, and in which the eacher scores or comments on the performance. We are talking about observation as a systematic, planned procedure for real-time, almost surreptitious recording of student verbal and nonverbal behavior. One of the objectives of such obscrvation is o assess students without their awareness (and possible consequent anxiety) of the observation so that the naturalness of their linguistic performance is maximized.

The list could be even more specific to suit the characteristics of students, the focus of a lesson or module, the objectives of a curriculum, and other factors.The list might expand, as well, to include other possible observed performance. In order to carry out classroom observation, it is of course important to take the following steps:

 1. Determine the specific objectives of the observation.
 2. Decide how many students will be observed at one time.
3. Set up the logistics for making unnoticed observations.
 4. Design a system for recording observed performances.
 5. Do not overestimate the number of different elements you can observe at one time-keep them very limited.
 6. Plan how many observations you will make.
 7. Determine specifically how you will use the results.
 Designing a system for observing is no simple task. Recording your observa- tions can take the form of anecdotal records, checklists, or rating scales. Anecdotal records should be as specific as possible in focusing on the objective of the obser- vation, but they are so varied in form that to suggest formats here would be coun- terproductive.Their very purpose is more note-taking than record-keeping. The key is to devise a system that maintains the principle of reliability as closely as possible.
 Checklists are a viable alternative for recording observation results. Some check- lists of student classroom performance, such as the COLT observation scheme devised by Spada and Fröhlich (1995), are elaborate grids referring to such variables as
 whole-class, group, and individual participation,
• content of the topic,
 • linguistic competence (form, function, discourse, sociolinguistic),
 • materials being used, and
• skill (listening, speaking, reading, writing),
with subcategories for each variable. The observer identifies an activity or episode as well as the starting time for each, and checks appropriate boxes along the grid Completing such a form in real time may present some difficulty with so many fac- tors to attend to at once.
 Checklists can also be quite simple, which is a better option for focusing on only a few factors within real time. On one occasion I assigned teachers the task of noting occurrences of student errors in third-person singular, plural, and ing mor- phemes across a period of six weeks. Their records needed to specify only the number of occurrences of each and whether each occurrence of the error was ignored, treated by the teacher, or self-corrected. Believe it or not, this was not an casy task! Simpły noticing errors is hard enough, but making entries on even a very simple checklist required careful attention. The checklist looked like this:
 Each of the 30-odd checklists that were eventually completed represented a two- hour class period and was filled in with "ticks" to show the occurrences and the follow-up in the appropriate cell.
Rating scales have also been suggested for recording observations. One type of rating scale asks teachers to indicate the frequency of occurrence of target perfor- mance on a separate frequency scale (always = 5; never = 1). Another is a holistic assessment scale, like the TWE scale described in the previous chapter or the OPI scale discussed in Chapter 7, that requires an overall assessment within a number of categories (for example, vocabulary usage, grammatical correctness, fluency).
If you scrutinize observations under the microscope of principles of assess- ment, you will probably find moderate practicality and reliability in this type of pro- cedure, especially if the objectives are kept very simple. Face validity and content validity are likely to get high marks since observations are likely to be integrated into the ongoing process of a course. Washback is only moderate if you do little follow- up on observing. Some obscrvations for research purposes may yield no washback whatever if the researcher simply disappears with the information and never com- municates anything back to the student. But a subsequent conference with a student Assessment can then yield very high washback as the student is made aware of empirical data on targeted performance. Authenticity is high because, if an observation goes rela tively unnoticed by the student, then there is little likelihood of contrived contexts or playacting.

SELF- AND PEER-ASSESSMENTS

 A conventional view of language assessment might consider the notion of self and peer-assessment as an absurd reversal of politically correct power relation ships. After all, how could learners who are still in the process of acquisition, es- pecially the early processes, be capable of rendering an accurate assessment their own performance? Nevertheless, a closer look at the acquisition of any skill reveals the importance, if not the nccessity, of self-assessment and the benefit of peer-assessment. What successful learner has not developed the ability to monitor his or her own performance and to use the data gathered for adjustments and cor- rections? Most successful learners extend the learning process well beyond the classroom and the presence of a teacher or tutor, autonomously mastering the ant of self-assessment. Where peers are available to render assessments, the advantage of such additional input is obvious.
Self-assessment derives its theoretical justification from a number of well- established principles of second language acquisition. The principle of autonomy stands out as one of the primary foundation stones of successful learning. The ability to set one's own goals both within and beyond the structure of a classroom curriculum, to pursue them without the presence of an external prod, and to inde- pendently monitor that pursuit are all keys to success. Developing intrinsic moti- vation that comes from a self-propelled desire to excel is at the top of the list ef successful acquisition of any set of skills.
Peerassessment appeals to similar principles, the most obvious of which is coop erative learning. Many people go through a whole regimen of education from kindergarten up through a graduate degree and never come to appreciate the value collaboration in learning-the benefit of a community of learners capable of teaching each other something. Peerassessment is simply one arm of a plethora of tasks and procedures within the domain of learner-centered and collaborative education.
 Researchers (such as Brown & Hudson, 1998) agree that the above theoretical underpinnings of self- and peer-assessment offer certain benefits: direct involvement of students in their own destiny, the encouragement of autonomy, and increased motivation because of their self-involvement . Of course, some noteworthy draw backs must also be taken into account. Subjectivity is a primary obstacle to over come. Students may be either too harsh on themselves or too self-flattering, or they may not have the necessary tools to make an accurate assessment. Also, especially in the case of direct assessments of performance (see below), they may not be abie to discern their own errors. In contrast, Bailey (1998) conducted a study in which learners showed moderately high correlations (between 58 and .64) between self rated oral production ability and scores on the OPI, which suggests that in th assessment of general competence, learners' self-assessments may be more accurat than one might suppose.

Types of Self- and Peer-Assessment

 It is important to distinguish among several different types of self- and peer-assessmen and to apply them accordingly. I have borrowed from widely accepted classifica tions of strategic options to create five categories of self- and peer-assessment (1) direct assessment of performance, (2) indirect assessment of performance (3) metacognitive assessment, (4) assessment of socioaffective factors , and (5) stu dent self-generated tests.
Brown's (1999) New Vistas series offers end-of chapter selfevaluation checkli that give students the opportunity to think about the extent to which they ha reached a desirable competency level in the specific objectives of the unit. Figure 10 shows a sample of this "checkpoint" feature. Through this technique, students a reminded of the communication skills they have been focusing on and are given chance to identify those that are essentially accomplished, those that are not yet f filled, and those that need more work. The teacher follow-up is to spend more tin on items on which a number of students checked "sometimes" or"not yet," or possib to individualize assistance to students working on their own points of challenge.
1.       Socioaffective assessment. Yet another type of self- and peer-assessme comes in the form of methods of examining affective factors in learning. Such sessment is quite different from looking at and planning linguistic aspects of acq sition. It requires looking at oneself through a psychological lens and may not diff greatly from self-assessment across a number of subject-matter areas or for any s of personal skills. When learners resoive to assess and improve motivation, to gau and lower their own anxiety, to find mental or emotional obstacles to learning ar then plan to overcome those barriers, an all-important socioaffective domain is voked. A checklist form of such items may look like many of the questionnaire iten in Brown (2002), in which test-takers must indicate preference for one stateme over the one on the opposite side:
process of constructing tests themselves. The traditional view of what a test is would never allow students to engage in test construction, but student-generated tests can be productive, intrinsically motivating, autonomy-building processes.
 Gorsuch (1998) found that student-generated quiz items transformed routine weekly quizzes into a collaborative and fulfilling experience. Students in smail groups were directed to create content questions on their reading passages and to collectively choose six vocabulary items for inclusion on the quiz. The process of creating questions and choosing lexical items served as a more powerful reinforce- ment of the reading than any teacher-designed quiz could ever be. To add further interest, Gorsuch directed students to keep records of their own scores to plot their progress through the term.
Murphey (1995), another champion of self- and peer-generated tests, success- fully employed the technique of directing students to generate their own lists of words, grammatical concepts, and content that they think are important over the course of a unit. The list is synthesized by Murphey into a list for review, and all items on the test come from the list. Students thereby have a voice in determining the content of tests. On other occasions, Murphey has used what he calls "interac- tive pair tests" in which students assess each other using a set of quiz items. One stu- dent's response aptly summarized the impact of this technique:

 We had a test today. But it was not a test, because we could study for it beforehand. i gave some questions to my partner and my partner gave me some questions. And we studenty decided what grade we should get. I hate tests, but I like this kind of test. So please don't give us a surprise test. I think, that kind of test that we did today is more useful for me than a surprise test because I study for it.
Many educators agree that one of the primary purposes in administering tests is to stimulate review and integration, which is exactly what student-generated testing does, but almost without awareness on the students' part that they are reviewing the material. I have seen a number of instances of teachers successfully facilitating students in the self-construction of tests. The process engenders intrinsic involvement in reviewing objectives and selecting and designing items for the final form of the test. The teacher of course needs to set certain parameters for such a project and be willing to assist learners in designing items.

Guidelines for Self- and Peer-Assessment

Self- and peerassessment are among the best possible formative types of assess- ment and possibly the most rewarding, but they must be carefully designed and administered for them to reach their potential. Four guidelines will help teachers bring this intrinsically motivating task into the classroom successfully.
1.       Tell students the purpose of the assessment. Self-assessment is a process that many students-especially those in traditional educational systems-will ini- tially find quite uncomfortable. They need to be sold on the concept. It is therefore essential that you carefully analyze the needs that will be met in offering both self- and peerassessment opportunities, and then convey this information to students.
2.        Define the task(s) clearly. Make sure the students know exactly what they are supposed to do. If you are offering a rating sheet or questionnaire, the task is not complex, but an open-ended journal entry could leave students perplexed about what to write. Guidelines and models will be of great help in clarifying the procedures.
3.       Encourage impartial evaluation of performance or ability. One of the greatest drawbacks to self-assessment is the threat of subjectivity. By showing stu- dents the advantage of honest, objective opinions, you can maximize the beneficial washback of self-assessments. Peer-assessments, too, are vulnerable to unreliability as students apply varying standards to their peers. Clear assessment criteria can go a long way toward encouraging objectivity.
4.       Ensure beneficial wasbback througb follow-up tasks. It is not enough to simply toss a self-checklist at students and then walk away. Systematic follow-up can be accomplished through further selfanalysis, journal reflection, written feedback from the teacher, conferencing with the teacher, purposeful goal-setting by the stu- dent, or any combination of the above.

A  Taxonomy of Self- and Peer-Assessment Tasks

 To sum up the possibilities for self- and peerassessment, it is helpful to consider a variety of tasks within cach of the four skills.

 Self- and peer-assessment tasks

Listening Tasks

listening to TV or radio broadcasts and checking comprehension with a partner listening to bilingual versions of a broadcast and checking comprehension asking when you don't understand something in pair or group work listening to an academic lecture and checking yourself on a "quiz" of the content setting goals for creating/increasing opportunities for listening

 Speaking Tasks

filling out student self-checklists and questionnaires using peer checklists and questionnaires rating someone's oral presentation (holistically) detecting pronunciation or grammar errors on a self- recording asking others for confirmation checks in conversational settings setting goals for creating/increasing opportunities for speaking

An evaluation of self- and peer-assessment according to our classic principles of assessment yiclds a pattern that is quite consistent with other alternatives to assessment that have been analyzed in this chapter. Practicality can achieve a mod- erate level with such procedures as checklists and questionnaires, while reliability risks remaining at a low level, given the variation within and across learners. Once students accept the notion that they can legitimately assess themsclves, then face validity can be raised from what might otherwise be a low level. Adherence to course objectives will maintain a high degree of content validity. Authenticity and washback both have very high potential because students are centering on their own linguistic needs and are receiving useful fecdback.

pany such a chart is that none of the evaluative "marks" should be considered permanent or unchangeable. In fact, the challenge that was presented at the begin ning of the chapter is reiterated here: take the "low" factors in the chart and create assessment procedures that raise those marks.

Perhaps it is now clear why "alternatives in assessment" is a more appropriate phrase than "alternative assessment." To set traditional testing and alternatives against each other is counterproductive. All kinds of assessment, from formal con ventional procedures to informal and possibly unconventional tasks, are needed to assemble information on students. The alternatives covered in this chapter may not be markedly different from some of the tasks described in the preceding four chap ters (assessing listening, speaking, reading, and writing). When we put all of this together, we have at our disposal an amazing array of possible assessment tasks for second language learners of English. The alternatives presented in this chapter simply expand that continuum of possibilities.

 EXERCISES

 [Note: () Individual work; (G) Group or pair work; (C) Wholeclass discussion.]
 1. (C) Using Brown and Hudson's (1998) 12 characteristics of alternatives in assessment (quoted at the beginning of the chapter), discuss the differences between traditional and "alternative" assessment. Some performance assess- ments are relatively traditional (oral interview, essay writing, demonstrations), yet they fit most of the criteria for alternatives in assessment. In this light, iden- tify a continuum of assessments, ranging from highly traditional to alternative.
2. (G) In a small group, refer to Figure 10.1, which depicts the relationship between practicality/reliability and authenticity/washback. With each group assigned to a separate skill area (L, S, R, W), select perhaps 10 or 12 techniques that were described earlier in this book and place them into this same graph. Show your graph to the rest of the class and explain.
 3. (G) In pairs or groups assigned to procure a sample of a portfolio from a teacher you know, or from a school you have some connection with, evaluate the portfolio on as many of the seven guidelines (pages 257-259) as possible. Present the portfolio and your evaluation to the rest of the class.
4. (G) In pairs or groups, follow the same procedure as #3 above for a journal.
5. a/C) If possible, observe a teacher-student conference or a student-student peer-assessment. The most common type of conference might be over a draft of an essay. Report back to the class on what you observed and offer an evalu- ation of its effectiveness.
6. (1/C) Plan to observe an ES/FL class. Select specific students to observe, and define the form of linguistic performance you will focus on. Such an observati could include attention to students' processing of the teacher's error treatmen Report your findings to the class.
7. (C) Look at the self-assessments in Figures 10.2, 10.3, and 10.4. Evaluate the effectiveness in terms of the guidelines offered in this chapter.
8. (G) At the end of the chapter, Table 10.1 offers a broad estimate of the exte to which the alternatives to assessment in this chapter measure up to basic principles of assessment. In pairs or small groups, each assigned to one of th six alternatives, decide whether you agree with these evaluations. Defend your decisions and report them to the rest of the class.

FOR YOUR FURTHER READING

Brown, J.D. (Ed.) (1998). New ways of classroom assessment. Alexandria, Teachers of English to Speakers of Other Languages. This volume in TESOL's "New Ways" series offers an array of nontraditional assessment procedures. Each procedure indicates its appropriate level, objective, class time required, and suggested preparation time. Încluded are examples of portfolios, journals, logs, conferences, and self- and peer- assessment. Alternatives to traditional assessment of listening, speaking, reading, and writing are also given. All techniques were contributed by teachers in varying contexts around the world.

 O'Malley, J. Michael, and Valdez Pierce, Lorraine. (1996.) Authentic assessment English language learners: Practical approaches for teachers. White Plains, N Addison-Wesley. This practical guide for teachers targets English Language Learners (ELLS) from K-12, but has applications beyond this context. It is a valuable col- lection of techniques and procedures for carrying out performance assess- ments that are authentic and that offer beneficial washback to learners. It contains reproducible checklists, rating scales, and charts that can be adapted to one's own context. It offers a comprehensive treatment of port- folios, journals, observations, and self- and peer-assessments.

TESOL Journal 5 (Autumn, 1995). Special Issue on Alternative Assessment. This entire issue is devoted to alternatives in assessment. Included are arti- cles from practicing teachers on portfolios, self-assessment, collaborative teacher assessment, test review activities, and general reflections on the benefits of assessment that promotes collaboration and reflection.



REFERENCES :


Brown H. Douglas san, 2003,Language assessment principles and classroom practices,Francisco,california. 


ASSESSMENT FOR MEETING 15

            Assessing grammar,vocabulary Chapter one Differing notions of ‘grammar’ for assessment Introduction   Even...