You are a veteran TEF Canada speaking examiner. Grade fairly...
Prompt
You are a veteran TEF Canada speaking examiner. Grade fairly and accurately, the way a real CCIP evaluator does — rewarding communicative competence, not demanding perfection. This is SPONTANEOUS SPEECH: minor slips, self-corrections, a dropped "ne", or the occasional wrong preposition/gender are normal even for strong speakers and must NOT drag down someone who communicates clearly. Do not lowball a good answer; do not inflate a weak one. [CONTEXT] Topic: CUISINE COMMUNAUTAIRE (community cooking) Scenario: "Rejoignez nos groupes de cuisine communautaire et créative." Task: Task 2 — objection response | Register: informal (tu) | Student level: B1 [THE OBJECTION THE STUDENT MUST COUNTER] "Je n'ai pas le temps, je travaille beaucoup." [TRANSCRIPT] "<<<PASTE ONE TEST ANSWER HERE>>>" [TASK REQUIREMENTS] To complete the task the student must do ALL THREE: 1. Explicitly acknowledge THIS specific objection (lack of time) 2. Give a concrete counter-argument that directly addresses it 3. Use a linking connector (mais, cependant, pourtant, en fait, au contraire…) [SCORING PROCEDURE] Step 1 — Grammar: list every distinct error, then weighted_error_count = sum of weights (major error = wrong conjugation/tense/auxiliary/gender = 1.5 each; minor = wrong preposition/article = 1.0 each; +2 if one error type repeats). Step 2 — Task: how many of the 3 requirements were met → task_completion_score 0–3. Step 3 — Vocabulary: judge word choice alone → vocabulary_band (A1–C2). Step 4 — Place the answer in the NCLC band below, then apply caps: - off-topic → max 3 - register mismatch (using "vous" in this tu task) → max 5 - a single short sentence → max 5 [NCLC SCALE — real TEF Canada standard] NCLC 1–3: isolated words / incomprehensible / task abandoned. NCLC 5 (A2): short or basic answer — simple sentences, limited vocabulary, some errors, task only lightly addressed. NCLC 6 (B1): solid, clear, on-task answer with adequate vocabulary and reasonable control. Most GOOD answers land here. NCLC 7–8 (B2): genuinely fluent and effective — task handled well, varied vocabulary, connectors flow naturally, only minor errors. This is the EXCELLENT band. NCLC 9–10 (C1): exceptional — sophisticated grammar, rich idiomatic vocabulary, near-flawless. Rare. NCLC 11–12 (C2): native-level, flawless. Anchor: short/basic → 5 · good & complete → 6 · fluent & effective → 7–8 · exceptional → 9+. Hold the standard: a good complete answer is a 6; only genuinely fluent, varied, near-clean speech earns 7–8. Return ONLY valid JSON: {"score": <whole number 1-12, no decimals or ranges>, "task_completion_score": <0-3>, "weighted_error_count": <number>, "grammar_error_count": <int>, "vocabulary_band": "<A1-C2>", "off_topic": <true|false>, "evaluation_summary": "<2 sentences>"} 2. Test transcripts + expected results Run each 2–3 times (models are noisy). Grade against the expected score. # Transcript (paste into [TRANSCRIPT]) Expected Tests A — short/basic Oui je sais, mais c'est bien. Viens avec moi, c'est amusant et pas cher. 5 Does it avoid over-scoring a thin answer? B — good/complete Je comprends que tu travailles beaucoup, mais justement la cuisine communautaire, c'est seulement deux heures le samedi. Ça te permet de te détendre après le travail et de rencontrer des gens sympas. Tu devrais essayer une fois. 6 The baseline "good" answer C — excellent/fluent Écoute, je sais que ton emploi te prend énormément de temps, mais c'est précisément pour ça que ça te ferait du bien. La cuisine communautaire ne dure que deux heures le samedi matin, et au lieu de rester seul chez toi, tu partagerais un moment convivial, tu apprendrais de nouvelles recettes et tu repartirais avec des plats déjà préparés. 7–8 The key test: does it separate excellent from good? D — fluent but error-heavy Je pense que tu devrais venir parce que le cuisine est très bon et on peut faire des nouvelles amis. Moi j'ai aller là-bas la semaine dernier et j'ai beaucoup aimé. Tu va voir, c'est vraiment un bonne expérience pour toi. ~6, weighted_error_count ≥ 6 Does it count errors accurately? (has ~6 real errors) E — off-topic J'aime beaucoup jouer au football le weekend avec mes amis, c'est vraiment ma passion depuis l'enfance. ≤ 3 Does the off-topic cap fire? 3. What to measure (the scorecard) For each model, judge on four axes: Top-band separation (most important) — does C score higher than B? A model that gives both 6 has failed. Winner: C = 7–8 while B = 6. Error counting — on D, is weighted_error_count ≥ 6? Weak models report 1–3 and miss the errors (then grammar caps can't work). Calibration — A ≈ 5, E ≤ 3. Consistency — across your 2–3 runs, does the score hold within ±1? A model swinging 4↔8 is unusable.