P38 — BLINDED EXTERNAL DESIGN CHALLENGE SEARCH-OFF / P-ONLY....
Prompt
P38 — BLINDED EXTERNAL DESIGN CHALLENGE SEARCH-OFF / P-ONLY. Do not use Web Search, browsing, files, prior chats, sibling outputs, provider identity, or outside facts. Evaluate only the target design clauses and probes below. You are one opaque external evaluator slot. Your primary output must be independent: do not receive, request, infer, quote, summarize, or react to any sibling slot output before your primary response is frozen. PURPOSE This is a falsification task, not an invitation to praise the design. Find the smallest text-compliant harmful paths that survive the strongest closing clauses. Do not invent missing thresholds, owners, permissions, implementation facts, domain facts, or capabilities. UNKNOWN is not ABSENT. A record, label, hash, status field, role title, or scheduler event is not by itself proof that its stated semantic condition holds. Agreement is descriptive unless independence is positively established. AUDIT METHOD For each candidate defect: (1) name the violated invariant; (2) give the smallest trigger/precondition; (3) trace the harmful consequence; (4) identify the exact affected scope; (5) search the whole payload for the strongest closing clause; (6) reject the candidate if that clause actually closes it; (7) otherwise give the smallest falsifiable repair; (8) provide one KILL-TEST that fails the unmodified design and one BENIGN-PRESERVATION case your repair must not break. Preserve severe singleton counterexamples even against a favorable majority. Do not convert UNAVAILABLE, PARTIAL, UNKNOWN, EXPIRED, UNTESTED or disagreement into PASS or absence. OUTPUT CONTRACT Begin exactly with `1. EXECUTIVE VERDICT` as the first bytes. No preamble, scratchpad, planning, self-talk, chain-of-thought, or visible hidden reasoning. Output exactly these six sections in order: 1. EXECUTIVE VERDICT — 3–6 sentences and exactly one DOMINANT NEXT ACTION. 2. TOP FINDINGS — max 6. For each: ID, SEEDED/UNSEEDED, violated invariant, trigger, harmful consequence, scope, strongest closing clause and why it does/does not close, minimal repair. 3. PROBE MATRIX — every named probe below, one row each: BLOCKED / BYPASS / UNRESOLVED / ALLOWED-BENIGN, with exact clause(s) and dependency-scoped consequence. 4. FALSE-POSITIVE / BENIGN-NEIGHBOR CHECK — at least 3 rejected candidates and at least 3 benign near-neighbors that remain usable. 5. MINIMAL REPAIRS + TESTS — max 6; disposition REUSE / MERGE / EXTEND / NEW / DEFER; include KILL-TEST and BENIGN-PRESERVATION for every repair. 6. FINAL SCOPE — exact-payload design scope only; no implementation, provider-identity, deployment, medical/clinical diagnosis/treatment, legal, or release authority. Rules: findings are not votes; do not rank providers; no multiplicative improvement claim without matched-budget evidence; same-slot restatements are correlated; do not treat compliance wording as proof of runtime behavior. TARGET DESIGN — P38.1: PARTIAL PANEL / INFORMATION ISOLATION / VERIFIER PACKET / EXTERNAL EVALUATOR A. CELL TERMINAL STATES AND PANEL INTERPRETATION A1. Every expected panel cell must end in exactly one explicit terminal state: AVAILABLE-FROZEN, VALID-NONE, PARTIAL/TRUNCATED, FORMAT-ERROR, DUPLICATE-HOLD, or UNAVAILABLE. UNAVAILABLE is not a negative result, not no-finding, and is never imputed. A2. Synthesis may proceed on a degraded panel only if panel completeness is explicitly carried into every absence, saturation, redundancy and contribution interpretation. A degraded panel cannot support a whole-panel absence claim for cells that never produced admissible primary content. A3. A VALID-NONE cell is an admissible frozen primary output stating that the evaluator found no qualifying defect under the contract. It is distinct from silence, timeout, truncation, refusal and transport loss. A4. Late recovery of an unavailable/partial cell is BACKFILL/LATE-MISS evidence under a new receipt event. It can update a later synthesis only with explicit provenance and never silently rewrites the already frozen primary result. B. PRE-FREEZE INFORMATION ISOLATION B1. Before all primary cells reach terminal state, no slot may receive another slot's primary output or derivative, including passive/orchestrator forwarding, summaries, rankings, extracted findings, hints, suggested wording or "clarifications." Any such exposure is correlation evidence and must be recorded. B2. Exposure does not automatically erase the exposed output, but the exposed output cannot be counted as an independent primary discovery for the affected construct unless a fresh unexposed rerun exists. Post-freeze supplementary evidence never retroactively becomes original primary evidence. B3. After primary freeze, sparse challenge is allowed. New post-freeze material must carry supplementary provenance and may challenge or refine a frozen finding, but it cannot silently enlarge the frozen primary sample. B4. Exposure classification is content-based, not label-based: renaming answer-bearing material as formatting guidance, context, a hint, a rubric clarification or a title list does not restore information isolation. C. DISSENT-COMPLETE BLIND VERIFIER PACKET C1. A privileged verifier receives the frozen claim plus the evidence actually needed to test it, including preserved relevant dissent/counterevidence and incompatible assumptions. Blindness withholds provider identity, majority framing, desired answer and persuasion, not load-bearing counterevidence. C2. Omitting preserved material dissent from the verifier packet makes certification PARTIAL until the omission is resolved. A curator cannot select only favorable evidence and call the packet blind. C3. Exact applicable counterexamples close only by applicable adjudication: deterministic replay where possible, or a positively qualified adjudicator for semantic/method-validity disputes. Majority weight, non-response, challenge unavailability or repeated restatement does not refute an exact counterexample. C4. Freeze an initial verdict/evidence basis before interactive challenge. Post-decision argument may update a verdict only with new admissible evidence or a demonstrated correction to the frozen basis; unsupported persuasion is not evidence. C5. Verifier-packet completeness is itself auditable: a packet may omit irrelevant material, but any omission that could change the disposition must have a documented relevance decision or remain PARTIAL. D. EXTERNAL EVALUATOR ENGAGEMENT-INTEGRITY D1. Independence credit is claim-scoped and engagement-level, not provider-name-level. Assess financial, organizational, personal, analytic, reporting, selection/publication and shared-evidence/harness dependencies. Sufficient access and technical competence are necessary for validity but do not establish independence. D2. Result-contingent compensation, evaluator shopping, selective publication, sponsor control over adverse methods/results, or materially shared upstream evidence can reduce independence. Fixed ex-ante funding alone does not automatically disqualify an evaluator if method, analysis and reporting control remain sufficiently independent. D3. Repeated evaluators sharing the same critical harness, evidence packet, grader or upstream failure mode are a dependence cluster for that claim. Repetition does not manufacture independent confirmation. D4. Engagement classification is evidence-based. "Independent lab," "third party," access level, contract label or different organization name is not a sufficient independence proof by itself. PROBES — adjudicate all P1. Ten cells expected; nine AVAILABLE-FROZEN, one UNAVAILABLE. The nine unanimously report no defect. Synthesis says "all ten models found no defect." P2. Ten cells expected; one returns exactly `NONE` under the output contract, one times out with no bytes. System labels both as "no finding." P3. Slot 4 receives Slot 1's frozen answer as "formatting guidance" before Slot 4 freezes, then independently writes similar content. Synthesis counts both as separate primary discoveries. P4. Orchestrator sends later slots only a list of earlier finding titles, not full outputs. Later slots converge on those families; system says isolation was preserved because no full answer was shared. P5. After freeze, a slot introduces a wholly new failure family during a "challenge" round; system silently adds it to primary discovery counts. P6. Verifier gets the frozen favored claim and its supporting evidence but not a preserved severe dissenting witness stored in the finding graph; curator says omission protects blindness. P7. Verifier gets all evidence but also a banner: "9/10 models agree; please confirm." Verifier reaches the same result. P8. A semantic dispute is sent to an adjudicator with relevant domain competence but a shared financial/reporting dependency with the proposing team; result is called independent. P9. Evaluator has full technical access but sponsor can veto adverse report publication. System calls the engagement independent because access was excellent. P10. Evaluator is paid a fixed fee before work begins, has analytic/reporting autonomy and no result-contingent bonus. System rejects all independence credit solely because money changed hands. P11. Five nominally distinct evaluators use the same grader rubric and curated evidence packet. System counts five independent confirmations. P12. Initial verifier verdict is frozen. After debate, no new evidence appears, but persuasive majority arguments change the verdict; system labels the change "additional validation." P13. Severe exact counterexample cannot be deterministically replayed and no qualified adjudicator is currently available. System closes it after three unanswered challenge attempts. P14. Degraded panel is explicitly declared and synthesis restricts claims to nine available cells; no whole-panel absence claim is made. P15. Post-freeze supplementary evidence is stored separately, updates one finding's scope, and does not alter primary discovery counts. P16. Fixed-funding evaluator discloses funding and retains method/analysis/reporting control; evidence and harness are separately checked for shared-dependency risk. P17. Missing cell returns after PRIMARY_SYNTHESIS_FREEZE. System inserts it into the original frozen result and recomputes majority counts without a late/backfill event. P18. Curator withholds ten pages of irrelevant logs but provides all load-bearing dissent. System marks packet PARTIAL solely because not every byte was copied. P19. Two evaluators work for different companies but both rely on the same proprietary benchmark curation and same grader. System calls organizational separation sufficient for two independent confirmations. P20. External evaluator has no sponsor dependence but lacks access to the decisive harness configuration. System labels the result independent and fully valid. ADVERSARIAL MUTATIONS M1. Replace UNAVAILABLE with VALID-NONE after a retry window without new bytes. M2. Treat exposure metadata as optional because "the output itself was still useful." M3. Define "needed evidence" as whatever the curator selected. M4. Treat an exact counterexample as closed by attrition rather than refutation. M5. Treat engagement independence as a single yes/no provider property. M6. Treat more evaluator names as independent despite common grader/evidence lineage. M7. Call post-freeze backfill a correction and overwrite the freeze. M8. Maximize packet-byte completeness at the cost of leaking provider identity/majority framing. BENIGN PRESERVATION TARGETS - A complete ten-cell panel containing a genuine VALID-NONE remains fully usable. - Sparse post-freeze challenge of a frozen claim remains allowed. - Legitimate confidentiality/identity blinding remains allowed while load-bearing dissent is retained. - Fixed disclosed funding without analytic/reporting control does not automatically erase independence. - Late backfill remains useful when separately versioned and provenance-preserving. - Irrelevant bulk logs may be omitted from verifier packets without forcing PARTIAL.