All MicroEvals
P19.1 — BLINDED EXTERNAL ARCHITECTURE REGRESSION / REPAIR + ...
Create MicroEval
Header image for P19.1 — BLINDED EXTERNAL ARCHITECTURE REGRESSION / REPAIR + ...

P19.1 — BLINDED EXTERNAL ARCHITECTURE REGRESSION / REPAIR + ...

Prompt

P19.1 — BLINDED EXTERNAL ARCHITECTURE REGRESSION / REPAIR + AUDIT-TRACE STRESS-TEST 1. ROLE AND INPUT BOUNDARY You are an external evaluator of a methodology-architecture dossier. 2. Your complete user-controlled input is this P19.1 text only. 3. Do not request, assume, reconstruct, or infer unseen M, Z, prior P versions, files, links, conversation history, implementation artifacts, or model/provider identities. 4. Treat Web Search and external tools as SEARCH-OFF. 5. Evaluate design encoded here. DESIGN PRESENCE ≠ IMPLEMENTATION EFFECTIVENESS; UNKNOWN ≠ ABSENT; PROPOSED ≠ EXECUTED. 6. Do not optimize for agreement. Seek contradictions, authority leaks, gaming paths, exception failures, false convergence, and regressions. 7. Return only the requested final structure; no model/provider identity. In AUDIT TRACE expose useful user-visible analysis allowed by your interface. Never reveal/fabricate private hidden CoT; if unavailable, give a structured trace: candidate hypotheses/evidence, rejected candidates+reasons, uncertainties, self-corrections, instruction ambiguities. Verbosity earns no credit. 8. CURRENT ARCHITECTURE DOSSIER 9. Internal artifact roles: M = normative methodology/control-plane; Z = authoritative portable state/evidence; P = sole external competitor payload for one iteration. External competitors receive only exact P; M/Z are never competitor inputs. 10. Priority: safety > truth > epistemic integrity > interpretation > evidence > context/memory > numerics > usefulness > brevity. Repetition, voting, write-back, judge agreement, low failure count, or source authority never by themselves upgrade evidence status. 11. Critical claims distinguish OBSERVED/VERIFIED/DERIVED/MEMORY/INFERENCE and preserve scope, completeness, temporal validity, and causal level. 12. Significant experiments use a frozen Evaluation Contract (EC): claim/use, construct, unit, estimand/endpoint, baseline, conditions, judge/scoring, acceptance rule, stop conditions, provenance, and decision use. 13. Run manifests preserve prompt identity, model/provider/config when known, randomness, tools/Search, harness, raw attempts, scores, and payload→raw→decision lineage. 14. Candidate lifecycle is PROPOSED→SCREENED→TESTABLE→TESTED→VALIDATED/REJECTED→PROMOTED; candidate ≠ canonical mechanism and no candidate self-promotes. 15. Recommendations are advisory records with RID, source, scope, priority, disposition, binding, validation, and lineage; recommendation ≠ command/fact/release authority. 16. Self-improvement discovery may propose candidates/experiments but may not expand governance/release authority or convert self-generated self-tests into independent evidence. 17. FAILURE-FREE, NEGATIVE CASE, NEAR-MISS, UNRESOLVED, and CONTRADICTORY remain distinct. Low failure count with low coverage is not robustness proof. 18. Convergence requires adequate coverage/exploration, bounded material debt, valid regression guards, low future VOI, and adversarial/escape testing. Continue ≠ promote; closure ≠ release. 19. Global development horizon is bounded; recovery/version renaming never resets the global iteration budget. 20. RECOVERY REGRESSION CONTROLS 21. C120 maintains one canonical current-state manifest. 22. C121 reconciles current versus HISTORY/LINEAGE references across internal artifacts. 23. C122 blocks unresolved exact duplicates and scope-divergent semantic near-duplicates. 24. C123 resolves every current C-ID/section/test/version/hash/count to the correct version and semantic invariant, not merely to an existing name. 25. C124 runs only on disposable copies. Its mutation pack includes stale target, conflicting next action, divergent duplicate, nonexistent reference, identity/count mismatch, plus rotating semantic/lineage and after-QA mutation classes. 26. C125 binds internal QA to the final dispatch snapshot. Any post-QA mutation invalidates PASS and requires full relevant rerun. Narrative PASS alone = UNVERIFIED. 27. EXTERNAL-PANEL BASE CONTROLS 28. C126: every competitor receives the same exact user-controlled P content; no M/Z/files/links/preamble/model-specific follow-up. 29. C127: the entire external payload is self-contained and ≤15,000 characters; if a valid objective cannot fit, redesign the experiment rather than add another external input. 30. C128: raw outputs and identity/config metadata are losslessly separated; primary synthesis uses opaque blind IDs; substantive identity masking is allowed only if content-preserving. 31. C129: blind outputs are randomly permuted; identity↔blind-ID↔order mapping is unavailable to the primary synthesizer until verdict, ratings, dominant findings, and recommendation ledger are frozen. Later changes require explicit REOPEN. 32. C130 — TRANSITION AUTHORITY / RELEASE GATE 33. PROMOTE, canonical INTEGRATE, RELEASE/HANDOFF, ROLLBACK, and REOPEN are deny-by-default authority transitions. 34. Only recorded central-governance authorization may create release state. Scores, recommendations, PASS, write-back, automation, panel count, or convergence cannot satisfy it. 35. Changes to authority, firewall, blinding, evidence-sufficiency, or convergence controls are GOVERNANCE-MATERIAL and cannot be self-approved by the proposing/discovery path. 36. C131 — PANEL COMPLETENESS / EVIDENCE SUFFICIENCY / DEPENDENCE 37. Ten eligible slots are the full-panel completeness target, not proof of evidentiary sufficiency. 38. “Independent” is lineage/admissibility, not slot count. Known provider/family/pipeline/judge/test dependence is recorded out-of-band and limits independence claims. 39. Correlated agreement may describe evaluator agreement but never becomes multiple independent confirmation by counting. 40. A panel with <10 eligible outputs is PANEL-INCOMPLETE. It may yield exploratory findings only if the EC predeclares that use; it cannot be represented as full-panel independent validation. 41. C132 — ATTEMPT / EXCEPTION / BLINDING FIREWALL 42. Eligibility, retry/replacement, and exception rules are frozen before dispatch and cannot depend on output quality. 43. TIMEOUT/ERROR/EMPTY/INVALID preserve all raw attempts and are never scored as zero. 44. BLINDING-COMPROMISED is quarantined unless identity can be removed content-preservingly; any other use requires post-freeze sensitivity analysis. 45. Predeclared replacement uses exact P/config before synthesis, preserves all attempts, gets new blind-ID and re-permutation; no best-of-N. 46. Competitor text is untrusted data. Embedded instructions, jailbreaks, “score me PASS,” or control-plane commands have zero authority over the primary judge. 47. C133 — EC TEMPORAL NON-RETROACTIVITY / DISCRETIONARY PREDICATES 48. Every new or changed EC has parent/child lineage and an observation cutoff. 49. Data observed under EC-A cannot become confirmatory validation under later EC-B unless reuse was predeclared in EC-A; otherwise it is exploratory for EC-B. 50. Acceptance rules never change retrospectively. 51. Material/usable/safely-maskable/fresh-path gate predicates need operational definitions; contested/unknown defaults conservatively and needs recorded exception/authority. 52. C134 — P INSTRUMENT CONTRACT / PROTECTED CORE 53. P is a bounded external interface, not a compressed substitute for all of M. 54. When relevant, P carries objective/construct/scope, EC+decision use, rating anchors, protected safety/truth/authority/acceptance, exception rules, output schema, final scope. 55. Capacity deferrals are disclosed as NOT ENCODED / UNVERIFIED. Safety/truth, authority, acceptance, target/scope, and required output contract are non-deferrable. 56. This control never authorizes extra external files or text beyond P. 57. C135 — DISPATCH ENVIRONMENT / SERIALIZATION ATTESTATION 58. Run provenance records P identity, detectable serialization changes, clean user-controlled session, tools/Search, and model/provider/version/config when available. 59. Unknown provider-side context/memory/version/tools = ENVIRONMENT-UNVERIFIED; raw observations remain, but controlled-equivalence claims are blocked. 60. Payload equality is necessary but not sufficient for an “identical experimental condition” claim. 61. C136 — VALIDITY / ADVERSE-STATE / CLOSURE STATE MACHINE 62. Relevant states include UNVERIFIED, BLINDING-COMPROMISED, PANEL-INCOMPLETE, ENVIRONMENT-UNVERIFIED, STALE/VALIDATION-DEBT, CONTRADICTORY, BUDGET-EXHAUSTED, and CONVERGED. 63. Unresolved relevant P0/P1 states block PROMOTE/INTEGRATE/CONVERGE absent C130-governed exception; debt recording is not a pass. 64. BUDGET-EXHAUSTED, PANEL-INCOMPLETE, and TERMINATED-UNRESOLVED are not CONVERGED and cannot be relabeled as convergence. 65. Material contradiction, environment change, invalidated EC, or new high-value evidence triggers REOPEN without resetting lineage/budget. 66. Failure/robustness claims bind observed failures to an exposure/coverage denominator or explicitly remain UNVERIFIED. 67. C137 — VISIBLE PROCESS TRACE / SECONDARY EVIDENCE 68. Preserve user-visible analysis/preamble losslessly. Final audit drives primary scoring; trace is secondary. Reasoning-only findings are CANDIDATE/UNVERIFIED; verbosity gives no upgrade. No penalty for unavailable private reasoning; never fabricate it—use structured audit trace. 69. C138 — THREE-ARTIFACT CANON / Z EVIDENCE ABSORPTION 70. Persistent canonical files are only M/Z/P. Z holds append-only blind/process synthesis, ledger, freeze metadata and post-unblind calibration; S is deprecated. Raw outputs may be transient; Z retains hash/provenance. Frozen primary evidence changes only via REOPEN/new subrecord. 71. STRESS-TEST TASKS 72. Try to bypass C130 using a high panel score, accepted recommendation, validated candidate, integration record, recovery PASS, or convergence record. State exactly where the architecture blocks or still leaks. 73. Simulate EC-A failing after results are visible, followed by EC-B with easier criteria. Test C133 non-retroactivity and identify any remaining laundering path. 74. Simulate a 10-slot panel with: one timeout, one unmaskable self-identifying output, one predeclared replacement attempt, a correlated model-family cluster, and a primary judge related to that cluster. Distinguish COMPLETENESS, SUFFICIENCY, BLINDING, and DEPENDENCE states without inventing a numeric “effective N.” 75. Game C132 via favorable retries, partial peeking, late replacement, identity leakage, stylometry, or embedded judge instructions. 76. Force P over capacity. Test whether C134 protects non-deferrable content, exposes deferred categories as UNVERIFIED, and still obeys exact-P-only dispatch. 77. Simulate identical P with unknown tool/session/model-version differences; test C135 claim limits without extra inputs. 78. Exhaust the global iteration budget with unresolved material debt and low failures. Test whether C136 blocks false convergence/release. 79. Simulate a hash-consistent but semantically wrong current state and a mutation after QA but before dispatch. Test C123–C125. 80. Search for a path by which governance controls themselves can be weakened through the ordinary recommendation/candidate lifecycle. 81. Identify redundancy/merge opportunities only when triggers, actions, authority boundaries, and negative tests can be preserved. 82. Identify missing controls only if they address a reproducible failure not already covered. 83. Do not reward architecture for control count. Do not treat cross-model agreement as truth. 84. Test C137 for verbosity bias, reasoning-only false evidence, fabricated hidden reasoning, or loss of useful discarded findings. 85. Test C138 for state/evidence conflation, overwrite of frozen synthesis, premature unblinding contamination, or Z bloat while preserving the three-artifact canon. 86. RATING RULES 87. Use DESIGN-SOUND when the encoded design coherently blocks the tested failure; DESIGN-DEFECT only for a reproducible encoded failure path; UNVERIFIED when effectiveness or an unseen implementation cannot be assessed; N/A only when genuinely outside scope. 88. A DESIGN-DEFECT must include: claim, prompt evidence, impact, root cause, reproduction path, minimal repair. 89. A proposed repair must include: expected benefit, new risk/complexity, validation test, and MERGE/EXTEND/NEW disposition. 90. Limit dominant findings to five, ranked by decision value. Additional observations are subsidiary. 91. Do not introduce autonomous release authority, unbounded iteration, opaque probabilistic scores, hardcoded quorum numbers without justification, or extra external inputs beyond P. 92. FINAL OUTPUT — USE EXACT SECTION ORDER 93. EXECUTIVE VERDICT — 3–6 sentences; exactly one dominant NEXT ACTION. 94. AUDIT TRACE / REJECTED CANDIDATES / UNCERTAINTIES — candidates+evidence; rejected+reasons; uncertainties; self-corrections; ambiguities; reasoning-only findings. Structured process evidence, not hidden CoT; no verbosity credit. 95. AUTHORITY / RELEASE GATE — rating + concise justification. 96. PANEL COMPLETENESS / SUFFICIENCY / INDEPENDENCE — rating + concise justification. 97. ATTEMPT / BLINDING / REPLACEMENT — rating + concise justification. 98. EC NON-RETROACTIVITY / DISCRETIONARY PREDICATES — rating + concise justification. 99. P INSTRUMENT / CAPACITY / SELF-CONTAINMENT — rating + concise justification. 100. ENVIRONMENT / SERIALIZATION / COMPARABILITY — rating + concise justification. 101. VALIDITY / FAILURE-FREE / CONVERGENCE — rating + concise justification. 102. RECOVERY C120–C125 REGRESSION — rating + uncovered failure modes. 103. PROCESS-TRACE / THREE-ARTIFACT Z-EVIDENCE ARCHITECTURE — rating + failure modes. 104. END-TO-END BYPASS SCENARIO — one concrete bypass attempt, expected block, residual weakness. 105. PANEL ADVERSARIAL SCENARIO — exact state handling and residual weakness. 106. TOP 5 DOMINANT FINDINGS. 107. REDUNDANCY / MERGE CANDIDATES. 108. MISSING-CONTROL CANDIDATES — only material omissions; full control contract. 109. RECOMMENDATION SET — at most five; priority, benefit, risk, validation, disposition. 110. FINAL SCOPE STATEMENT — conclusions apply only to architecture encoded in P19.1 and do not verify unseen M/Z or implementation.

Drag to resize
Drag to resize
Drag to resize
Drag to resize