
P 18.2 — BLINDED EXTERNAL ARCHITECTURE STRESS-TEST STATUS: C...
Prompt
P 18.2 — BLINDED EXTERNAL ARCHITECTURE STRESS-TEST STATUS: CANDIDATE EXTERNAL PANEL PAYLOAD — ITERATION 18 SCOPE: Evaluate only the methodology architecture dossier encoded in this prompt. No other artifact is part of your input. MANDATORY INPUT BOUNDARY 1. Your complete user-controlled input is this P18.2 text only. Do not request, assume, reconstruct, or claim to inspect M, Z, prior P versions, files, links, conversation history, hidden context, or external attachments. 2. This is a design-level stress-test, not file-level verification of unseen implementation. Exact implementation is UNVERIFIED unless encoded below. 3. Do not infer missing details from names, numbering, prior versions, or likely project conventions. UNKNOWN ≠ ABSENT; PROPOSED ≠ EXECUTED; DESIGN PRESENCE ≠ EFFECTIVENESS. 4. Do not identify or mention your model/provider identity in the answer. Start directly with the requested output structure. 5. Treat Web Search and external tools as SEARCH-OFF for this task. Do not simulate search results or cite sources you did not actually access. 6. Do not optimize for agreement. Search for contradictions, authority leaks, blind spots, redundancy, gaming paths, and false convergence. ARCHITECTURE DOSSIER — CURRENT DESIGN 7. The system has three internal artifact roles: M = normative methodology/control-plane; Z = authoritative portable state/evidence; P = the single external competitor payload for one iteration. Conversation history is not required as source-of-truth. 8. External dispatch is payload-isolated: each competitor receives exactly current P and no other user-controlled text/file. M, Z, recovery/QA material, preambles, attachments, and follow-ups are forbidden. 9. P is self-contained; complete user-controlled external payload ≤15,000 UTF-8 characters, working target ≤14,200. It may summarize internal architecture but cannot depend on unseen M/Z. 10. Internal continuity is separate from competitor dispatch: internal work may use M+Z+P, but that bundle is never the external competitor payload. 11. Priority is safety > truth > epistemic integrity > interpretation > evidence > context/memory > numerics > usefulness > brevity. Repetition, voting, or write-back never upgrades evidence status. 12. Claim discipline distinguishes observed/verified/derived/memory/inference states and forbids association→causality, mechanism→outcome, missing evidence→absence, or authority→support shortcuts. 13. Significant experiments use a frozen EVALUATION CONTRACT: intended claim/use, construct, unit, estimand/endpoint, baseline/reference, conditions, judge/scoring, acceptance rule, and stop conditions. Acceptance criteria do not silently change after seeing results. 14. Run manifests preserve model/provider, configuration, prompt identity, randomness, tools/Search, judge/harness, and source→payload→raw→score→decision provenance. 15. Fixed panel = 10 stable slots. Track execution/response states separately. Full-panel adjudication needs 10 usable independent outputs unless the evaluation contract says otherwise. 16. Preserve raw outputs losslessly with identity map. Before primary synthesis/scoring, sequester model/provider labels and assign opaque blind IDs. 17. Randomly permute blinded outputs. Keep identity↔blind-ID↔order map separate from the primary judge until verdict, ratings, dominant findings, and recommendation ledger are frozen. 18. Remove provider/model metadata from blind copies; preserve raw originals. If substantive text self-identifies and cannot be safely masked, record BLINDING-COMPROMISED. 19. Unblind only after primary synthesis is frozen. Later model-pattern analysis cannot alter it without explicit reopen record. 20. Artifact/dependency logic separates COMPETITOR-TEST-PAYLOAD, CENTRAL-AUDIT-BUNDLE, and RELEASE/HANDOFF-BUNDLE. A dependency is blocking only for the experiment that actually requires it. 21. Closed-loop recommendations get RID, source, scope, priority, disposition and, if accepted, implementation/test binding + validation. They are inputs, not commands/facts/release authority. 22. Mechanism governance uses canonical/candidate registries, fingerprints, relation decisions (SAME/EXTENSION/DISTINCT/SPLIT/MERGE/SUPERSEDE/NEW) and lineage; renaming alone creates nothing. 23. Candidate lifecycle: PROPOSED→SCREENED→REFINED→TESTING→DECISION→PROMOTED/DEFERRED/REJECTED/MERGED/SUPERSEDED. No self-promotion; transitions need evidence and decision record. 24. Implementation-depth control requires more than a label: operative rule, trigger/input, action/output, scope/invariant, authority/interaction boundary, and validation signature. 25. Flow: EVIDENCE→CLASSIFY→FAILURE/COVERAGE/UNCERTAINTY→CANDIDATE→MERGE→VOI→EXPERIMENT→VALIDATION→ADJUDICATION→INTEGRATION/ROLLBACK→MONITOR/REVALIDATE→CONVERGE/REOPEN. 26. OBSERVATION/CANDIDATE/EXPERIMENT/VALIDATION/CHANGE are data-contract interfaces sharing IDs, provenance, scope, dependencies, evidence state, lifecycle, lineage. 27. Architecture-health controls track missing layers/interfaces, authority bypass, orphan lineage, stale/deferred debt, diversity collapse, mechanism bloat, and evidence→decision traceability. Health is not inferred from control count. 28. Dynamic experiment discovery creates/changes EC-ID for a material decision gap; reword-only classes merge/reject unless objective/discriminator/validation path materially differs. 29. Experiment portfolio balances additive, prune/merge, boundary, adversarial, evaluator-integrity, and regression candidates. Selection favors net decision value rather than score/count growth. 30. VOI routing may choose experiment, fresh holdout, adversarial challenge, ablation, research, or human escalation. Missing data = UNVERIFIED, never zero. 31. Capacity governance scans every P; allocation records core/selected/deferred content, utilization/slack/reserve/rationale. Duplicate/cosmetic text has zero utilization benefit. 32. FAILURE-FREE is only a data subsystem: negative cases, near-misses, unresolved cases, and contradictions stay distinct. Low failure count is not robustness proof and never auto-promotes. 33. External standards/recommendation radar records source/date/scope/authority/supersession/conflict/recheck, but external authority never substitutes for claim support. 34. Self-Improvement Firewall separates discovery autonomy from governance/release autonomy. Material improvements require risk-matched challenge or fresh-path evidence; self-generated self-tests do not become independent validation. 35. Drift/change-impact tracks direct/indirect/interaction effects and validation debt. Timing ≠ causality; cross-branch transfer preserves lineage/revalidation. 36. Convergence uses marginal-value decline, coverage/exploration confidence, and escape/adversarial tests. Low failures + low coverage/exploration is false convergence. Continue ≠ promote; closure ≠ release. 37. C120–C125 enforce canonical state, current/history reconciliation, duplicate/divergence + reference checks, five pre-dispatch mutations, and exact-bound QA. Narrative QA PASS alone = UNVERIFIED. 38. C126–C129 enforce exact-P-only payload, P self-containment, blinded/sequestered intake, random permutation, and post-freeze unblinding. 39. Recovery/version renaming does not reset the global iteration budget. Internal repair is not independent external evidence. STRESS-TEST TASKS 40. Audit whether the architecture above has a coherent authority model. Identify any path where discovery, recommendation, scoring, self-test, aggregation, or convergence could acquire release authority indirectly. 41. Audit the evidence semantics. Find any place where repetition, consensus, judge agreement, internal QA, external authority, or absence of failures could be mistaken for stronger evidence than warranted. 42. Audit whether exact-P-only external dispatch is sufficiently separated from internal M/Z/P continuity to prevent hidden context, unequal inputs, or dependency confusion. 43. Audit self-containment under 15,000 chars: identify essential omitted information and low-value capacity use. 44. Audit blindness against prestige/order bias, self-identification, metadata leakage, replacement runs, invalid/timeout outputs, premature unblinding. 45. Propose permutation/unblinding that preserves provenance while hiding identity and original order before freeze. 46. Audit whether 10 fixed slots plus full-panel sufficiency can create liveness or selection problems. Distinguish panel completeness from evidentiary sufficiency. 47. Audit recommendation closure for silent loss, accepted-but-unbound items, recommendation-as-command, and recommendation-driven authority expansion. 48. Audit registry/lifecycle controls for mechanism proliferation, merge errors, duplicate names, premature IMPLEMENTED status, and untestable implementation signatures. 49. Audit experiment discovery for candidate flood, reward hacking, diversity collapse, confirmation bias, low-VOI churn, and proxy optimization. 50. Audit VOI/capacity for pseudo-quantification, missing-data abuse, padding, unsafe compression, and utilization-over-information incentives. 51. Audit FAILURE-FREE/near-miss logic for false robustness, denominator neglect, weak exploration, hidden negatives, and classifier circularity. 52. Audit self-improvement/evaluator integrity for self-confirmation, judge contamination, shared-family dependence, drift, and circular criteria. 53. Audit drift/revalidation for stale evidence, interactions, negative transfer, environment change, and scope/validity overclaim. 54. Audit convergence for premature closure, low exploration, suppressed contradictory evidence, budget exhaustion masquerading as convergence, and reopening failures. 55. Audit C120–C125 conceptually: identify recovery defects not covered by stale target, conflicting next action, divergent duplicate, nonexistent reference, or hash/count mismatch. 56. Audit C126–C129 conceptually: identify how exact-P-only dispatch, self-containment, anonymization, permutation, and delayed unblinding could still fail or be gamed. 57. Run one end-to-end adversarial scenario through observation→candidate→experiment→validation→integration→monitoring/convergence; show block or bypass. 58. Run a panel scenario with high-status model, weak model, timeout, self-identifying output, replacement run; determine blind/permuted handling without fabricating evidence. 59. Distinguish DESIGN DEFECT from IMPLEMENTATION-UNVERIFIED. Do not fail the unseen implementation merely because this prompt omits implementation evidence. 60. Identify redundant controls; merges must preserve trigger, action, scope, authority, evidence contract, and negative tests. 61. Identify missing controls only when the omission is decision-relevant and not already covered by an existing mechanism family. 62. New controls require prevented failure, trigger/input, action/output, scope/invariant, authority boundary, validation signature. 63. Prefer architecture-level repair for recurring failure families. Do not recommend string patches when the failure is systemic. 64. Do not propose extra external inputs beyond P. Any recommendation that requires giving competitors M, Z, files, links, hidden context, or follow-up clarification violates the payload constraint. 65. Keep the design bounded: do not add autonomous release authority, unbounded iteration, or opaque probabilistic scores unless necessary and governable. RATING RULES 66. Use DESIGN-SOUND / DESIGN-DEFECT / UNVERIFIED / N/A for domain ratings. Reserve BLOCKED only for a task that cannot be assessed from the architecture dossier actually present here. 67. DESIGN-DEFECT requires claim, prompt evidence, impact, root cause, reproduction path, minimal architecture repair. 68. Repairs require benefit, new risk/complexity, validation test, and merge/new-family disposition. 69. Agreement with this architecture is not evidence. A clean verdict requires explicit adversarial coverage, not merely absence of obvious contradictions. 70. Limit dominant findings to five. Additional observations may be listed as subsidiary items but must not displace higher-risk findings. FINAL OUTPUT — USE EXACT SECTION ORDER 71. EXECUTIVE VERDICT — 3–6 sentences; exactly one dominant NEXT ACTION. 72. AUTHORITY / EVIDENCE SEMANTICS — rating + concise justification. 73. EXTERNAL PAYLOAD ISOLATION / SELF-CONTAINMENT — rating + concise justification. 74. BLINDING / PERMUTATION / UNBLINDING — rating + concise justification. 75. ARTIFACT / DATA-CONTRACT / TRACEABILITY ARCHITECTURE — rating + concise justification. 76. RECOMMENDATION / REGISTRY / LIFECYCLE — rating + concise justification. 77. EXPERIMENT DISCOVERY / VOI / CAPACITY — rating + concise justification. 78. FAILURE-FREE / SELF-IMPROVEMENT / EVALUATOR INTEGRITY — rating + concise justification. 79. DRIFT / REVALIDATION / CONVERGENCE — rating + concise justification. 80. RECOVERY CONTROLS C120–C125 — rating + uncovered recovery failure modes. 81. EXTERNAL-PANEL CONTROLS C126–C129 — rating + uncovered panel failure modes. 82. END-TO-END ADVERSARIAL SCENARIO — bypass path, expected block, residual weakness. 83. PANEL ADVERSARIAL SCENARIO — exact blind/permutation handling and residual weakness. 84. TOP 5 DOMINANT FINDINGS — numbered, reproducible, highest decision value first. 85. REDUNDANCY / MERGE CANDIDATES — only material merges, with preserved invariants. 86. MISSING-CONTROL CANDIDATES — only material omissions; include full control contract. 87. RECOMMENDATION SET — at most five candidate recommendations, each with priority, benefit, risk, validation, and merge/new-control disposition. 88. FINAL SCOPE STATEMENT — explicitly state that your conclusions apply only to the architecture dossier encoded in P18.2 and do not verify unseen M/Z implementation.