P37.2b PROMPT 1 - ELIGIBILITY STATE / ATOMIC EPOCH / COMPOSI...
Prompt
P37.2b PROMPT 1 - ELIGIBILITY STATE / ATOMIC EPOCH / COMPOSITIONAL ASSURANCE / TYPED EVALUATION DELIVERY CONTRACT You are one blinded external evaluator slot. Audit only this exact payload. SEARCH-OFF / P-ONLY / DESIGN-AUDIT. Do not use Web/Search, tools, prior chat, sibling P, M/Z, provider identity, hidden platform state, or outside facts. Do not infer missing facts. UNKNOWN != ABSENT; PARTIAL != COMPLETE. This is design review, not runtime proof. IMPORTANT UTILIZATION CONTRACT Do NOT print scratchpad, chain-of-thought, planning, self-talk, or a preamble. Begin exactly with "1. EXECUTIVE VERDICT". Use the payload as a five-lens conclusion audit. These are correlated passes inside one model, never five independent votes. Each lens may emit at most TWO candidate findings; emit NONE rather than filler. Later lenses must not merely restate an earlier candidate; if the same defect survives another lens, add that lens as provenance to the same canonical finding. L1 EXPLORE: seek a genuinely distinct failure mechanism or under-covered invariant. L2 ADVERSARY: construct smallest harmful paths still compliant via relabeling, timing, stale state, partial evidence, correlation, version rollover, waiver stacking, resource starvation, scope narrowing, adaptation or self-reference. L3 VERIFY: search the whole payload for closing clauses; reject false positives; test a benign near-neighbor and UNKNOWN/PARTIAL consequence. L4 SYSTEMS: test cross-clause interactions, concurrency/epochs, failure recovery, monitor-methodology coupling, recursion/adaptation and handoff. L5 SIMPLIFY/META: strongest disconfirmation; merge/retire redundant controls; detect metric/authority gaming; propose the smallest falsifiable repair. Do not count agreement, confidence, repeated wording or same-model lens recurrence as truth credit. An exact applicable counterexample may outweigh nine favorable slots. Confidence is diagnostic unless the payload positively qualifies calibration. MAX-UTILIZATION / ANTI-ANCHOR ADDENDUM Stress tests A-T are SEEDED PROBES, not a list of expected defects. After the seeded matrix, perform one WILDCARD search on the least-covered axis in your own audit among timing/concurrency, scope/transfer, authority, evidence independence, evaluator dependence, rollback/recovery, version/currentness, resource/starvation, adaptation/self-reference, and simplification/retirement. A wildcard candidate survives only if it is not a semantic duplicate of A-T or another retained finding. Tag retained findings SEEDED or UNSEEDED. Before retaining any candidate, perform the strongest-closing-clause search; if the payload closes it, REJECT rather than weaken the standard. For every recommended repair give both a KILL-TEST that the repair must block and a BENIGN-PRESERVATION test that it must still allow. Use NONE rather than quota filling. METHOD For each retained defect: cite exact clause; give the smallest text-compliant harmful path; search the whole payload for a closing clause; test one benign near-neighbor; state dependency-scoped consequence; classify REUSE / MERGE / EXTEND / NEW / REJECT / DEFER. Prefer REUSE -> MERGE -> EXTEND -> NEW. Do not invent thresholds, owners, independence, implementation, authority, completeness, environment state, clinical recommendations or current facts. A record/diagnostic alone is not a gate or authority. Blocking is dependency-scoped; unrelated safe/read-only/protective work stays available unless dependent. REQUIRED OUTPUT - EXACT ORDER 1. EXECUTIVE VERDICT - 3-6 sentences; exactly one DOMINANT NEXT ACTION. 2. FIVE-LENS FINDING LEDGER - L1..L5; max two concise candidate IDs per lens; tag SEEDED/UNSEEDED; cross-reference duplicates; NONE allowed. 3. AUDIT TRACE / REJECTED CANDIDATES / UNCERTAINTIES. 4. SECTION RATINGS - one sentence per target section: DESIGN-SOUND / DESIGN-DEFECT / UNVERIFIED / N/A. 5. STRESS-TEST MATRIX - rows A-T: BLOCKED / BYPASS / UNRESOLVED / ALLOWED-BENIGN; exact clauses + dependency-scoped consequence. 6. TOP DOMINANT FINDINGS - max 6; fingerprint = violated invariant + trigger + harmful consequence + scope + closing-clause status + minimal repair; include lens IDs and SEEDED/UNSEEDED. 7. WILDCARD + CROSS-INTERACTION / ADAPTATION / RECURSION CHECKS - report the distinct surviving wildcard or NONE; include correlation/stale-version tests where relevant. 8. REDUNDANCY / MERGE / RETIRE + CONCISE RESIDUAL REGISTER. 9. RECOMMENDATION SET - max 6; disposition + minimal patch + KILL-TEST + BENIGN-PRESERVATION test; no generic advice. 10. FINAL SCOPE STATEMENT - exact-payload design scope; no implementation, medical, legal, deployment or release authority. TARGET DESIGN A. EVIDENCE-TO-ELIGIBILITY STATE A1. A declaration, diagnostic, score, record or audit note is never permission by itself. Every load-bearing claim/action/dependency has a typed state: VERIFIED/ELIGIBLE, UNVERIFIED, UNKNOWN, PARTIAL, FAILED/CONTRADICTED, EXPIRED/STALE, HOLD, WAIVED/ACCEPTED-RESIDUAL, RETIRED/NOT-APPLICABLE. UNKNOWN != ABSENT; PARTIAL != COMPLETE. A2. Gate record binds subject/version/scope, evidence identity/currentness, criterion, state, affected dependents, consequence for non-verified states, owner/adjudicator, revisit/expiry, authority boundary and triggering witness. Non-verified state blocks/narrows only dependent claims; unrelated safe/read-only/protective work stays available. A3. Work-reducing labels are eligibility calls. BLOCKED needs a named externally checkable dependency or positively evidenced resource/authority barrier. DEFERRED needs owner/revisit/consequence. INCOMPATIBLE needs scope/interface reason. Valid stop: hard resource/tool/context exhaustion, named external dependency or recorded measured low marginal value. Convenience/unlabeled BLOCKED cannot manufacture saturation. B. ATOMIC EPOCH / EXACT-P FREEZE / HANDOFF B1. Every state-mutating run reads STATE-EPOCH and commits only if the base epoch is current or an explicit non-conflicting merge/rebase succeeds. Stale writes cannot overwrite newer M/Z, findings/debts, monitor state or result intake. Obligations and exact counterexample witnesses are monotone until authorized disposition. B2. A P part is FROZEN only if final serialization/identity is valid, char and UTF-8 byte limits pass, no-filler/usefulness passes, independent/direct second byte verification matches the committed bytes, SHA-256 reproduces and manifest binds the bytes. Mismatch -> FAILED-FREEZE/NOT-FROZEN; dispatch barred; divergent witness retained. Dispatch transition occurs only when every expected part and manifest passes one atomic predicate. B3. LOSSLESS handoff = semantic manifest complete AND bytes valid. Required state classes include current M/Z/P identities, P lifecycle/result status, expected/received result set, open findings/errors/debts/waivers, monitor config/runtime-verification, mission/improvement summary, protected baselines, scheduler/receipt limitations, PROJECT_NEXT_ACTION, CONTINUATION_VOI and restore instructions; explicit empty class is allowed. Receiver reconstruction/currentness/schema check is part of verification. C. COMPOSITIONAL ASSURANCE / WAIVER / PROOF C1. Required proof/discharge graph must be well-founded. An obligation cannot be DISCHARGED solely by a circular assume-guarantee cycle; cycles require an independently grounded root/environment fact or remain UNVERIFIED-PENDING/HOLD for dependents. C2. Waivers/residuals/exceptions are evaluated compositionally along each dependency path. Individually valid waivers cannot collectively remove all defense-in-depth without an independently governed compositional decision. Scope, version, cumulative age, renewal lineage, compensating controls and breach consequence are explicit. C3. A compensating control counts only when its relevant implementation is positively evidenced current and armed/healthy. Stale support or changed dependency/evaluator/version invalidates dependent discharge until requalified. Fallback/rollback requires compatible state/schema migration and armed trigger, not a tombstone or stale historical success. D. TYPED EVALUATOR / CALIBRATION / NO-GOLD D1. TYPE-SAFE/SCHEMA-VALID is representation evidence, not truth or decision validity. Typed outputs remain inside evaluator-bias/adversarial/currentness gates. A universal-superiority claim needs qualified alternatives under matched relevant system envelope. D2. Calibration is claim-, subgroup-, version- and operating-envelope specific. UNKNOWN/expired calibration does not authorize threshold-as-risk/truth. Severe exact subgroup failures cannot be hidden by aggregate calibration. Router cost claims include fallback/review/retry/capacity paths, not retained-success cases only. D3. AGREEMENT/CONSENSUS != ACCURACY unless anchored by qualified gold/partial-gold or explicit justified assumptions appropriate to the claim. Shared rubric/source/model/harness dependencies are correlated. A no-gold reference/mean of strong models is assumption-bound. Exact high-severity counterevidence remains visible; confidence is not weight unless qualified. E. SELECTION / CURRENTNESS / ADAPTATION E1. Evidence used to generate, tune, select, route or stop cannot silently become independent confirmation for the same promoted claim. Threshold/rubric/alternative selection lineage is recorded; post-selection evidence has the scope appropriate to its exposure. E2. Model/version/alias identity is load-bearing where evidence currentness depends on it. A moving alias/version change triggers scoped requalification; old evidence cannot silently attach to a new system state. E3. Known failure-family representability remains explicit. High mutation kill rate cannot hide a known load-bearing family that the operator set cannot express; disposition is named test/operator, verified compensating control, bounded accepted residual/debt or HOLD. F. AUTHORITY / BASELINES / RECOVERY F1. Protected acceptance criteria/evaluators/authority baselines have a versioned v0 before first load-bearing use. Author/control lineage is recorded. Beneficiary-authored/controlled v0 is proposal-only until the required independent adjudication; UNKNOWN lineage preserves protection. Protection-reducing revisions also require independent authority. F2. Retry/fallback/rebase/rollback cannot erase the triggering witness. A stale runner or failed repair cannot close its own debt by overwriting state. Requalification explicitly binds which previous obligations are discharged, inherited or reopened. F3. Blocking and HOLD are dependency-scoped. High-impact uncertainty may hold the affected action, but a record of uncertainty cannot globally stop unrelated safe/read-only/protective paths or invent medical/legal/deployment authority. STRESS TESTS A-T A. Saturation record labels a ready high-risk branch BLOCKED with no named dependency and stops. B. A named external dependency blocks one lane; all unrelated ready read-only lanes are halted too. C. Two pracuj runs overlap; older runner commits after newer M/Z and erases one new open debt. D. Result intake under a stale epoch overwrites newer received-set status. E. P part first count is 14920 bytes; independent second verification differs; system treats recount as advisory and dispatches. F. All P hashes validate, but handoff omits RESULTS-RECEIVED and open-waiver state; receiver calls it lossless. G. A and B proof obligations mutually discharge each other with no independently grounded root. H. Two individually valid waivers remove primary and secondary safeguards on the same dependency path; no compositional check fires. I. Waiver cites compensating control whose documented existence is current but live arming/health is UNKNOWN. J. Old fallback passed six months ago; schema changed; current refactor deletes new fallback because tombstone exists. K. Typed JSON validates perfectly; system concludes answer is correct and safe. L. Aggregate evaluator calibration is good, but severe subgroup is poor; subgroup witness averaged away. M. Router looks 98% accurate only because 70% hard cases are handed off; system claims cheap router solved the task. N. Ten model judges agree under a shared rubric/source pool; system calls agreement identified truth. O. A strong-model mean with no gold is promoted as gold standard because models are advanced. P. Threshold is chosen on outcomes, then the same cases validate that threshold without selection qualification. Q. Model alias moves to a new version; old calibration and bias probes remain VERIFIED automatically. R. Known catastrophic failure family has no mutation operator; suite kills 100% of generated mutants and claims complete coverage. S. Beneficiary authors weak initial v0 acceptance baseline and it becomes load-bearing without independent adjudication. T. Repair retry succeeds only after discarding the original failing witness; system closes the defect as fixed. SPECIAL PROBES 1. Create partial/staggered external result intake during a concurrent result-return run; test epoch/currentness and completeness states. 2. Create a semantically complete but byte-corrupt handoff and a byte-valid but semantically incomplete handoff; both must be distinguished. 3. Create circular proof discharge with three nodes and one stale version edge. 4. Stack three bounded waivers that are individually valid but collectively eliminate a protected invariant. 5. Create evaluator-version rollover where only style changes are claimed; decide what evidence is sufficient to preserve versus reopen validation. 6. Create k/part-count adaptation oscillation driven by marginal fill change; check whether evidence-backed hysteresis belongs in freeze/issue governance without freezing real adaptation. 7. Create a benign named-tool-limit stop and benign low-marginal-value stop; repairs must preserve both. 8. Test command-surface semantics: typed/clicked/copied execution must enter the same governed state; presentation alone is not execution authority. 9. WILDCARD METAMORPHIC: mutate one load-bearing dimension (UNKNOWN/ABSENT, epoch n/n+1, scope narrow/wide, active/stale compensating control, or waiver order) while holding irrelevant text constant; report only a distinct surviving defect.
Response not available