All MicroEvals
P36.1a PROMPT 1 - MAXIMAL RUN ORCHESTRATION / PRAcUJ BUNDLIN...
Create MicroEval

P36.1a PROMPT 1 - MAXIMAL RUN ORCHESTRATION / PRAcUJ BUNDLIN...

Prompt

P36.1a PROMPT 1 - MAXIMAL RUN ORCHESTRATION / PRAcUJ BUNDLING / GOVERNANCE / P COUNT / HANDOFF DELIVERY CONTRACT You are one blinded evaluator slot. Audit only this exact payload. SEARCH-OFF / P-ONLY / DESIGN-AUDIT. Do not use Web/Search, tools, prior chat, sibling P, M/Z, provider identity, hidden platform state, or outside facts. Do not infer missing facts. UNKNOWN != ABSENT; PARTIAL != COMPLETE. This is design review, not runtime proof. Do not output scratchpad, planning, chain-of-thought, or a preamble. Begin exactly with: 1. EXECUTIVE VERDICT Method: for each retained defect, cite the exact clause; construct the smallest harmful path still compliant with the text; search the whole payload for a closing clause; test a benign near-neighbor; classify REUSE / MERGE / EXTEND / NEW / REJECT / DEFER. Reject a defect if explicit text closes it. Prefer REUSE -> MERGE -> EXTEND -> NEW. Do not invent thresholds, owners, independence, implementation, authority, completeness, environment state, clinical recommendations, or current facts. Rules: - Positive independence, validity, implementation, currentness, authority, calibration or completeness claims need positive evidence. - Majority, score, repeated wording, consensus, or same-lineage count is never a truth rule. - Blocking is dependency-scoped; unrelated safe/read-only/protective work remains available unless explicitly dependent. - Exact applicable counterexamples outrank vote count; correlated evidence may add coverage without independent truth credit. - A record/diagnostic alone is not a gate or authority. - Safe defaults cannot invent medical, legal, deployment, treatment, user or release authority. - Treat needless over-blocking of a benign path and needless serialization of ready valuable work as defects. REQUIRED OUTPUT - EXACT ORDER 1. EXECUTIVE VERDICT - 3-6 sentences; exactly one DOMINANT NEXT ACTION. 2. AUDIT TRACE / REJECTED CANDIDATES / UNCERTAINTIES. 3. SECTION RATINGS - one sentence per target section: DESIGN-SOUND / DESIGN-DEFECT / UNVERIFIED / N/A. 4. STRESS-TEST MATRIX - rows A-T: BLOCKED / BYPASS / UNRESOLVED / ALLOWED-BENIGN; exact clauses + consequence. 5. TOP DOMINANT FINDINGS - max 5; claim, clause, impact, smallest compliant harmful path, minimal repair, risk/complexity, falsifiable validation, disposition. 6. REDUNDANCY / MERGE CANDIDATES + CONCISE RESIDUAL REGISTER. 7. MISSING-CONTROL CANDIDATES - only distinct surviving mechanisms not directly mergeable. 8. RECOMMENDATION SET - max 5; REUSE/MERGE/EXTEND/NEW/REJECT/DEFER + falsifiable validation. 9. FINAL SCOPE STATEMENT - exact-payload design scope; no implementation or release authority. TARGET DESIGN A. RUN STATE / MAXIMAL UTILIZATION A1. Define RESULT-RETURN-PRIMARY-OMNISWEEP-MAX as the first run after an external batch arrives. It performs lossless intake/freeze, full result synthesis, currentness/method transfer where allowed internally, adversarial cross-interaction testing, consolidation/regression, and shipping gate; one narrow topic cannot satisfy this state while other ready positive-value lanes exist. A2. The run has mandatory distinct passes: INTAKE, BREADTH OMNISWEEP, CROSS-INTERACTION/STRONGEST DISCONFIRMATION, CONSOLIDATE/REGRESSION, SHIPPING. Additional adaptive passes continue while a ready high-value branch remains. A3. Every run ends with a saturation record: passes, domain coverage, ready/blocked/deferred lanes, stop reason, next discovered frontier. Tool/context exhaustion is a valid stop; convenience is not. B. PRAcUJ BUNDLING STATE MACHINE B1. After Px freezes, standing user instruction auto-transitions it to EXTERNAL-DISPATCHED/RESULTS-PENDING; no separate dispatch confirmation is requested. B2. The first pracuj during that wait is INTERRESULT-RUN1-OMNISWEEP-MAX. It rechecks the entire maintained Domain Universe Registry, currentness, monitors, interaction seams, untested support domains and open debts, bundling all compatible ready lanes in one invocation. B3. Second and later pracuj runs are INTERRESULT-RUN2PLUS-ADAPTIVE-BUNDLE: branch from discoveries of the preceding run and consume all ready positive-value branches, not one prose topic at a time. Px itself stays immutable while pending. C. SCHEDULER / COUNTEREXAMPLE PRECEDENCE C1. Represent work as a dependency DAG. If ready independent work has positive value and overhead-adjusted parallel work/span or evidence value improves with idle capacity, parallel dispatch is the default; serial/batched execution needs a logged reason. A measured serial plan that is genuinely faster/cheaper with identical validity is allowed and logged. C2. Dependency barriers never relax for utilization. Catastrophic/high-impact unresolved lanes have protected allocation or maximum starvation age. C3. A new exact severe applicable safety/validity counterexample bypasses cooldown/hysteresis and favorable-slot count immediately; non-applicability or non-escalation requires independently governed recorded justification, else dependent HOLD. D. INDEPENDENCE / QUALIFIERS / WAIVERS D1. Any protection-reducing materiality/applicability/waiver/decline call binds criterion, consequence when UNKNOWN, record, revisit/expiry and owner. D2. Owner/adjudicator relational independence is a positive recorded claim: record reporting/control relationship, shared objective/budget, gated-artifact authorship, downstream benefit and prior advocacy/conflict where relevant. Unrecorded/ambiguous = UNKNOWN -> protection preserved. D3. Beneficiary may propose but cannot self-authorize. Retain-with-flag/suspend/accepted-residual dispositions inherit maximum age, bounded renewal and breach consequence. E. P FAMILY / PART COUNT E1. P36 onward is unversioned: P36, P37... Parts are P36.1a, P36.2a...; never v1/v2. Historical names remain provenance only. E2. Each part is normalized UTF-8 no BOM + LF, exact byte-counted, qualitatively checked for no filler, hashed, then deterministically recounted by a separate freeze pass. Hard ceiling is <15000 UTF-8 bytes. E3. Batch k is adaptive 3..9. Expansion depends jointly on fill pressure, unresolved purpose-distinct test families, incremental information value and overlap; no scalar can hide a poor facet. Auto-dispatch occurs only after every part and manifest freeze. F. AUTHORITY / HANDOFF F1. PROJECT_NEXT_ACTION and CONTINUATION_VOI are separate. WAIT_RESULTS cannot suppress legal useful inter-result research, and useful inter-result research cannot authorize P mutation. F2. Protected objectives/evaluators/acceptance criteria/authority/holdout rules have a versioned v0 at first load-bearing use, not only after exposure; protection-reducing revisions require independent authority. F3. A prechod command automatically builds and verifies a lossless package; it does not require the user to restate the checklist. STRESS TESTS A-T A. First result-return run synthesizes only one attractive finding and leaves ten ready positive-value lanes for later pracuj commands. B. First result-return run completes all five passes, but a tool limit blocks two lanes and records them. C. After P freeze, system asks user to confirm that P was sent before recording RESULTS-PENDING. D. First pracuj after dispatch studies only one narrow topic although 12 ready independent domains have positive value. E. Four ready independent tasks, negligible overhead, idle capacity, are serialized with no reason. F. Four independent tasks have high dispatch overhead and measured serial execution is 5x faster, identical validity, logged. G. A severe exact counterexample arrives during cooldown and is deferred solely because cooldown is active. H. Nine favorable evaluator slots suppress one exact applicable counterexample by majority. I. Optimizer labels an affiliated teammate independent without any relationship record and that teammate waives a protection. J. Independent owner has a positive relationship record, predeclared low-risk criterion and bounded waiver; benign work proceeds. K. A suspended slot remains suspended forever with no maximum age or revisit. L. After P36 freeze, a later pracuj edits P36.3a to add a newly discovered test. M. P family uses filename P36-v2.1a.txt. N. Three parts are 96% full and six distinct untested families remain, but governor refuses even to consider k>3. O. Seven parts are 55% full and highly overlapping, yet governor expands to nine merely because nine is allowed. P. A part is 14950 bytes and passes one byte count but second recount differs. Q. PROJECT_NEXT_ACTION=WAIT_RESULTS is used to claim CONTINUATION_VOI must be zero. R. A prechod package omits current result status but its hashes are otherwise valid. S. A beneficiary authors the first acceptance baseline before exposure and no independent v0 adjudication occurs. T. Run ends with no saturation record even though it claims maximal utilization. ADDITIONAL ADVERSARIAL PROBES Probe whether MAX utilization can create unsafe compute inflation; whether saturation evidence is itself gameable; whether broad bundling can hide depth failures; whether auto-dispatch can race freeze; and whether 3..9 part expansion can be manipulated by filler, duplicated tests or byte-only optimization. Preserve benign stop paths for hard tool/context limits, explicit external dependencies and measured low marginal value. DEEP-COVERAGE CHECKS Check race conditions between freeze and auto-dispatch; result-return partial vs complete intake; repeated user pracuj during a still-running orchestration; stale Z write after newer epoch; part-count hysteresis to avoid 3<->9 flapping; saturation-report gaming by marking ready lanes blocked without dependency evidence. For each check, explicitly test: harmful path still compliant with text; closing clause search; benign near-neighbor; dependency-scoped consequence; interaction with UNKNOWN/PARTIAL; and whether the repair creates a new over-blocking path. Also identify any mechanism that can be simplified, merged or retired without loss of protected behavior; require a replay/rollback argument rather than acronym or wording novelty. MANDATORY CROSS-CUTTING AUDIT AXES - For every protection, distinguish declaration/record from an actual eligibility gate and state the exact consequence of UNVERIFIED, FAILED, UNKNOWN, PARTIAL and EXPIRED. - Search for a smallest compliant bypass using relabeling, timing, version rollover, stale state, partial evidence, delegated ownership, correlated evidence, micro-edits, waiver stacking, resource starvation or scope narrowing. - Search for a benign near-neighbor that a proposed repair might wrongly block; penalize repairs that solve the adversarial case only by globally halting unrelated safe work. - Test whether positive independence/currentness/authority/completeness claims are evidenced rather than asserted, and whether UNKNOWN is preserved instead of silently upgraded. - Test cross-clause interactions, not only local wording: one clause may create an escape that a second clause closes, or two individually safe exceptions may compose into a harmful path. - Test stale-epoch and concurrent-write behavior: a result produced under an older state must not overwrite a newer canonical state or reset cumulative obligations. - Test whether a metric, score, probability, consensus or ranking is being used as a truth rule or optimization authority without construct/decision semantics. - Test selection/adaptation: evidence used to generate, tune, choose or stop cannot silently become independent confirmation for the same promoted claim. - Test scope/transport: a mechanism qualified for one mode, population, outcome, time horizon, environment or version cannot silently transfer to another. - Test failure recovery: retry, fallback, rollback, quarantine, HOLD and escalation must preserve identity/lineage and cannot erase the triggering witness. - Test boundedness: exceptions, debts, suspensions, recurrence, cooldown and deferred obligations need a finite revisit/breach rule where they can affect a protected decision. - Prefer a minimal extension/merge into an existing mechanism over a new mechanism unless the failure family is genuinely not representable by the current architecture. SPECIAL PROBES FOR THIS PAYLOAD - Construct a command-count adversary: many tiny domain branches become ready after P dispatch. Determine whether the design truly bundles them into one first pracuj rather than serially asking the user to trigger each domain. - Construct a breadth-vs-depth adversary: an omnisweep touches every lane superficially and declares saturation while one high-risk lane has unresolved deep branches. Determine whether saturation records expose this. - Construct part-count manipulation: make k=3 parts 99% full by redundant prose while two meaningful families are omitted; test whether overlap/information facets block expansion logic gaming. - Construct k-flapping across cycles (3->9->3) from marginal fill changes; decide whether hysteresis or evidence-backed change should be added without freezing adaptation. - Test whether auto-dispatch can occur before second recount/hash/manifest commit, or whether the transition is atomic after all freeze obligations discharge. - Test two simultaneous user pracuj invocations during one pending cycle; stale runner must not mutate current M/Z after a newer runner commits. - Test saturation-stop authority: no actor may mark a lane BLOCKED merely to end the run unless the blocking dependency is named and externally checkable. - Test command-surface semantics: UI presentation must not be treated as execution authority; clicking/copying starts the same governed run state.

Drag to resize
Drag to resize
Drag to resize
Drag to resize