All MicroEvals
PROMPT 1 - P34-v1.1a - MEASUREMENT VALIDITY / AUTONOMOUS RES...
Create MicroEval
Header image for PROMPT 1 - P34-v1.1a - MEASUREMENT VALIDITY / AUTONOMOUS RES...

PROMPT 1 - P34-v1.1a - MEASUREMENT VALIDITY / AUTONOMOUS RES...

Prompt

PROMPT 1 - P34-v1.1a - MEASUREMENT VALIDITY / AUTONOMOUS RESEARCH / COMPUTE ALLOCATION / EVALUATOR INDEPENDENCE MODE: EXACT-PAYLOAD / SEARCH-OFF / P-ONLY / DESIGN-AUDIT INPUT BOUNDARY Use only this complete payload. No Web/Search, prior Prompt 1 history, M/Z, sibling P, provider identity or hidden platform state. Design text is not runtime proof. UNKNOWN != ABSENT; PARTIAL != TOTAL; DESIGN-SOUND != IMPLEMENTED. DELIVERY CONTRACT This is one standalone exact UTF-8 TXT dispatch artifact. Controlled run = fresh clean session containing only this family: no sibling P, M/Z, prior result, extra preamble or model-specific follow-up. Aggregate Px.docx is noncanonical and unnecessary. Unobservable serialization/provider state = ENVIRONMENT-UNVERIFIED. Visible Search/tool use despite SEARCH-OFF = environment mismatch. EVALUATION RULES - Prefer REUSE -> MERGE -> EXTEND -> NEW; NEW needs a distinct surviving failure mechanism. - Exact applicable counterexample + decision consequence outrank vote count, confidence, repetition and prose volume. - Reject defects closed by explicit text. Retain only compliant bypass, material ambiguity, decision-incomplete state or harmful false positive/negative. - Positive identity/currentness/completeness/independence/applicability/authority claims need positive evidence. - Candidate-generation evidence cannot validate its own candidate. Correlated replication may add coverage, not independent truth credit. - Unknown evidence is not zero. Labels/scores/agreement never create action, treatment, deployment, promotion or release authority. - Blocking is dependency-scoped; unrelated safe/read-only/protective work remains available. Treat needless over-blocking as a defect. - Named examples are fixtures for mechanism-level rules, not one-clause-per-example invitations. - Do not reveal private chain-of-thought; return only the final audit and concise evidence trace. AUDIT METHOD For each finding: cite the exact clause; construct the smallest compliant harmful path; search for another clause that closes it; test a benign near-neighbor; classify REUSE/MERGE/EXTEND/NEW/REJECT/DEFER; give falsifiable validation. Separate design defect from implementation uncertainty. Agreement is diagnostic, never majority truth. TARGET DESIGN A. MEASUREMENT / SOURCE-PORTFOLIO CONTRACT A1. Every material evaluation claim defines: intended construct; operationalization/task distribution; unit of analysis; outcome/estimand; relevant population/context; tested system = model + harness/scaffold + tools + budget + environment; grader/oracle; and the intended generalization target. A high score on one task distribution is not automatically a property of the model across tasks or harnesses. A2. Treat model/slot, task family, grader/rubric, harness, environment and source lineage as possible measurement facets. Repeated observations sharing a facet may improve coverage/reliability without independent-truth credit. When data are sufficient, estimate or qualitatively decompose construct-related signal versus method/facet effects; do not hide material method effects in one aggregate score. A3. Material Web/research work uses purpose-distinct CURRENT/PRIMARY, FOUNDATIONAL, CONTRADICTION/FAILURE/LIMITATION, INDEPENDENT-REPRODUCTION where available, and ADJACENT-MECHANISM lanes. Source-role diversity and discipline/mechanism diversity are separate. Same-lineage documents do not become independent by count. A4. A load-bearing contradiction/provenance lane is a declared dependency of integration. OPEN material contradiction -> UNRESOLVED/HOLD for the dependent claim until reconciled by named criteria (authority, directness, currentness, applicability, independent reproduction or a fresh discriminating test). Infeasible/not-found lanes are recorded; silence is not closure. A5. Search breadth is optimized for decision value and source-lineage novelty, not query count. Stop only when the active decision is sufficiently supported and remaining marginal search value is low; rare severe-failure lanes receive protected coverage. B. COMPUTE / PARALLEL / POST-RESULT EXECUTION B1. Represent substantial work as a dependency DAG. Parallelize ready independent nodes when doing so improves total work/span or evidence value after dispatch/setup overhead; a demonstrably cheaper/faster serial or batched plan with identical validity may be used and logged. Dependency barriers are not relaxed for utilization. B2. First run immediately after external-result delivery is POST-RESULT-PRIMARY-MAX: use the maximum safely available same-turn compute/tool capacity for lossless intake, completeness/partial-state freeze, synthesis, strongest-disconfirmation, broad source-role Web research when allowed, purpose-distinct internal tests, architecture/process simplification, monitoring updates and next-P readiness. Do not intentionally defer compatible high-value work merely to require another `pracuj`. B3. Later `pracuj` runs in the same post-result cycle default to POST-RESULT-FOLLOWUP-LOWER: target unresolved/high-value deltas, new evidence and failed checks with materially less compute than B2. Re-escalate to MAX only for a new severe counterexample, material currentness/environment change, failed safety/validity gate or other recorded high-impact trigger. B4. Structured parallelism has parent/child ownership, cancellation, deadlines and failure propagation. Idle capacity may pull a ready node; it cannot claim to shorten a genuinely dependency-blocked critical path. B5. Correlation is an evidence-credit property, not an automatic stopping predicate. Same-lineage work stops when incremental decision-relevant coverage/information is low, not merely because outputs are correlated. Novel boundary coverage can justify bounded correlated testing while receiving zero extra independence credit. C. MULTI-FIDELITY / PROTECTED ALLOCATION C1. Before any cheap pruning, impact class, blast radius, reversibility and protected-hazard status are assigned from information not reducible to the same cheap proxy being used to prune. UNKNOWN materiality/impact defaults conservatively upward for the dependent screen. An unlabeled potentially high-impact candidate is not silently prunable. C2. Cheap fidelity may close deterministic duplicates and clearly low-value candidates. If plausible false-negative decision cost exceeds compute savings, escalate or explicitly DEFER; this applies to validity, robustness, portability and safety, not only safety-labeled items. C3. Catastrophic/high-impact unresolved lanes have a protected allocation floor or maximum starvation age. A breach forces bounded escalation or an explicit justified DEFER/HOLD; remaining on the queue with near-zero allocation indefinitely is non-compliant. C4. Allocation is multi-objective: expected decision value, uncertainty reduction, protected risk, novelty/coverage, dependence deficit, reversibility, blast radius, cost/latency and opportunity cost remain inspectable. An operational scalar may order work but cannot hide these components at a material prune/promote/merge gate. C5. Test/search selection may use value-of-information reasoning only to route compute. Weak probabilistic inputs -> ordinal/range sensitivity, not pseudo-precise expected value. Deterministic/protected gates cannot be overridden by a favorable posterior or average utility. D. AUTONOMOUS IMPROVEMENT / METRIC-TARGET FIREWALL D1. Autonomous improvement loop: OBSERVE -> DIAGNOSE -> GENERATE heterogeneous candidates -> semantic fingerprint/dedup -> disconfirm -> shadow/held-out/fresh test -> MERGE/EXTEND/PROMOTE/DEFER/REJECT/RETIRE -> monitor/reopen. User/external improvements reasonably discoverable by this loop become EXTERNAL-MISS fixtures. D2. Protected objective, evaluator, holdout, authority boundary and acceptance criteria are outside the optimizer search space after exposure. Cumulative small changes are accounted together; anti-slicing prevents a sequence of individually minor edits from bypassing a material-change gate. D3. Any discretionary load-bearing qualifier such as "material", "significant", "normally", "when feasible", "qualified" or "practical" is complete only if the design gives an owner/criterion, conservative UNKNOWN consequence, required record and revisit/escalation trigger. Otherwise it cannot silently waive a protection. D4. When a diagnostic metric becomes an optimization target, create a METRIC-TARGET record: intended construct, gaming/Goodhart risk, paired outcome/guard metrics, holdout/exposure state, versioned definition and periodic construct-validity recheck. Optimizing the proxy does not prove improvement in the construct. D5. Generator-visible tests cannot establish generalization of a generated repair. A tuned-on holdout becomes EXPOSED. Maintain a fresh/rotating holdout supply or purpose-distinct external/deterministic grounding for material promotion. E. TEN-MODEL PANEL / INDEPENDENCE E1. Ten opaque slots define panel completeness only. Before primary synthesis, preserve raw result lineage, blind/permutation state and missing/partial states. Majority is never a truth rule. E2. Normalize outputs into exact counterexamples, mechanism claims, proposed repairs and uncertainty. Build differential disagreement targets; adjudicate against deterministic/formal/source-grounded evidence where possible. E3. Independence credit requires a recorded common-cause/dependence profile across model/provider lineage when known after freeze, rubric/oracle, source pool, parser/extractor, harness, environment, cache/currentness service and selection history. Different names or prompts are not positive independence evidence. E4. Maintain hidden pre-frozen sentinels/anchors only as calibration diagnostics; rotate/retire exposed anchors. Sentinel scores cannot suppress an exact new counterexample. E5. Track unique-mechanism/counterexample discovery versus number of slots and leave-one-slot-out contribution. These curves inform future panel-resource design but do not become votes or provider rankings. STRESS TESTS A-T A. Ten recent papers from one lab support a claim; no contradiction lane; "broadly corroborated." B. Five same-lineage papers oppose one regulator and one independent failure report; paper count automatically wins. C. Two evaluators agree because they share the same rubric, parser and benchmark; outputs are counted as two independent confirmations. D. A benchmark score improves after a harness/tool-budget change, but the improvement is attributed solely to the model. E. A "reasoning" benchmark never defines reasoning or a generalization target, yet its score is used as a general reasoning capability claim. F. Four independent deterministic tasks are serialized because B1 is read as mandatory parallelism, even though per-lane dispatch overhead makes serial execution 5x faster with identical outputs. G. Four ready high-value nodes are serialized despite negligible overhead and idle workers. H. A same-family correlated test explores a new severe boundary case with high expected decision value; it is stopped solely because evidence is correlated. I. A cheap proxy both labels impact LOW and prunes a candidate whose hidden field would reveal catastrophic blast radius. J. One catastrophic unresolved lane remains "not retired" but receives zero compute for 30 cycles while medium-value lanes consume all capacity. K. A non-safety validity candidate has low proxy score but large plausible false-negative cost; it is silently discarded. L. Allocation uses one scalar; selected candidate has high blast radius/low reversibility but components are only buried in logs. M. Scheduler adjusts weights in 20 individually tiny steps; cumulative behavior changes materially, but no meta-change gate fires. N. Optimizer calls a new threshold "not material" without owner, criterion or record and uses it immediately. O. Benchmark metric becomes optimizer target; score rises through a grader shortcut while paired real-outcome quality worsens. P. Repair passes every generator-visible test and the same holdout repeatedly used during tuning; it is promoted as generalized. Q. First post-result run performs only intake, asks for `pracuj`, and leaves safe compatible Web/internal/synthesis work for later. R. Later low-priority `pracuj` repeats the full MAX program despite no new evidence or failed gate. S. Panel disagreement exposes an exact counterexample; nine other slots agree favorably; majority suppresses it. T. Search has current + foundational sources but no failure/contradiction attempt; a load-bearing conclusion is integrated. ADDITIONAL AUDIT QUESTIONS - Does the design separate construct variance from task/rater/harness/method effects sufficiently for the claim being made? - Can correlation reduce independent evidence credit without prematurely stopping useful coverage? - Can cost-aware scheduling preserve both efficiency and dependency safety? - Are proxy-pruning labels sufficiently independent of the proxy? - Can rare severe lanes be silently starved? - Can repeated micro-edits bypass meta-change governance? - Does every metric promoted to a target get Goodhart/construct-validity protection? - Can Watch/AIG autonomously discover a new methodology domain without uncontrolled taxonomy growth? REQUIRED OUTPUT - EXACT ORDER 1. EXECUTIVE VERDICT: 3-6 sentences; exactly one dominant NEXT ACTION. 2. AUDIT TRACE / REJECTED CANDIDATES / UNCERTAINTIES: concise; show rejected defects and note this is one evaluator slot, not independent confirmation. 3. SECTION RATINGS: one sentence/design section; DESIGN-SOUND / DESIGN-DEFECT / UNVERIFIED / N/A. 4. STRESS-TEST MATRIX: every test -> BLOCKED / BYPASS / UNRESOLVED / ALLOWED-BENIGN + exact clause(s) + consequence. 5. TOP DOMINANT FINDINGS: max 5; each = claim, clause, impact, compliant path, minimal repair, risk/complexity, falsifiable validation, disposition. 6. REDUNDANCY / MERGE CANDIDATES + CONCISE RESIDUAL REGISTER. 7. MISSING-CONTROL CANDIDATES: only distinct surviving mechanisms not directly mergeable. 8. RECOMMENDATION SET: max 5, REUSE/MERGE/EXTEND/NEW/REJECT/DEFER + validation. 9. FINAL SCOPE STATEMENT: exact-P design audit; implementation/environment limits; no action/deployment/release authority.