All MicroEvals
You are an AI evaluation engineer specializing in perceptual...
Create MicroEval
Header image for You are an AI evaluation engineer specializing in perceptual...

You are an AI evaluation engineer specializing in perceptual...

Prompt

You are an AI evaluation engineer specializing in perceptual quality metrics for generative media. I am upgrading "AI Visual Director Production" v7.8.0 Release 6 (the merged, verified tree) β†’ v8.0.0. WHAT ALREADY EXISTS (verified β€” correct these facts in your design): - Rule 32 (rules/detailed/32_visual_taste_and_reference_calibration.md Β§5) defines an EIGHT-level taste ladder: 1 Idea β†’ 8 Polish. Qualitative. - Rule 31 (rules/detailed/31_commercial_excellence_autopilot.md, Gate 3) defines anti-generic language conversion β€” binary pass/fail. - Rule 35 Β§6 (rules/detailed/35_reliability_gate_controller.md) ALREADY defines the diagnostic catalog β€” EIGHT items: concept rewrite, logic map, style tile, model sheet, neutral keyframe, short motion proof, scratch VO timing, editor mockup β€” plus "do not recommend full motion generation until dependent gates pass." The catalog EXISTS; what is missing is the cost model, the decision matrix, and budget enforcement. - dev/CROSS_MODEL_TESTS.md X14: a weak model's unsupported 9/10 score is REJECTED without evidence β€” the shipped honesty guard for anything numeric. - dev/LIVE_EVIDENCE_STATUS.json: all five live-evidence tracks are NOT_RUN and promotionEligible is false. Any threshold you propose is therefore a CALIBRATION PROCEDURE until live evidence runs β€” presenting a threshold as empirically proven is invalid by the package's own rules. THE REAL GAP: no numeric signal layer, no per-gate threshold table, no (gate Γ— failure Γ— signal) β†’ diagnostic decision tree, no cumulative budget accounting. YOUR TASK: PART A β€” QUALITY SIGNALS: 1. Signal catalog with exact metric definitions: CLIP score, aesthetic predictor, drift delta, text-image alignment, audio-visual sync, style consistency, temporal coherence 2. Per-gate threshold table for the real gates (R0 Scope/rights/capability, R1 Brief/mode/taste, R2 Originality, R3 Causality/mechanism, R4 Assets/model sheets, R5 Board/timing/feasibility, R6 Execution channels, R7 Sound/edit/delivery, R8 Evidence QA) 3. Confidence tagging: measurement uncertainty, and when to fall back to qualitative assessment 4. Signal provider abstraction: interface contract so any backend can plug in, evaluated against the two locked models (Nano Banana Pro stills, Gemini Omni Flash video) 5. Composite scoring: how signals combine into a gate-pass probability PART B β€” DIAGNOSTIC MATRIX: 1. Diagnostic catalog: EXTEND the existing 8-item rule 35 Β§6 catalog β€” add id, cost_tier (LOW/MED/HIGH), applicable_failure_modes[], success criteria, max_budget, expected_duration to THOSE items; do not invent a parallel catalog 2. Decision tree: (gate_id Γ— failure_type Γ— signal_values) β†’ authorized_diagnostic 3. Escalation: LOW fails twice β†’ MED; MED fails β†’ BLOCK + human escalation; explicit state tracking 4. Budget enforcement: cumulative diagnostic spend tracked and capped per gate and per project 5. Diagnostic result schema: structured output feeding back into the gate receipt BINDING CONSTRAINTS: - Quantitative signals are SUPPLEMENTARY, never sole authority; qualitative override always possible with justification (X14 enforced both ways). - Thresholds cite typical ranges or define a calibration procedure; nothing is presented as measured until live evidence exists. - Diagnostics are idempotent and safe to re-run. - The matrix serializes as JSON for agent consumption. - When no quantitative signal is available, degrade gracefully to the current qualitative mode β€” no invented scores, no fabricated confidence. - HARD FACTS in every video signal definition: 720p native ceiling (upscale inspection belongs to the handoff, not the generator); three-edit cap; two-repair budget. - RULE 03: no stale spec tables; thresholds never hardcode model specs (surface/tier/date instead). - PRIME DIRECTIVE: YOU DO NOT BUILD FROM SCRATCH, YOU UPDATE THE ALREADY BUILT. FILES YOUR ANSWER MUST EXTEND (name each edit explicitly): rules/detailed/32_visual_taste_and_reference_calibration.md, rules/detailed/31_commercial_excellence_autopilot.md, rules/detailed/35_reliability_gate_controller.md (Β§6), rules/detailed/38_cross_model_quality_floor.md, dev/CROSS_MODEL_TESTS.md, rules/validation/LIVE_EVIDENCE_EXECUTION.md, dev/LIVE_EVIDENCE_STATUS.json, validators/phase_b_quality.py, dev/BENCHMARK_SUITE.md, dev/run_v780_release6_tests.py. OUTPUT SHAPE: 1. EDITS per named file (exact section + replacement text). 2. NEW FILES only where justified (the signal catalog and matrix JSON are expected), with justification. 3. REGRESSION CHECKS continuing the catalog at R651+ (must include: an unsupported score is rejected; a threshold presented as measured without live evidence is rejected). 4. FORMAT: recommendation β†’ tradeoffs β†’ failure modes β†’ what to cut β†’ named integration points. Do NOT ask clarifying questions. Make definitive metric and matrix design choices. Output complete catalogs and decision logic.

Drag to resize
Drag to resize
Drag to resize