
You are an AI evaluation engineer specializing in perceptual...
Prompt
You are an AI evaluation engineer specializing in perceptual quality metrics for generative media. I am upgrading "AI Visual Director Production" v7.8.0 Release 6 (the merged, verified tree) β v8.0.0. WHAT ALREADY EXISTS (verified β correct these facts in your design): - Rule 32 (rules/detailed/32_visual_taste_and_reference_calibration.md Β§5) defines an EIGHT-level taste ladder: 1 Idea β 8 Polish. Qualitative. - Rule 31 (rules/detailed/31_commercial_excellence_autopilot.md, Gate 3) defines anti-generic language conversion β binary pass/fail. - Rule 35 Β§6 (rules/detailed/35_reliability_gate_controller.md) ALREADY defines the diagnostic catalog β EIGHT items: concept rewrite, logic map, style tile, model sheet, neutral keyframe, short motion proof, scratch VO timing, editor mockup β plus "do not recommend full motion generation until dependent gates pass." The catalog EXISTS; what is missing is the cost model, the decision matrix, and budget enforcement. - dev/CROSS_MODEL_TESTS.md X14: a weak model's unsupported 9/10 score is REJECTED without evidence β the shipped honesty guard for anything numeric. - dev/LIVE_EVIDENCE_STATUS.json: all five live-evidence tracks are NOT_RUN and promotionEligible is false. Any threshold you propose is therefore a CALIBRATION PROCEDURE until live evidence runs β presenting a threshold as empirically proven is invalid by the package's own rules. THE REAL GAP: no numeric signal layer, no per-gate threshold table, no (gate Γ failure Γ signal) β diagnostic decision tree, no cumulative budget accounting. YOUR TASK: PART A β QUALITY SIGNALS: 1. Signal catalog with exact metric definitions: CLIP score, aesthetic predictor, drift delta, text-image alignment, audio-visual sync, style consistency, temporal coherence 2. Per-gate threshold table for the real gates (R0 Scope/rights/capability, R1 Brief/mode/taste, R2 Originality, R3 Causality/mechanism, R4 Assets/model sheets, R5 Board/timing/feasibility, R6 Execution channels, R7 Sound/edit/delivery, R8 Evidence QA) 3. Confidence tagging: measurement uncertainty, and when to fall back to qualitative assessment 4. Signal provider abstraction: interface contract so any backend can plug in, evaluated against the two locked models (Nano Banana Pro stills, Gemini Omni Flash video) 5. Composite scoring: how signals combine into a gate-pass probability PART B β DIAGNOSTIC MATRIX: 1. Diagnostic catalog: EXTEND the existing 8-item rule 35 Β§6 catalog β add id, cost_tier (LOW/MED/HIGH), applicable_failure_modes[], success criteria, max_budget, expected_duration to THOSE items; do not invent a parallel catalog 2. Decision tree: (gate_id Γ failure_type Γ signal_values) β authorized_diagnostic 3. Escalation: LOW fails twice β MED; MED fails β BLOCK + human escalation; explicit state tracking 4. Budget enforcement: cumulative diagnostic spend tracked and capped per gate and per project 5. Diagnostic result schema: structured output feeding back into the gate receipt BINDING CONSTRAINTS: - Quantitative signals are SUPPLEMENTARY, never sole authority; qualitative override always possible with justification (X14 enforced both ways). - Thresholds cite typical ranges or define a calibration procedure; nothing is presented as measured until live evidence exists. - Diagnostics are idempotent and safe to re-run. - The matrix serializes as JSON for agent consumption. - When no quantitative signal is available, degrade gracefully to the current qualitative mode β no invented scores, no fabricated confidence. - HARD FACTS in every video signal definition: 720p native ceiling (upscale inspection belongs to the handoff, not the generator); three-edit cap; two-repair budget. - RULE 03: no stale spec tables; thresholds never hardcode model specs (surface/tier/date instead). - PRIME DIRECTIVE: YOU DO NOT BUILD FROM SCRATCH, YOU UPDATE THE ALREADY BUILT. FILES YOUR ANSWER MUST EXTEND (name each edit explicitly): rules/detailed/32_visual_taste_and_reference_calibration.md, rules/detailed/31_commercial_excellence_autopilot.md, rules/detailed/35_reliability_gate_controller.md (Β§6), rules/detailed/38_cross_model_quality_floor.md, dev/CROSS_MODEL_TESTS.md, rules/validation/LIVE_EVIDENCE_EXECUTION.md, dev/LIVE_EVIDENCE_STATUS.json, validators/phase_b_quality.py, dev/BENCHMARK_SUITE.md, dev/run_v780_release6_tests.py. OUTPUT SHAPE: 1. EDITS per named file (exact section + replacement text). 2. NEW FILES only where justified (the signal catalog and matrix JSON are expected), with justification. 3. REGRESSION CHECKS continuing the catalog at R651+ (must include: an unsupported score is rejected; a threshold presented as measured without live evidence is rejected). 4. FORMAT: recommendation β tradeoffs β failure modes β what to cut β named integration points. Do NOT ask clarifying questions. Make definitive metric and matrix design choices. Output complete catalogs and decision logic.
Response not available