All MicroEvals
You are being consulted as a world-class expert to stress-te...
Create MicroEval
Header image for You are being consulted as a world-class expert to stress-te...

You are being consulted as a world-class expert to stress-te...

Prompt

You are being consulted as a world-class expert to stress-test and architecturally harden an AI Visual Director skill system (v7.8.0-release-6). This system is a structured set of markdown rules designed to force AI models to produce professional-grade film/video direction outputs instead of generic content. We have just merged "retention-aware hook integration" into three core files, but forensic audit reveals the structure is conceptually correct yet execution-fragile when consumed by LLMs under token pressure. YOUR ROLE: You are a Cognitive Systems Architect + Senior Film Editor + AI Evaluation Researcher. Your job is NOT to praise the work. Your job is to find where this system will silently fail in production with real AI models, and to design solutions that are robust to known LLM failure modes. CONTEXT: Read the following three updated rule files carefully. They cross-reference each other and form the retention-hook subsystem: 23_hooks_and_packaging.md (owns the retention spec) 14_faceless_longform.md (applies retention to chapter-based video) 08_directing_and_editing_grammar.md (ties retention to shot planning and coverage) [# 14 Faceless and Long-Form Video ## Channel/show kit Lock only production-relevant constants: ```text Audience promise | recurring format | mode/style token | narrator token | caption/type system | music palette | chapter grammar | thumbnail system ``` ## Script-led workflow ```text research/fact-check β†’ script β†’ chapter map β†’ scratch/final VO β†’ storyboard/animatic β†’ chapter-by-chapter visuals β†’ edit/sound β†’ packaging β†’ evidence-based review ``` Do not generate one image per sentence literally. Mix evidence, visual metaphor, maps/diagrams, archive-like reconstruction with disclosure, details, reactions, and breathing room. ## Retention without formula spam Each chapter needs a question, development, and payoff. Re-hook when the argument genuinely turnsβ€”not at an arbitrary universal interval. Use real audience data only after publication; do not invent retention predictions. ## Continuity at length - Plan by chapters, not a flat hundred-clip list. - Maintain approved bibles and chapter-local look tokens. - Review drift at the end of each chapter. - Use last-frame chaining only for continuous action. - Keep voice, names, dates, and factual claims in a source log. - Use `19` session handoff when work spans sessions. ## Mode C faceless route For diagram/explainer-led channels, combine this file with `15`, `09`, `10`, `12`, and `23`. Motion graphics require designed scenes, not generic B-roll. # 08 Directing and Editing Grammar Camera movement is only one part of directing. Build shots for story function, spatial clarity, and the cut. ## Shot card ```text Shot ID: Story function: [orient / reveal / prove / react / transition / payoff] Shot size + angle: Lens perspective: Subject and action: Screen direction / eyeline: Camera movement: Foreground / midground / background: Entry and exit frame: Cut motivation: Audio cue: Reference IDs / generation mode: ``` ## Shot sizes and functions - **Extreme wide / wide:** geography, scale, isolation, world rules. - **Medium:** behavior, interaction, product use. - **Close-up:** emotion, evidence, mechanism. - **Extreme close-up / insert:** tactile detail, proof, transition bridge. - **Reaction:** meaning after an event; often more valuable than another action shot. Use progression intentionally. A sequence of equally framed medium shots feels generated rather than directed. ## Perspective - Wide perspective emphasizes space and movement but can distort near objects. - Normal perspective reads naturally and is flexible for coverage. - Long perspective compresses space and isolates details or faces. - Macro reveals material truth but loses geography. Keep perspective coherent across coverage unless the change expresses a story turn. ## Movement families - **Locked:** lets action, performance, or design carry attention. - **Push/pull:** changes emotional proximity or reveals context. - **Pan/tilt:** redirects attention within a stable position. - **Track/follow:** shares subject motion and screen direction. - **Orbit/arc:** reveals form or changes relationship to background. - **Rise/drop/crane:** reveals scale or hierarchy. - **Handheld/POV:** subjective immediacy; control amplitude and purpose. - **Focus/optical:** transfers attention without moving the camera. Name direction, speed, endpoint, and what stays stable. ## Continuity rules For continuous scenes: - establish geography before complex coverage; - preserve left/right screen direction across cuts; - match eyelines and prop/hand state; - respect the action axis unless the crossing is shown; - cut on action or motivated sound when helpful; - carry the previous endpoint into the next start frame; - use inserts/cutaways to bridge unavoidable continuity defects. ## Edit grammar Cut because one of these changes: - information; - emotion/reaction; - action phase; - location/time; - scale/detail; - sound beat; - visual match or contrast. Do not cut merely because a generated clip ended. ## Coverage templates ### 15-second product ad ```text problem/evidence close-up β†’ tactile reveal β†’ mechanism/use β†’ reaction/result β†’ clean hero/CTA ``` ### Dialogue beat ```text establishing two-shot β†’ speaker medium/close β†’ listener reaction β†’ insert if relevant β†’ changed two-shot ``` ### Cinematic scene ```text geography β†’ objective β†’ obstacle evidence β†’ action/reaction β†’ turn β†’ exit image ``` ### Animation Prioritize readable silhouette, staging, anticipation, action, reaction, and hold. Camera movement is optional; character staging can carry the shot. ## Intentional rule breaking ```text Default being broken: Creative reason: Risk introduced: Control/coverage fallback: ``` # 23 Hooks, Packaging, and First-Frame Design A hook is a concept-specific promise or tensionβ€”not a stock phrase. ## Hook construction ```text Audience state β†’ specific tension/desire β†’ visible evidence β†’ information gap β†’ honest payoff ``` Generate at least three **mechanically different** options: - evidence-first; - contradiction; - sensory transformation; - question with visible stakes; - character objective/obstacle; - result first, mechanism withheld. Avoid empty formulas such as β€œstop scrolling,” β€œnobody is talking about this,” or β€œthe industry doesn’t want you to know” unless the statement is true, necessary, and specific. ## Hook test card ```text Hook | visual evidence | read-aloud time | format fit | promise paid off? | clichΓ© risk | select/reject ``` ## Packaging system Title, thumbnail/cover, and opening should add different information while making one honest promise. ### Thumbnail/cover - one dominant read at small size; - concept-specific object, face, or event; - controlled contrast and negative space; - minimal editable text only when it adds meaning; - no misleading scale, fake result, or unsupported claim. ### First frame Use `18`’s first-frame check. A static frame is valid when its stillness creates tension, luxury, scale, or contrast; do not ban static openings universally. ## A/B discipline Change one packaging variable, choose the decision metric before looking, record sample/context, and call results local until repeated. Do not fabricate predicted CTR or retention. ] KNOWN FAILURE MODES WE HAVE ALREADY IDENTIFIED (DO NOT REPEAT THESE AS FINDINGS β€” SOLVE THEM): Circular cross-references cause reference loops under token pressure "≀15s or 10%" latency ceiling is unmeasurable without script timing annotations "Mechanically different" re-hooks lack an enumerated taxonomy β†’ AI produces semantic clones Payoff audit is prose, not structured output β†’ becomes optional Anti-patterns are negative constraints β†’ LLMs ignore prohibitions Session handoff schema doesn’t carry hook state across sessions No mandatory workflow gate forces execution before visual generation WHAT I NEED FROM YOU (ANSWER ALL FIVE): Q1 β€” EXECUTION ENFORCEMENT ARCHITECTURE: Design the exact structured artifact template (field names, enumerated option lists, validation rules) for the Hook-Retention Audit that makes compliance machine-verifiable rather than prose-aspirational. Specify which fields must be present vs. conditional. Define what constitutes a VALID vs. INVALID entry for each field. This template must work inside a single LLM turn without external tooling. Q2 β€” TAXONOMY OF MECHANICAL DIFFERENCE: Provide a complete, mutually exclusive, collectively exhaustive taxonomy of re-hook mechanical types that an AI can select from via constrained decoding or structured output. For each type, give: (a) precise definition, (b) what makes it mechanically distinct from every other type, (c) one valid example, (d) one invalid example that looks similar but fails the distinction test. Minimum 6 types. Maximum 10. Q3 β€” CROSS-REFERENCE RESILIENCE PATTERN: Propose the exact structural pattern that eliminates circular dependency risk while preserving necessary cross-file coordination. Should we use hub-spoke, layered inheritance, inline expansion, or something else? Provide the concrete rewrite instructions for all three files showing exactly what stays, what moves, and what gets replaced with a pointer. Justify why your pattern survives token-window truncation better than the current structure. Q4 β€” TEMPORAL GROUNDING WITHOUT EXTERNAL TOOLS: LLMs cannot accurately estimate runtime from text alone. Design a self-contained temporal annotation protocol that can be embedded directly into script/chapter-map output format such that promise-to-payoff latency becomes computable from the text artifact itself, without requiring the model to "guess" durations. Specify the annotation syntax, placement rules, and how cumulative timestamps propagate through nested structures (chapters β†’ segments β†’ beats). Q5 β€” ADVERSARIAL EVALUATION RUBRIC: Create a 10-point scoring rubric that an evaluator model can apply to ANY output produced using this skill to detect silent compliance failures. Each point must be: (a) binary or 3-tier scored (not subjective), (b) grounded in observable output features (not intent), (c) weighted by production impact, (d) calibrated so that a superficially compliant but substantively hollow output scores ≀4/10. Include the exact phrasing the evaluator should use for each criterion. CONSTRAINTS ON YOUR RESPONSE: Do NOT summarize the files back to me. I wrote them. Do NOT offer generic advice like "add more examples" or "use chain-of-thought." Be specific to this system's architecture. Do NOT assume external tools, APIs, or multi-turn verification. Solutions must work within a single structured LLM response. If you believe a problem is unsolvable within these constraints, say so explicitly and explain WHY, then propose the closest achievable approximation. Prioritize solutions that survive model upgrades, context window changes, and varying capability tiers. Fragile cleverness is worse than robust simplicity. I am building this to be the definitive AI visual direction skill system. Treat this as a peer-level architectural review, not a tutoring session.

Drag to resize
Drag to resize
Drag to resize
Drag to resize