
ctrl
Prompt
# Post-CL0 Architecture Review β Round A: Diagnosis Before Design You are reviewing a CS2 human-behavior imitation/control stack after completion of its first live closed-loop integration phase, **CL0**. Your task is **diagnosis before design**. Do **not** propose the next roadmap yet. Determine what the new evidence actually establishes, what behavioral problem has been exposed, and how it should update the main architectural hypotheses. ## 1. Established CL0 evidence The following live causal chain has now been demonstrated end-to-end: **authoritative Source 2 state β native feedback β exact causal W8 state reconstruction β frozen neural predictor β behavioral reference β deterministic feedback controller β native motor actuation β Source 2 physical response β subsequent feedback** Established results: - Ground planar SysID is sufficient for the tested regime: - command-speed identity through 150 u/s; - ~240 u/s saturation; - ~904 u/sΒ² post-kick acceleration; - planar isotropy; - 64 Hz integration. - The deterministic controller passed a **15/15 live trajectory matrix**, including tracking, settling, directional reversals, and scheduled jump takeoff. - Native feedback/actuation transport is causal, bounded, and cadence-safe. - Frozen predictor weights, normalization, and reference conversion are bit-exact. - Exact causal W8 history reconstruction is established. - Delay sensitivity is characterized; a later native onset loses a physical actuation tick, and bounded software compensation did not materially recover it. - The final full-stack smoke executed **16/16 predictor-driven commands** with: - 0 cadence misses; - 0 expiries; - 0 transport faults; - exact causal accounting. ### Final smoke observation The initial W8 history was completely stationary. The frozen H1 imitation predictor produced approximately: - reference velocity: **(-0.292, +0.224) u/s** - resulting controller command magnitude: **~0.46 u/s** This was far below the effective grounded movement/friction scale, so Source 2 remained stationary. Subsequent stationary feedback produced essentially the same small prediction again. Therefore: > **Control/integration closure is established. Autonomous behavioral continuation is not.** Do **not** assume this observation proves that goal conditioning, retraining, recurrence, planning, trajectory chunks, exploration, or any other particular remedy is required. ## 2. Relevant prior evidence Protected learned references currently include separate **H1, H64, and H192** specialists. Earlier project evidence: - simple long dense-history GRU experiments were negative; - tested low-resolution/repaired POV branches were negative for the movement/trajectory task; - direct predictorβmotor interpretation was abandoned in favor of predictor/referenceβfeedback-controller separation; - recurrence remains conditional, not prohibited; - retraining has not been earned merely by CL0; - tactical, world-model, and RL branches remain evidence-gated. Previous independent architecture reviews proposed these as **hypotheses**, not decisions: 1. coherent semantic trajectory/reference chunks; 2. multimodal continuation with persistent behavioral modes; 3. flat shared latent vs hierarchical/cascade latent vs coherent joint continuation without an explicit latent; 4. H1 remaining near-deterministic conditional on a higher-level continuation; 5. goal- or endpoint-conditioned behavioral prediction; 6. finite action/response history; 7. recurrent adaptation only if hidden state is empirically required; 8. event-tokenized longer context; 9. inverse-action identifiability rather than naΓ―ve keyboard reconstruction; 10. tactical entity-set models; 11. kNN/retrieval baselines; 12. eventual action-conditioned world models. ## 3. Questions ### A. What exactly did CL0 establish? Separate clearly: - established causal/control facts; - behavioral observations; - unresolved questions. Do not infer more than the evidence warrants. ### B. What is the stationary-loop phenomenon? Give a differential diagnosis. Consider, but do not limit yourself to: - correct conditional mean under the training distribution; - multimodal averaging; - sparse/off-support initialization; - missing intention or goal information; - missing temporal/action state; - rollout/covariate shift; - reference-interface mismatch; - an artificial initialization condition with little relevance to ordinary play. Do **not** collapse these into βthe model needs goals.β ### C. Update the previous hypotheses For each relevant hypothesis, classify it as: **strengthened / weakened / unchanged / currently untestable** Explain why CL0 changes or does not change its plausibility. Pay particular attention to: - three separate horizon outputs vs coherent reference chunks; - point vs multimodal continuation; - flat shared mode vs hierarchical/cascade structure; - goal/endpoint conditioning; - finite action history; - recurrence; - retraining; - whether H64/H192 could help where H1 is locally near zero. ### D. What evidence would distinguish the competing explanations? Propose **diagnostic measurements, not a roadmap**. Prefer measurements that: - reuse frozen models; - reuse existing datasets/evidence; - do not require retraining; - isolate one hypothesis at a time. For each important diagnosis, state what result would support or falsify it. ### E. What would be premature? List architectural decisions that should specifically **not** be made yet. ### F. What important hypothesis, if any, is missing from the previous reviews? ## 4. Constraints - Do not inherit any historical post-CL0 roadmap. - Do not assume CL1 should be next. - Do not assume retraining is required. - Do not conflate controller competence with behavioral competence. - Do not treat the stationary smoke as proof of a single architectural diagnosis. - Preserve H1/H64/H192 as protected scientific controls unless evidence says otherwise. - Source 2 remains physical truth. - Human-information causality must remain strict. - Do not recommend RL/self-play unless preceding imitation/control approaches are shown inadequate. ## 5. Return format Return: 1. **CL0 interpretation** 2. **stationary-loop differential diagnosis** 3. **updated hypothesis table** 4. **highest-value discriminating measurements** 5. **premature conclusions to avoid** 6. **current architecture-neutral scientific judgment** **Do not provide a successor roadmap yet.**