average goal usage
Prompt
Fair local measurement of R1 (v212 vs v212 + R1 only) Workspace: D:\G Documents\Game Dev\PROJECTFOLDER. Task: docs/ai-orchestrator/tasks/2026-09-27-performance-repairs/. Start from experiments/fair-r1-20260930/, which holds the baseline and R1-only handlers, the guard and the harness. This goal replaces the preregistration and pass rules of the 30 September fair round. Don't reclassify that round's runs. Under the frozen synthetic load, does R1 cut server network cost by at least 5%, without making the server or the receiving clients do measurably more work? R1 here means 20 Hz coalescing of the aim relay, applied to v212 and nothing else. This measures transport cost, not production lag; the local server holds 60 Hz in every arm. Production comes later. - Separate launches per arm. Every launch was a fresh chance for startup failures, Adonis kicks, window differences and drift. Nine launches gave zero valid comparisons. - Instead: interleave both arms inside one Play session. - Add a diagnostic switch to the server relay so one session runs baseline and R1 windows in turn: B, R, B, R, B, R, B. The server, clients, rigs and windows stay the same throughout. Drain and pause between windows so no queued work carries over. - Both code paths must be identical to the baseline and R1-only handlers; only the selector is new. The switch exists only in the test copy. - Focus. Client frame time depends on which window has focus, because background Studio windows run at 15 fps, and controlling focus failed repeatedly. - Instead: nothing that can fail R1 depends on focus or window size. Log frame time and focus, but only report them, and spend no effort controlling focus or window size. - Noise. A three-run noise range measured in advance was narrower than the baseline's own drift. - Instead: noise is the spread of all the baseline windows, interleaved with the R1 windows. A metric counts as worse only if it's worse beyond that spread and by more than 5%. - Ping. Loopback raw ping (1β5 ms) and Adonis ping (21β25 ms run to run) are noise on this PC. - Instead: report both; neither can fail R1. - Adonis kicks. Adonis r10004 kicks lost runs. - Instead: if Adonis kicks or fails twice, disable the Adonis loader in the test copy. That affects both arms equally, since R1 doesn't touch Adonis. Report Adonis ping as unavailable. - Disk. C: filled up. - Instead: the guard checks C: (no launch below 5 GB, stop below 2 GB). Run tools/cleanup-roblox-test-cache.ps1 after each session. - Side work and paperwork. Side work grew before any result existed (OBS, 12 unused files), and paperwork and reconciliation ate the budget. - Instead: follow the guardrails below. 1. Arms: v212, and v212 + R1 only. Nothing else differs. 2. Workload: the frozen calibration workload: 2 clients plus 38 synthetic rigs, offering 60 updates a second. Each window settles, measures for a fixed time, then drains, identically every time. 3. Headline (must pass): server send KB/s and outgoing aim relay calls per second. R1 must be lower by at least 5% and by more than the baseline spread. 4. Can fail R1: - Server heartbeat, mean and worst 10%. - The receiving clients' aim-handler work: CPU time per second inside the aim receiver, and tweens created per second. - Each of these fails R1 only if it's worse beyond the baseline spread and by more than 5%. - Correctness also fails R1: - After each drain, every accepted input is accounted for: relayed, coalesced, or counted stale. - Stale must be 0 under the unchanged load. - Every rig's final pose equals its last offered pose, in both arms. - R1 windows show no new errors. 5. Reported only: client frame times and focus, raw ping, Adonis ping. 6. Valid window: the workload completed, both clients stayed connected, the guard didn't stop it, and the correctness checks ran. Focus and window size don't affect validity. 7. Replicate: the verdict needs the B/R/B/R/B/R/B pattern in at least two separate sessions, reaching the same conclusion. If they disagree, run a third. 8. Pre-register the analysis in a short file before the first launch, and don't change it afterwards. 9. Unchanged: - Memory floors: admission 10,240 MiB; stops at 2,560 MiB physical or 4,096 MiB commit. - The C: disk guard. - Production stays read-only. No publishing, no player data. The OOM task stays closed. Everything not listed above is yours: how to build the switch, window lengths and gaps, instrumentation, harness fixes, worker setup and models, and stopping early once the answer is clear. If the in-session design proves unworkable (two attempts, or about 3 points spent without a valid session), fall back to separate launches. Use the same metrics and rules, interleaved B/R/B/R/B/R/B. - Don't build ahead. Do no work for a later step until the current question has its answer. - Side-effort tripwire. If a side effort costs about 2 points, or fails twice, without producing a measurement, stop it. Write one line on why, then take the simpler route. - No new installs or OS changes. Never wait on a system dialog. - Light paperwork. One closure.json status per launch, one analysis, and one independent check at the verdict. No per-step dispositions, and no reconciliation beyond confirming the test processes are gone. - Lean root. Workers parse the logs and return summary tables. - Never wait on me. Write questions down with the default you took, then continue. 1. Floating actor. Explain the floating actor and odd gun angle. They also appear in v212, so check the harness first. 2. Look check. Check how R1 looks from the observer's view: aiming, switching, unequip, death and respawn. - Use matched frame sequences from Studio screen capture. - A fresh reviewer judges them by what a player would notice. No OBS and no audio. 3. Publish candidate. Build v212 + R1 only, with no switch and no instrumentation, and record its hash. There is no usage stop. Keep working until the work is done or the allowance runs out. Running out cuts you off mid-step with no chance to wrap up, so stay ready for that at all times: - Every launch has its external guard with the memory floors, the C: disk check and a total time limit. If you're cut off mid-run, the guard still stops and cleans up the test processes. - Before each launch, bring the handoff, STATE.md and RESUME.md up to date with everything so far. - Never leave a half-applied edit. Save a restorable copy before any change to a place file, and make each change in one batch. - Late launches are fine. If one is cut off, its raw logs and guard record still show what happened. - RESUME.md must say what a fresh chat does first: 1. Check that no test processes are running. 2. Run the cache cleanup. 3. Check the last session's records. 4. Then continue. Keep these at the top of reports/MORNING_REPORT.md and update them after every session. Don't wait until the end. - The verdict so far. - A table for each window and session: arm, send, relays, heartbeat mean and worst 10%, client handler time and tweens. Add the reported-only columns: frame time, focus and pings. - The baseline spread for each metric. - What's next. - Usage spent so far. Keep STATE.md and RESUME.md current at the same points.
Response not available