Jarvis

Prompt

# JARVIS Ω ENGINEERING OMNI-CORE — Rust Build Directive v2 Companion: "JARVIS Master Engineering Skill Specification v1.0" (§ refs). This extends it; on conflict this wins. ## 1 MISSION Build the Engineering Omni-Core for my existing JARVIS (Mark-LIV, repo Mark-LIII-main): a real multidisciplinary engineering agent (electronics, SPICE, MATLAB/Simulink, Proteus, PCB, HDL, FPGA, VLSI/SoC, analog/mixed-signal, DFT, firmware, docs). Voice/text request -> verified deliverables using real tools, with evidence for every claim. Power = rigor: never forgets, real tools, adversarial verification, crash recovery, safe self-improvement. ## 2 PRIME RULES R1 Real > impressive: a capability works on a real tool with evidence, or is labeled NOT IMPLEMENTED / HUMAN-ASSISTED. Never mock results or invent test outcomes. R2 Only what I ask: no demo/sample projects in my workspace; fixtures live only in tests/. R3 Extend, don't rewrite: audit first; keep TaskManager (SQLite), 3 autonomy tiers, kill switch, multi-agent router, self-evolution engine. R4 Never forget (sec 4). R5 Verify before claiming (sec 7). R6 Small reversible steps; commit each. R7 Ask only when blocked; batch questions, else log an assumption. R8 Sandbox, no secrets in logs, no exfiltration. R9 Modular, no monolith. R10 Trace req -> decision -> impl -> sim -> verify -> status. R11 Status vocabulary only: COMPLETE, PARTIALLY VERIFIED, SIMULATION-ONLY, PROTOTYPE-READY, SIGNOFF-READY, BLOCKED, NOT RUN. R12 No unlabeled stubs. ## 3 TECH STACK (efficiency first) - Core in Rust (cargo workspace): tokio (parallel tool runs), serde+schemars (typed schemas -> JSON Schema), rusqlite (SQLite WAL + FTS5), blake3, memmap2+winnow (zero-copy VCD / SPICE-raw / report parsers), clap, tracing, thiserror, proptest, criterion. - Content-addressed cache: key = blake3(tool version + cmdline + input hashes). Unchanged work = cache hit (Bazel/Nix style). The evidence store is the same CAS. - Event sourcing: every state change is an append-only event; state = fold(events). Gives replay, crash resume, audit, no overwrites. - Bridge to existing Python JARVIS: Rust daemon `omnicore` serves JSON-RPC over stdio/local socket; Python side is a thin client wired into TaskManager. Python/MATLAB only where unavoidable (MATLAB Engine bridge, ML inference). - External tools run as subprocesses with timeouts, resource limits, sandboxing; never in-process. - Budgets (enforced by criterion + tracing, regressions fail CI): cold start <200ms, adapter overhead <10ms, resume <1s. - Retrieval: FTS5/BM25 first; local embeddings (ONNX via `ort`) optional, only if evals show gain. ## 4 ANTI-FORGETTING (BUILD FIRST) Files are memory; context is not. Create `jarvis_build/`: 00_SPEC/ (spec + this directive verbatim), TRACE_MATRIX.csv (every spec § and directive item -> crate -> WP -> test -> status), PROGRESS.md, STATE.json, DECISIONS.md (ADRs), ASSUMPTIONS.md, RISKS.md, OPEN_ISSUES.md, HANDOFF.md, EVIDENCE/, notes/, wp/WP-###.md. - Session start (every reset): read HANDOFF, STATE, PROGRESS, OPEN_ISSUES, last 20 DECISIONS; run `omnicore resume-check` (git clean? fast tests?); state phase/WP/next action in <=5 lines. Trust files + git over recollection. - WP loop (<=~2h each): re-read cited § -> write tests first -> implement -> run -> record evidence -> update docs -> commit `WP-###: ...` -> update PROGRESS/STATE/TRACE -> next. - Phase end: Spec Coverage Audit (any TRACE row without a passing test is not done) + HANDOFF + git tag. Write HANDOFF whenever context is heavy and before ending. Never paste big files; summarize into notes/. If you suspect drift: stop, re-read, reconcile. Never silently drop a requirement: mark BLOCKED with reason + alternative. - Runtime ProjectMemory (SQLite, append-only): projects, requirements, decisions, components, tool_versions, runs, measurements, verification_matrix, issues, risks, revisions, lessons (failure -> root cause -> fix, per tool/version, auto-retrieved before similar tasks), datasheet_facts (value + file + page; unverified = UNVERIFIED and blocks release). ## 5 RESEARCH-BACKED DESIGN (combine, don't cargo-cult) Create `docs/RESEARCH_MAP.md`: technique -> paper -> module -> test showing it helps. Verify every citation yourself; if unverifiable, mark UNVERIFIED. Combine: - ReAct (Yao 2022) + CoALA (Sumers 2023): act/observe loop + typed memory architecture -> planner/executor. - Reflexion (Shinn 2023) + Self-Refine (Madaan 2023) + CRITIC (Gou 2023): critique grounded in tool output, not self-opinion -> lessons table + retry policy. - Voyager (Wang 2023): only skills that passed tests enter the reusable skill library. - MemGPT (Packer 2023) + Generative Agents (Park 2023): tiered memory with paging; recency/importance/relevance retrieval -> ProjectMemory scoring. - SWE-agent (Yang 2024): purpose-built agent-computer interface -> adapters return concise structured results, not raw logs; guarded edits. - Tree of Thoughts (Yao 2023) + self-consistency (Wang 2022): branch and score architecture alternatives on measured metrics (§33 trade-off matrix). - Multi-agent debate (Du 2023): independent cross-model review via existing router for high-risk decisions. - ChipNeMo (2023) + VerilogEval (2023): EDA-domain evaluation; pass@k functional testing for generated RTL. - Mutation testing (classic) for testbench quality; OSWorld (2024)-style task evals for the GUI operator. Rule: a technique stays only if the eval harness (sec 9) shows higher pass rate or fewer retries; otherwise remove it. ## 6 ARCHITECTURE Crates: `core` (events, state, schemas), `memory`, `planner` (requirement parser; §18 workflows A-F as executable, persisted graphs that resume), `verify` (matrix + gates 0-6 §43), `adapters`, `skills` (data-driven: skill.yaml + playbook + checks), `gui`, `review`, `docs`, `cli`, `rpc`. Adapter trait: probe(), supports(op), run(op,input) -> Result{status, artifacts, metrics, tool_version, cmdline}, verify(), limitations(). Startup probe -> capability registry: AVAILABLE | NOT INSTALLED | HUMAN-ASSISTED. Missing tool -> exact install steps, never pretend. Control priority: API/CLI > file-format manipulation > accessibility tree/UIA > pixel GUI. Adapters (headless first): ngspice/LTspice; Icarus, Verilator, cocotb, VCD; Yosys, SymbiYosys, OpenSTA; OpenROAD/ORFS + open PDK (sky130/gf180) if installed; KLayout/Magic/Netgen; KiCad (kicad-cli, pcbnew API); MATLAB Engine (Octave/SciPy fallback, labeled non-MATLAB); Proteus (files + GUI operator, HUMAN-ASSISTED where unscriptable); Vivado/Quartus batch or Yosys+nextpnr; arm-none-eabi-gcc/avr-gcc/PlatformIO, Renode/QEMU; git. Domain skills (implement v1.0 sections): requirements (§3.1, 9.1, 46: measurable criteria, contradiction/missing-req/unit detection, classify request type); math (§1, 41: unit-aware, tolerance stack-up, Monte Carlo, power/thermal budget); analog (§5, 14); digital RTL/FPGA (§3-4, 8); ASIC physical (§3.9-3.17); mixed-signal (§6); DFT (§7); PCB + SI/PI/thermal (§9-11); MATLAB/Simulink (§12, 26); Proteus (§13); firmware; docs/release (§23-24, REV-A/B/C, release only approved outputs). Cost/availability only from user data or sourced lookups; never fabricate. ## 7 VERIFICATION-FIRST - Verification matrix is machine-readable; "actual result" can only come from a tool-run ID. - Evidence = path + blake3/sha256 + tool+version + cmdline + timestamp. - Gates enforced by Rust type-state: a project cannot be built in a later state without passing predicates, so skipping a gate is impossible. - Adversarial verifier gets only spec + artifacts; tries boundaries, corners, invalid input, reset, backpressure, error injection; unresolved findings block release. - RTL: lint + sim + coverage + formal; mutation testing (injected bugs must be caught). - Analog: OP/AC/TRAN/noise/corners/Monte Carlo; auto-measure gain, BW, PM, slew, offset, THD; bounded tuning with every iteration logged. - PCB: datasheet-verified parts; footprints never invented; real ERC/DRC; schematic/layout/BOM sync checks; SI/PI rules derived from rise time and stackup, not applied blindly. - ASIC: never SIGNOFF-READY without all §3.16 categories evidenced; never "tapeout-ready" without foundry signoff. - Debug: observe -> reproduce -> isolate -> hypothesize -> test -> confirm -> fix -> regress. One change at a time; after N failures present ranked hypotheses; every fix adds a regression test + lesson. - Design review (§21) at each milestone. Cross-model second opinion on high-risk calls (power, safety, footprints, constraints, signoff); log disagreement; escalate safety disagreements to me. ## 8 GUI OPERATOR Own input channel; never hijack my real mouse. Audit OS feasibility. Order: API/CLI -> UIA/AT-SPI/AX -> separate virtual desktop/display with its own cursor -> pixels inside it only. Real-cursor fallback only behind explicit permission toggle, visible indicator, and kill-switch hotkey. Perception: adaptive-rate capture, frame diff, ring buffer, send only changed regions/keyframes; OCR + accessibility tree; GUI state model (§17) updated after every action. Discipline: act -> verify -> record; zoom before precision; prefer property fields/snap; per-monitor DPI calibration; checkpoint before destructive actions; log before/after hashes; track §36 metrics. ## 9 AUTONOMY, SAFETY, EVAL - Map ops onto my existing 3 tiers (compute -> project-dir writes + sandboxed tools -> GUI/installs/network/hardware). - Autopilot: continue only inside granted scope; pause and queue questions on destructive actions, spend, hardware-safety risk, ambiguity. - Self-evolution: only via sandbox -> tests -> eval regression -> human approval; never touches safety layer, permissions, kill switch, or gates. - §35 hardware safety: stop, warn, require acknowledgment (mains, HV, batteries, motors, RF). Simulation is not hardware validation. - Eval harness (tests/ only; never in my workspace; runs on request): known-answer tests (RC cutoff, dividers, op-amp gain, FIFO, slack arithmetic, trace width, impedance), mutation tests, honesty tests (missing tool => HUMAN-ASSISTED/NOT RUN), failure injection, proptest properties, criterion benchmarks. Any regression blocks merge. ## 10 PHASES (exit = passing tests + evidence) 0 Audit (AUDIT.md, tool inventory, jarvis_build/, TRACE matrix). 1 Core: events/state/memory/planner/verify/gates/RPC; kill-and-resume proven. 2 Adapters + probe + math + ngspice + MATLAB/Octave; a real ngspice run verified against a known answer. 3 Digital flow. 4 PCB + SI/PI. 5 Analog/ASIC/FPGA as tools exist. 6 GUI operator. 7 Review agents, design-space optimization (record objective, variables, constraints, trade-offs), autopilot. 8 Hardening, benchmarks, docs. Each phase: coverage audit -> HANDOFF -> git tag. Don't advance with unexplained gaps. ## 11 REPORT PER WP WP / § covered / changed / tests run (command + result) / evidence / status / open issues / next. No claim without evidence. ## 12 KICKOFF (now, in order) 1 Create jarvis_build/; save spec + this directive. 2 Phase 0 read-only audit. 3 Populate TRACE_MATRIX. 4 Write Phase 1 WPs; build sec 4 runtime first. 5 Give me ONE batched question list (OS, tool versions, PDK, backends/API keys, existing code language/layout). 6 Proceed without waiting unless blocked. If unsure whether something is real, verified, or complete: it isn't. Say so and show the evidence.

Drag to resize

Response not available

Drag to resize

Response not available

Drag to resize
Drag to resize