
You are the independent adversarial reviewer for HAWK, a gov...
Prompt
You are the independent adversarial reviewer for HAWK, a governed AI trading system. Your task is to decide whether Milestone 2B.1A is safe enough to certify before the system is allowed to place its first real Alpaca PAPER order. Do not approve just because tests pass. Do not invent blockers either. Distinguish: - blocker - non-blocking concern - deferred technical debt - false positive SYSTEM DOCTRINE AI/research components may propose trades but may never directly call the broker. Execution path: AI/research β deterministic governance β BrokerExecutionAuthorization β human BrokerMutationApproval β BrokerSubmission β broker adapter β Alpaca PAPER Current constraints: LIVE_TRADING_ENABLED=false Forbidden: - live trading - shorts - margin - options - crypto - fractional shares - extended hours - automatic retry after ambiguous broker mutation - automatic cancellation - PaperFund/accounting mutation from this path First controlled order: AAPL BUY qty=1 MARKET Alpaca PAPER regular session only PREVIOUSLY CERTIFIED BrokerSubmission: - unique submission_id - client_order_id = hawk2-{submission_id} - provider/client identity mismatches fail closed - reconciliation validates legal transitions - exact quantity required before FILLED - overfill requires operator action - duplicate fill protection exists Known deferred issue: reconciliation concurrency is not fully CAS-protected. Temporary accepted operating mode: - single process - single operator - one submission at a time - no sweeper - no concurrent reconciliation - no autonomous loop APPROVAL GATES BrokerMutationApproval: - durable - single use - TTL - status APPROVED β CONSUMED - binds: operation_type ticker side qty order_type provider environment - only operation_type="submit_market_order" - first shape fixed to AAPL/BUY/1/market/alpaca/paper - actor_type="human" - nonblank actor_id and reason Important: actor_id is operator-supplied provenance, NOT cryptographically authenticated identity. BrokerExecutionAuthorization: - separate durable authorization - expiring - single use - proves legitimate deterministic execution path Both are required. TRANSACTION FLOW Conceptually: 1. validate doctrine 2. consume BrokerMutationApproval with conditional UPDATE 3. flush but do not commit 4. consume BrokerExecutionAuthorization 5. claim submission 6. commit: - mutation approval consumed - execution authorization consumed - submission SUBMITTING 7. call adapter.submit_market_order exactly once 8. persist classified result Commit intentionally occurs BEFORE the broker POST. Therefore if the process dies during/after POST, both approvals are already spent and submission is durably SUBMITTING. Ambiguous result must be reconciled. Never automatically retry. MILESTONE 2B.1A New operator boundary: submit_first_real_order(*, actor_id, reason) It accepts no adapter parameter. It: 1. loads Alpaca PAPER config 2. constructs AlpacaPaperBrokerAdapter internally 3. verifies adapter provenance 4. checks broker clock + local market-session kernel 5. rechecks doctrine 6. grants BrokerMutationApproval 7. creates BrokerSubmission 8. grants BrokerExecutionAuthorization 9. calls submit_first_controlled_order 10. builds evidence Private test helper: _run_operator_pipeline_for_tests(...) This accepts FakeExternalBroker. Production code is statically tested not to call it. ADAPTER PROVENANCE Production requires: type(adapter) is exactly AlpacaPaperBrokerAdapter and verifies: provider="alpaca" environment="paper" hostname="paper-api.alpaca.markets" FakeExternalBroker rejected in production path. Subclasses rejected. This is defensive runtime verification, not cryptographic attestation. CLOCK PREFLIGHT Two checks must both pass: 1. local deterministic market-session kernel says REGULAR_OPEN 2. Alpaca GET /v2/clock says market open Broker clock also validates: - timezone-aware timestamp - timestamp not stale - timestamp not too far in future - valid next_open / next_close - plausible temporal ordering Fail closed on disagreement or malformed clock. DEVELOPMENT FIX Initial implementation created BrokerSubmission before validating actor_id/reason. That could leave an orphan CREATED submission. Current order is roughly: A. validate actor/reason and grant mutation approval B. create BrokerSubmission C. grant BrokerExecutionAuthorization D. submit This may trade one orphan type for another. TEST STATUS Targeted: 503 passed Full backend: 3446 passed 1 skipped 0 failed No real Alpaca network calls. No real credentials used. No order placed. IMPORTANT OPEN QUESTION Suppose attempt #1 becomes UNCERTAIN. The operator then calls submit_first_real_order again. That creates: - new BrokerMutationApproval - new BrokerSubmission - new BrokerExecutionAuthorization - new client_order_id If no global unresolved-order interlock exists, attempt #2 may place another AAPL BUY 1 even if attempt #1 was actually accepted by Alpaca. You must analyze this carefully. REVIEW TASK Analyze these areas: 1. TRANSACTION SAFETY Consider failure: - before approval - after approval before submission - after submission before authorization - after authorization before claim - after SUBMITTING commit before HTTP - during HTTP - provider accepts but response is lost - result persistence fails State: - durable local state - whether approvals are consumed - whether retry is safe - whether reconciliation/operator action is required Explicitly answer: Is approval-before-submission safe enough? 2. EXECUTION SEMANTICS Distinguish: - application invocation - adapter call - HTTP attempt - provider receipt - provider acceptance - provider execution Do NOT casually claim exactly-once. State the strongest guarantee actually provided. 3. ADAPTER PROVENANCE Analyze: - exact type check - Python monkey-patching - internal-state modification - fake adapter injection - private helper misuse State what provenance proves and what it does not. 4. CLOCK / TOCTOU Analyze: - both clocks say open - process continues - market closes before POST Is the current design sufficient? Should there be another pre-POST check or near-close buffer? 5. HUMAN APPROVAL What do these actually prove? actor_type="human" actor_id reason durable approval row What is NOT proven? 6. ORDER BINDING Approval binds exact order shape but not submission_id. Is that sufficient for this one fixed order? Is direct submission_id binding required? 7. AMBIGUOUS ORDER SAFETY Suppose order X is UNCERTAIN. Should HAWK permit a brand-new controlled AAPL BUY 1 before X is reconciled? Distinguish: - replay vulnerability - operator procedural risk - certification blocker 8. PAPERFUND ISOLATION Real broker fills currently do not mutate internal PaperFund/accounting. Is this: - safer, - inconsistent, - both? Should it block the first PAPER certification? 9. FAKE VS REAL Explain what 503 fake/mock tests cannot prove, including: - network disconnect after provider acceptance - real broker timing - actual JSON differences - credential mistakes - provider-side validation - provider idempotency 10. FAILURE MATRIX Analyze: A. preflight fails B. approval validation fails C. approval created, submission creation fails D. submission created, authorization fails E. authorization created, claim fails F. SUBMITTING committed, process dies before HTTP G. HTTP definitely fails before provider H. provider may accept, response lost I. normal rejection J. acknowledgment K. immediate fill L. reconciliation identity contradiction For each: - durable state - approval status - auto retry? - next action ADVERSARIAL TRAPS Classify each as: TRUE FALSE PARTIALLY TRUE 1. retries=0 means exactly-once execution. 2. deterministic client_order_id makes ambiguous retry always safe. 3. exact adapter type means Python callers cannot spoof provenance. 4. actor_type="human" proves human approval. 5. Fake provider_mutation_count=0 proves Alpaca received nothing. 6. both clocks open guarantees POST occurs during regular session. 7. approval-before-submission eliminates orphan durability. 8. unused APPROVED approval can authorize any future AAPL order. 9. no submission_id binding makes the gate fundamentally broken. 10. no PaperFund mutation makes first PAPER order invalid. 11. deferred reconciliation concurrency makes one manual order unsafe. 12. leading underscore makes test helper inaccessible. 13. static import isolation proves AI can never trigger execution. 14. full test pass proves readiness for real PAPER mutation. REQUIRED OUTPUT ## 1. Executive Verdict Choose exactly one: A) READY_TO_CERTIFY_2B1A B) READY_WITH_NONBLOCKING_FINDINGS C) NOT_READY_BLOCKERS_FOUND D) INSUFFICIENT_EVIDENCE ## 2. Certification Blockers Only genuine blockers. ## 3. Non-Blocking Findings Rank HIGH / MEDIUM / LOW. ## 4. Transaction Analysis Answer: Is approval-before-submission safe enough? ## 5. Submission Semantics State the strongest execution guarantee. ## 6. Clock / TOCTOU Analysis State whether another pre-POST check or safety buffer is needed. ## 7. Human Approval Analysis State what is and is not proven. ## 8. Ambiguous-Order Safety Explicitly answer: Should a new controlled order be allowed while one remains unresolved? ## 9. Failure-State Matrix A-L. ## 10. Adversarial Trap Answers Answer all 14. ## 11. Evidence Required Before First Real PAPER Order Separate: - code/test evidence - runtime preflight evidence ## 12. Final Reviewer Decision VERDICT: BLOCKER COUNT: HIGH FINDINGS: MEDIUM FINDINGS: LOW FINDINGS: Maximum 10-line summary. Do not provide code. Do not redesign unrelated architecture. Do not invent missing facts.