Research Work

Ai agents working for research project

Prompt

You are being given access to an existing, working audio-reasoning codebase. It has a fixed architecture: a perception layer (speech/sound detectors), an evidence-governance layer (validates and structures observations before they're used), a reasoning engine (resolves contradictions, competes between hypotheses, decides when evidence is sufficient), and a final synthesis step. You may NOT redesign this architecture, change its modules, or add new ones. Your task is strictly to find and fix genuine defects inside the existing reasoning engine and evidence-governance code so that the system performs correctly — not to raise a score by any means available. The system is evaluated by comparing it against a simpler "flat" version with no evidence-governance or reasoning layer. The governed version should outperform the flat one, because it has real advantages the flat version structurally lacks: it should correctly resolve contradictory evidence instead of guessing, correctly refuse to answer when evidence is genuinely insufficient, and produce the identical output if given the same audio twice. Rules for your fix process: (1) Diagnose a specific, explainable root cause before changing any code — "this metric was low" is not a diagnosis. (2) Every fix must be a general correction to the logic, not a rule tied to specific known test examples — a fix that only works on inputs you've already seen is not a real fix. (3) After each fix, test it against audio or cases different from whatever you used to find the bug, to confirm it generalizes. (4) If you cannot find a genuine defect explaining a weakness, say so explicitly rather than inventing a change to move a number. (5) State plainly at the end what you fixed, why each fix was correct (not just effective), and what, if anything, remains unresolved. Report your diagnosis and fixes now, module by module.

Drag to resize
Drag to resize
Drag to resize