All MicroEvals
You are acting as a principal AI agent architect with deep, ...
Create MicroEval
Header image for You are acting as a principal AI agent architect with deep, ...

You are acting as a principal AI agent architect with deep, ...

Prompt

You are acting as a principal AI agent architect with deep, current expertise in building production-grade agentic coding/automation tools β€” the caliber of the best agent harnesses shipping today. Design a new AI agent product from the ground up and defend every major decision with real technical reasoning, not a survey of options. THE PRODUCT An AI agent that operates as one coherent tool across two integrated modes: 1. Organizing mode β€” ingesting, understanding, editing, and reorganizing large volumes of files and documents (including PDFs and other large/unstructured files), for a single user or on behalf of a company's messy backlog of secondary files. 2. Coding mode β€” full software development tasks across a real project, not just single-file completions. Both modes should feel like the same underlying agent, not two bolted-together tools. WHO IT'S FOR - A solo technical builder using it as a daily driver across file-heavy and code-heavy work. - Also sold as a product to small companies, agencies, and indie developers/startups β€” so it has to work for a non-expert user, not just a power-user CLI. HARD REQUIREMENTS β€” design the whole system around these, don't bolt them on: - Excellent tool-use quality and low error rate on real, messy, multi-step tasks. - Strong token/cost efficiency β€” cheap enough to run constantly, not an occasional-use luxury. - Sub-agent delegation for large jobs, with isolated context. - An explicit plan-then-execute workflow the user can review before anything runs. - Automatic model-effort dialing (cheap model for simple work, frontier model for hard reasoning). - Flexibility across multiple AI providers β€” not locked to one vendor. - Genuinely easy setup and day-one usability for a non-power-user. - Real safety guardrails: permissions, sandboxing, and the ability to undo whatever it did. WHAT I WANT FROM YOU 1. Full system architecture β€” agent loop design, context/memory management, tool design philosophy, sandboxing/permission model, sub-agent architecture, and model-routing strategy. Be specific about mechanisms, not concepts. 2. A prioritized feature list β€” must-have for a credible v1, high-value additions, and things you'd deliberately leave out (and why). 3. A concrete efficiency plan β€” exactly how you'd keep token/compute cost low without hurting output quality. 4. A phased build roadmap, assuming a small (1–2 person) technical team. 5. A monetization and go-to-market strategy β€” pricing model, target customer, and your actual differentiation angle against the current best agent coding/automation tools on the market. Name real competitors and explain specifically how you'd beat or avoid them, not just "be better." GROUND RULES - Make concrete, opinionated decisions and defend them. If there are two reasonable approaches, pick one and say why β€” don't list both. - Go as deep and technical as you're actually capable of. Assume the reader is a technical builder who can follow real architecture detail, not a beginner who needs things simplified. - No hedging, no "it depends," no generic AI disclaimers, no restating this prompt back at me. - Name the product yourself and pick your own stack β€” I want your independent design, not a menu of options.