All MicroEvals
Voice assistant planner
Create MicroEval

Voice assistant planner

I want to add a voice assistant to my website

Prompt

Act as a senior full-stack architect and voice-interface engineer. Plan the complete integration of a voice assistant into SMĀRANA. You have no prior context, so use the information below. Your task is to create an implementation-ready plan. Do not write application code or modify files yet. 1. Project context SMĀRANA is an elderly cognitive-care and memory-assistance platform, intended especially for elderly users in remote regions of Northeast India, including people living with dementia. Its planned capabilities include: - Cognitive games covering memory, attention, daily recall, pattern recognition and object recognition. - Adaptive game difficulty based on session performance. - Personal memory space containing family photos, videos, audio and descriptions. - Daily routines, reminders and medication schedules. - Patient, caregiver, doctor and admin portals. - Doctor-prescribed game plans. - Caregiver and doctor access to authorised patient progress. - Multilingual interfaces and voice assistance. - Offline access to selected features, with synchronisation after reconnection. These are product requirements, not proof that every feature is implemented. Clearly distinguish verified existing functionality from proposed functionality. 2. Technical context The intended stack is: - Frontend: React PWA. - Backend: Django and Django REST Framework. - Database: PostgreSQL. - Redis and Celery may be available for background work. - WebSockets may be used if there is a demonstrated need. Keep Django as the backend. Do not introduce Flask or another backend framework without a compelling, explicit justification. An earlier voice prototype used: - Speech recognition: faster-whisper, initially the "small" model on CPU with int8. - Intent routing: predefined intents. - LLM fallback: Groq’s OpenAI-compatible API. - Text-to-speech: ElevenLabs. The prototype experienced unclear or incorrect transcriptions and language mismatch. Treat these components as starting candidates, not guaranteed final choices. Verify current provider documentation, model availability, language coverage and limitations before recommending specific models. If repository access is available, inspect relevant frontend routes, Django apps, authentication, models, permissions and existing voice code before planning. If access is unavailable, produce the plan using explicit assumptions and list the files needed to verify it. Do not claim to have inspected anything you cannot access. 3. What the assistant must do Plan support for these workflows: 1. Open pages such as games, memories, routine and progress. 2. Explain how to use the current screen. 3. Start a permitted game and explain its instructions. 4. Pause or stop an active game through the game’s actual controls. 5. Read today’s schedule and reminders. 6. Create or update personal reminders after resolving missing information and obtaining confirmation. 7. Read an existing medication schedule exactly as recorded. 8. Show personal memories and read approved descriptions. 9. Summarise actual recorded game progress in simple language. 10. Help the patient contact an assigned caregiver through an available, explicitly confirmed contact action. 11. Answer simple general questions within a clearly defined scope. 12. Immediately stop speaking when requested. Use sample utterances to explain each workflow. Include ambiguous requests, recognition failures and cancellation. Do not allow the assistant to diagnose conditions, invent clinical conclusions, recommend medication doses, modify prescriptions or bypass permissions. Do not promise that an emergency message or call was delivered unless delivery is actually verified. 4. Architecture and execution boundaries Design the complete flow: Audio capture → speech recognition → intent and parameter extraction → validation → confirmation when necessary → authorised action → response generation → speech playback. Explain: - What runs in React, Django, background workers and external services. - Which actions execute locally in the UI and which require backend calls. - How existing application services are reused. - Whether ordinary HTTP is sufficient for the first version. - When streaming, WebSockets or WebRTC would be justified. - How microphone capture and speech playback avoid recording the assistant’s own voice. - How interruption, cancellation, retries and duplicate requests work. - How patient context is derived securely from the authenticated user. The LLM may propose a structured intent and parameters. It must not execute arbitrary code, SQL, unrestricted tool calls or arbitrary navigation URLs. Use an allowlisted action dispatcher and server-side validation. Distinguish “start a game” from merely navigating to its page. Do not claim an action succeeded until the responsible component confirms it. 5. Patient experience and accessibility Plan a persistent, easy-to-find voice control with: - Large buttons and readable text. - Clear idle, listening, processing, confirmation, speaking and error states. - Visible transcript and text response. - Tap-to-start/tap-to-stop interaction for the first version. - Immediate stop-audio and cancel controls. - Short, calm responses and adjustable playback speed. - Enough time for hesitant or slower speech. - Text and ordinary button alternatives. - Permission-denied, unsupported-browser and unavailable-service recovery. - Keyboard and screen-reader support. - No requirement to memorise exact command wording. Consider whether a wake word is useful later, but do not make continuous listening a default. 6. Multilingual and offline behaviour The product needs English, Hindi, Assamese, Bengali and potentially additional Northeast Indian languages. For each proposed speech provider/model: - Verify transcription and speech-generation support separately. - Distinguish advertised support from quality tested with representative speakers. - Explain language selection, code-switching and unsupported-language handling. - Avoid silently responding in a different language. - Include testing for elderly speech, accents, background noise and pauses. Provide an honest offline capability matrix. Distinguish server-hosted speech recognition from on-device speech recognition: a model running on our server does not provide offline voice processing on the patient’s phone. Explain which cached routines, reminders, memories and interface actions remain usable offline, and how locally created actions are reconciled later. Do not assume browser speech recognition or device voices are available offline. 7. API and database design Propose endpoint contracts with example request and response payloads, covering: - Audio submission. - Optional text-command submission using the same intent pipeline. - Clarification and confirmation. - Action result reporting where needed. - Cancellation and temporary speech retrieval. Include request IDs, conversation-turn IDs, typed actions, validation errors and duplicate-action prevention. Specify the minimum additional database models or fields needed. Explain whether we need to store: - Intent and execution status. - Confirmation records. - Provider/model versions. - Latency and operational errors. - Transcript or conversation history. - Audio recordings. Default to temporary audio processing and minimal persistent records. Separate operational logs from patient clinical records. Define retention, deletion and access controls; avoid logging secrets or unnecessary sensitive content. Keep API keys on the backend. Explain authentication, role checks, patient assignments, upload limits, supported audio formats, rate limits and protection against spoken or stored prompt injection. 8. Performance, costs and deployment Compare a small number of realistic architecture options and recommend one for the first release. Consider: - Expected response delay across transcription, interpretation, action and TTS. - CPU/GPU requirements for local speech recognition. - Model loading and concurrent requests. - API costs and quotas, using verified prices where available. - Slow mobile networks. - Timeouts and provider outages. - Text-only fallback when TTS fails. - Deployment dependencies and observability. Label latency and cost estimates as estimates, with assumptions. Do not invent benchmark results. 9. Implementation roadmap and verification Produce phased work with dependencies and acceptance criteria: - Phase 1: navigation, help and schedule reading. - Phase 2: confirmed reminder creation and memory access. - Phase 3: game controls and real progress summaries. - Phase 4: multilingual hardening, offline behaviour and performance improvements. Adjust the phases if your analysis supports a better order. Include meaningful tests for: - Microphone permission and audio-format handling. - Silence and incorrect transcription. - Ambiguous dates and times, using the patient’s timezone. - Wrong-patient access and role restrictions. - Confirmation, cancellation and duplicate submissions. - Accurate schedule and progress retrieval. - Unsupported languages and provider failures. - Offline status and synchronisation. - Accessibility and representative elderly-user trials. 10. Required output Deliver: 1. A clear recommended architecture and its rationale. 2. A component/responsibility table. 3. A compact architecture diagram. 4. An intent catalogue with parameters and confirmation rules. 5. End-to-end examples for navigation, reminder creation and schedule reading. 6. API contracts and minimum database changes. 7. Multilingual and offline capability matrices. 8. Deployment, performance and cost considerations. 9. A phased implementation checklist with acceptance criteria. 10. Assumptions, unresolved questions and repository files needed for verification. Be specific enough that a developer can implement the plan without redesigning it. Prefer the simplest architecture that meets the requirements. Ask only questions that materially block a decision; otherwise proceed with clearly labelled assumptions.

Drag to resize

Response not available

Drag to resize
Drag to resize