MicroEvals
Public evaluations
Showing 6101-6120 of 10517

![This is input from if node:
"subject": "[ZUNANJI] Vaš paket ...](/_next/image?url=https%3A%2F%2Fartificialanalysiscdn.com%2Fmicro-evals%2F476517bfae524609a3e6e22e9f81e58d.jpg&w=3840&q=75)

Question for the currently not existing metric of GDP growth adjusted for the (government) deficit in the same period, akin to how there is inflation-adjusted GDP growth figures (nominal vs real). While seeming simple, pretty much all AIs instead automatically snap to the similar sounding but completely unrelated existing concept from their training data, of how much GDP growth does each percentage point of Deficit buy - resuting in a 100% failure of their task.








I dunno just trying some stuff.

Production-oriented evaluation for StoreAgent's customer-facing WhatsApp LLM. Tests intent/routing accuracy, entity extraction, grounded commerce reasoning, multi-turn context, Arabic/English/code-switching quality, conversational naturalness, clarification behavior, hallucination resistance, and safe proposed actions. Prices, stock, payment state, order state and irreversible actions remain deterministic backend truth and must never be invented or independently changed by the model.
