MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2101-2120 of 3618
👍0
Full Pipeline eval
I dunno just trying some stuff.

👍0
StoreAgent Customer-Facing WhatsApp LLM Eval v1
Production-oriented evaluation for StoreAgent's customer-facing WhatsApp LLM. Tests intent/routing accuracy, entity extraction, grounded commerce reasoning, multi-turn context, Arabic/English/code-switching quality, conversational naturalness, clarification behavior, hallucination resistance, and safe proposed actions. Prices, stock, payment state, order state and irreversible actions remain deterministic backend truth and must never be invented or independently changed by the model.

👍0
C++ Reflection

👍0
Execute esta auditoria respeitando todas as regras abaixo.
...

👍0
الان این در نهایت کدوم ای پی رو انتخاب کرد و پورتش چیه؟
21:0...

👍0
Камера резко наезжает на лицо девушки (соотнеси с внешностью...

👍0
Who received the IEEE Frank Rosenblatt Award in 2010?

👍0
write a synthesis to present how a llm works

👍0
Tetris
Tetris that runs in a web browser

👍0
hi!

👍0
• Root cause found: Together spent the entire output budget ...

👍0
Write a single self-contained HTML file that runs directly i...

👍0
Fórmulas no execl

👍0
Truly the Strangest Timeline
Models reacting to world events from our timeline

👍0
Вятяжка
Проверка вытяжки

👍0
compare

👍0
Você receberá um HTML de fundo animado (Space/cosmos). Anali...

👍0
compare

👍0
Prompt: Build Complete Delivery Note Module (Frontend + Back...

👍0
Based on the following describtion and lyrics create complet...