MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3721-3740 of 6437

👍0
Jsi expertní vědecký, medicínský a Evidence-Based Medicine a...

👍0
aoue

👍0
Consider yourself as a physics expert. Imagine that there is...

👍0
P21.2 — EVALUATOR / EC-COMMITMENT / REPLICATED-TRIAL / SHADO...

👍0
Complete the following Python function:
```python
from typi...

👍0
Kansızlığa karşı Demir eksikliğine karşı elma ve siyah üzüm ...
👍0
Full Pipeline eval
I dunno just trying some stuff.

👍0
StoreAgent Customer-Facing WhatsApp LLM Eval v1
Production-oriented evaluation for StoreAgent's customer-facing WhatsApp LLM. Tests intent/routing accuracy, entity extraction, grounded commerce reasoning, multi-turn context, Arabic/English/code-switching quality, conversational naturalness, clarification behavior, hallucination resistance, and safe proposed actions. Prices, stock, payment state, order state and irreversible actions remain deterministic backend truth and must never be invented or independently changed by the model.

👍0
Make chrome

👍0
C++ Reflection

👍0
Execute esta auditoria respeitando todas as regras abaixo.
...

👍0
الان این در نهایت کدوم ای پی رو انتخاب کرد و پورتش چیه؟
21:0...

👍0
Mathe

👍0
Камера резко наезжает на лицо девушки (соотнеси с внешностью...

👍0
Who received the IEEE Frank Rosenblatt Award in 2010?

👍0
we are a furniture brand who have come across an organizatio...
👍0
角色设定:
你是一位拥有理论物理博士学位的高级图形学工程师,精通广义相对论(特别是旋转带电荷黑洞的Kerr-Newman...

👍0
write a synthesis to present how a llm works

👍0
Jsi nezávislý red-team evaluátor řídicího promptu LLM. Odpov...

👍0
create a very very very very very very advance modren and am...