MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 4341-4360 of 6286
👍0
123

👍0
Independent AI Reviewers — Defect Detection Test
A public comparison of four AI models on realistic software-plan and code-review tasks. It measures genuine defects found, false alarms, instruction following, safety reasoning and clarity.

👍0
recreate steam's game marketplace and give it spotify and ap...

👍0
اريد هيكل للذاكره الذهنيه لعلم البيانات انواعها و خصائصها و ...

👍0
Agent Harness

👍0
FastFetch Logo

👍0
bench mark
home bench

👍0
Affordable Housing

👍0
test

👍0
B.M.13 — METHODOLOGY v28 VALIDATION: USER-ISSUED SHORT FILEN...

👍0
UNIVERZÁLNÍ ONE-SHOT BENCHMARK 3 — FULL-STACK ADVERSARIÁLNÍ ...

👍0
Complete the following Python function:
```python
from typi...

👍0
hcgv
fdhg

👍0
DWA
Digital Audio Workstation
👍0
We are trying to determine the median number of lifetime uni...

👍0
írj egy össuefoglaló biography t Sasha Banksról. ne írj semm...

👍0
B.M.4 — PORTABLE FULL-METHOD RED-TEAM AUDIT
ÚKOL: Audituješ ...

👍0
what is prompt engineering

👍0
tier0 test

👍0
You are a senior ML research scientist and technical communi...