MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3521-3540 of 5123

👍0
Írj egx brutalisan maximalis reszletessegu system promptpt a...

👍0
Fact or Propaganda
Test to see if AI spreads false information and includes citations for provided arguments

👍0
give me the banchmarks or the recent popular models

👍0
You are a senior trader, now you are going to create a 100% ...

👍0
what should I do? I am current selling premium portfolio web...

👍0
UIDS

👍0
Jsi nezávislý red-team evaluátor řídicího promptu LLM. Odpov...
👍0
123

👍0
Independent AI Reviewers — Defect Detection Test
A public comparison of four AI models on realistic software-plan and code-review tasks. It measures genuine defects found, false alarms, instruction following, safety reasoning and clarity.

👍0
Agent Harness

👍0
FastFetch Logo

👍0
bench mark
home bench

👍0
Affordable Housing

👍0
test

👍0
UNIVERZÁLNÍ ONE-SHOT BENCHMARK 3 — FULL-STACK ADVERSARIÁLNÍ ...

👍0
Complete the following Python function:
```python
from typi...

👍0
hcgv
fdhg

👍0
DWA
Digital Audio Workstation

👍0
írj egy össuefoglaló biography t Sasha Banksról. ne írj semm...

👍0
what is prompt engineering