MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 6141-6160 of 8828

👍0
Jsi nezávislý red-team evaluátor řídicího promptu LLM. Odpov...

👍0
Host shows: disk 92% full, memory 96% used incl 14 GiB cache...
👍0
123

👍0
Independent AI Reviewers — Defect Detection Test
A public comparison of four AI models on realistic software-plan and code-review tasks. It measures genuine defects found, false alarms, instruction following, safety reasoning and clarity.

👍0
recreate steam's game marketplace and give it spotify and ap...

👍0
اريد هيكل للذاكره الذهنيه لعلم البيانات انواعها و خصائصها و ...

👍0
Agent Harness

👍0
FastFetch Logo
👍0
Create a single-file index.html using Three.js with OrbitCon...

👍0
You are the Administrative Services Manager of the Administr...
👍0
Company financial year is from July to June, total 12 months...

👍0
Préface Jubilon
Mode d’emploi humoristique

👍0
bench mark
home bench

👍0
adwa
awda

👍0
Affordable Housing

👍0
test

👍0
B.M.13 — METHODOLOGY v28 VALIDATION: USER-ISSUED SHORT FILEN...

👍0
UNIVERZÁLNÍ ONE-SHOT BENCHMARK 3 — FULL-STACK ADVERSARIÁLNÍ ...

👍0
Complete the following Python function:
```python
from typi...

👍0
hcgv
fdhg