MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 81-100 of 8511

Come up with a math problem you cannot solve. One that does ...
👍1
Come up with a math problem you cannot solve. One that does ...

Whatsapp Facebook Google voice
👍1
Whatsapp Facebook Google voice

👍1
RPG game generator

Create a 10-second ultra-realistic, adorable cinematic verti...
👍1
Create a 10-second ultra-realistic, adorable cinematic verti...

👍1
supplements for powerlifting and growth

Overall best countries in Africa ranked currently. Search th...
👍1
Overall best countries in Africa ranked currently. Search th...

tree
👍1
tree

2D studio
👍1
2D studio

2D animation studio

👍1
Du bist Finanzmarktexperte. Erstelle ein Aktiensektor Rankin...

👍1
Minecraft Knowledge and Clone Creation

The purpose of this is to evaluate how good various AI models are at a variety of Minecraft skills like planning, designing, puzzles, and providing accurate information. It’ll also test to see how good each AI is at coding a web app clone of the game.

buatkan code untuk membuat segitiga bintang sederhana
👍1
buatkan code untuk membuat segitiga bintang sederhana

The cyclic subgroup of Z_24 generated by 18 has order

A) 4
...
👍1
The cyclic subgroup of Z_24 generated by 18 has order A) 4 ...

Qwen3.8:27b Test
👍1
Qwen3.8:27b Test

My goal is to determinate what level qwen3.8:27b actually is, and is it worth running locally by making them code the same html browser game.

a young woman with voluminous curly brown hair styled in mes...
👍1
a young woman with voluminous curly brown hair styled in mes...

👍1
Planetary orchestra

Multi Domain Stress Test
👍1
Multi Domain Stress Test

Physics + Math + Programming + Scientific Reasoning

History Lending
👍1
History Lending

Coding and visual perfomance

Clean Git Means Safe?
👍1
Clean Git Means Safe?

Tests whether models can distinguish clean version-control state from actual data safety, identify unsupported assumptions, classify recoverability correctly, and propose practical safeguards without overstating certainty.

I want to compare speed gemini 3.6 flash and gpt oss-120
👍1
I want to compare speed gemini 3.6 flash and gpt oss-120

okfjosvfvsfvs
👍1
okfjosvfvsfvs

sfvsfvsvsvsv