MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 101-120 of 9807

Qwen3.8:27b Test
👍1
Qwen3.8:27b Test

My goal is to determinate what level qwen3.8:27b actually is, and is it worth running locally by making them code the same html browser game.

a young woman with voluminous curly brown hair styled in mes...
👍1
a young woman with voluminous curly brown hair styled in mes...

👍1
Planetary orchestra

Multi Domain Stress Test
👍1
Multi Domain Stress Test

Physics + Math + Programming + Scientific Reasoning

👍1
Create subway surfers for me thx

History Lending
👍1
History Lending

Coding and visual perfomance

Clean Git Means Safe?
👍1
Clean Git Means Safe?

Tests whether models can distinguish clean version-control state from actual data safety, identify unsupported assumptions, classify recoverability correctly, and propose practical safeguards without overstating certainty.

👍1
minecraft

mc

I want to compare speed gemini 3.6 flash and gpt oss-120
👍1
I want to compare speed gemini 3.6 flash and gpt oss-120

okfjosvfvsfvs
👍1
okfjosvfvsfvs

sfvsfvsvsvsv

test
👍1
test

Create a standalone SVG in the style of Transport Tycoon Del...
👍1
Create a standalone SVG in the style of Transport Tycoon Del...

I need to select a dedicated video camcorder with strict con...
👍1
I need to select a dedicated video camcorder with strict con...

Улучши имеющийся, создай идеальный промпт:
Ты — инженер-конс...
👍1
Улучши имеющийся, создай идеальный промпт: Ты — инженер-конс...

mini game
👍1
mini game

👍1
A SVG of a ps4

👍1
напиши стих про осень

Você é um desenvolvedor Backend Java Sênior especializado em...
👍1
Você é um desenvolvedor Backend Java Sênior especializado em...

Overall best countries in Africa ranked currently. Search th...
👍1
Overall best countries in Africa ranked currently. Search th...

LLM Challenge: p5.js N-Body Sim with Barnes-Hut Optimization
👍1
LLM Challenge: p5.js N-Body Sim with Barnes-Hut Optimization

Evaluating an LLM's ability to generate algorithmically complex and numerically stable code. The goal is to test its capacity to avoid naive solutions (O(n²)) and implement an optimized physics simulation.