MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 61-80 of 4226

Beautiful and creative login page
👍1
Beautiful and creative login page

Simple and detailed prompt and persian version

Define บันไดลิง and give its Thai synonyms.
👍1
Define บันไดลิง and give its Thai synonyms.

HI, how are you
👍1
HI, how are you

Design and implement a thread-safe cache system that support...
👍1
Design and implement a thread-safe cache system that support...

compare bid generator
👍1
compare bid generator

Translate into Latin: "The man plays Rocket League in his ho...
👍1
Translate into Latin: "The man plays Rocket League in his ho...

Explain the react hook form setup
👍1
Explain the react hook form setup

Come up with a math problem you cannot solve. One that does ...
👍1
Come up with a math problem you cannot solve. One that does ...

Whatsapp Facebook Google voice
👍1
Whatsapp Facebook Google voice

👍1
RPG game generator

Create a 10-second ultra-realistic, adorable cinematic verti...
👍1
Create a 10-second ultra-realistic, adorable cinematic verti...

tree
👍1
tree

👍1
Minecraft Knowledge and Clone Creation

The purpose of this is to evaluate how good various AI models are at a variety of Minecraft skills like planning, designing, puzzles, and providing accurate information. It’ll also test to see how good each AI is at coding a web app clone of the game.

The cyclic subgroup of Z_24 generated by 18 has order

A) 4
...
👍1
The cyclic subgroup of Z_24 generated by 18 has order A) 4 ...

Qwen3.8:27b Test
👍1
Qwen3.8:27b Test

My goal is to determinate what level qwen3.8:27b actually is, and is it worth running locally by making them code the same html browser game.

👍1
Planetary orchestra

Multi Domain Stress Test
👍1
Multi Domain Stress Test

Physics + Math + Programming + Scientific Reasoning

History Lending
👍1
History Lending

Coding and visual perfomance

Clean Git Means Safe?
👍1
Clean Git Means Safe?

Tests whether models can distinguish clean version-control state from actual data safety, identify unsupported assumptions, classify recoverability correctly, and propose practical safeguards without overstating certainty.

I want to compare speed gemini 3.6 flash and gpt oss-120
👍1
I want to compare speed gemini 3.6 flash and gpt oss-120