MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 221-240 of 8352

Technical Hybrid stuff
👍1
Technical Hybrid stuff

basically going to see how models function in looking up parts

I am reverse-engineering a synthetic OTC binary options pric...
👍1
I am reverse-engineering a synthetic OTC binary options pric...

Black hole
👍1
Black hole

real black hole via three.js

👍1
You are an administrative operations lead in a government de...

Test for web dev and graphics
👍1
Test for web dev and graphics

This test is a copy of "https://www.youtube.com/watch?v=d4bVpUL9Hao" the tests run in the YouTube video to test the WebDEV and animation capabilities of AI models.

Suhaib proftolio
👍1
Suhaib proftolio

Production-quality portfolio with smooth interactive 3D elements, CraftWorld showcase, animations and responsive UI.

Fable 5.1   vs  gpt 6.1 astra
👍1
Fable 5.1 vs gpt 6.1 astra

Testing Intelligence
👍1
Testing Intelligence

构造一个最大值小于等于m的序列,满足任意2数的gcd不重复,求最长长度
👍1
构造一个最大值小于等于m的序列,满足任意2数的gcd不重复,求最长长度

👍1
Overall best countries in Africa ranked currently

👍1
Eval 3

👍1
fishing simulator 3d

Sme math question
👍1
Sme math question

👍1
THe hello promp

Web Dev w/ Non-Detailed Prompts
👍1
Web Dev w/ Non-Detailed Prompts

This microeval will evaluate how well leading coding models can develop different types of web apps based on prompts that aren't ultra-specific, and just ask for the overall concept. This is to test how much LLMs have evolved in terms of design skills and common knowledge coding choices. There is one detailed prompt to see if it enhances the quality.

Trading Instruction Fidelity (TIF-5)
👍1
Trading Instruction Fidelity (TIF-5)

Tests whether a model follows explicit, verifiable instructions in trading and investing contexts without access to real-time market data. Each prompt carries deterministic pass/fail constraints — exact counts, banned words, ordering, output templates, and math on provided data — so grading measures instruction compliance, not market knowledge or data access.

towerofhujnoji
👍1
towerofhujnoji

test
👍1
test

👍1
p5.js physics_New

fusion generator
👍1
fusion generator

Generator