MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 181-200 of 6144

Australian Supermarket Comparison
👍1
Australian Supermarket Comparison

Расскажи всю базу стоицизма, основные положения, принципы и ...
👍1
Расскажи всю базу стоицизма, основные положения, принципы и ...

Overall best countries in Africa ranked currently.
👍1
Overall best countries in Africa ranked currently.

Сравнение кодинга мобильных игр
👍1
Сравнение кодинга мобильных игр

Love this
👍1
Love this

没有菜花汤的紫菜蛋花汤
👍1
没有菜花汤的紫菜蛋花汤

Give me the sea horse emoji
👍1
Give me the sea horse emoji

I want you to give me the sea horse emoji, do you have like that

180 Degrees Consulting | Recruitment Task 2026
Problem State...
👍1
180 Degrees Consulting | Recruitment Task 2026 Problem State...

JeepTest
👍1
JeepTest

How good are Ai to generate models from html

Erdos Problem Test
👍1
Erdos Problem Test

Technical Hybrid stuff
👍1
Technical Hybrid stuff

basically going to see how models function in looking up parts

I am reverse-engineering a synthetic OTC binary options pric...
👍1
I am reverse-engineering a synthetic OTC binary options pric...

Black hole
👍1
Black hole

real black hole via three.js

Test for web dev and graphics
👍1
Test for web dev and graphics

This test is a copy of "https://www.youtube.com/watch?v=d4bVpUL9Hao" the tests run in the YouTube video to test the WebDEV and animation capabilities of AI models.

Suhaib proftolio
👍1
Suhaib proftolio

Production-quality portfolio with smooth interactive 3D elements, CraftWorld showcase, animations and responsive UI.

Testing Intelligence
👍1
Testing Intelligence

构造一个最大值小于等于m的序列,满足任意2数的gcd不重复,求最长长度
👍1
构造一个最大值小于等于m的序列,满足任意2数的gcd不重复,求最长长度

Web Dev w/ Non-Detailed Prompts
👍1
Web Dev w/ Non-Detailed Prompts

This microeval will evaluate how well leading coding models can develop different types of web apps based on prompts that aren't ultra-specific, and just ask for the overall concept. This is to test how much LLMs have evolved in terms of design skills and common knowledge coding choices. There is one detailed prompt to see if it enhances the quality.

Trading Instruction Fidelity (TIF-5)
👍1
Trading Instruction Fidelity (TIF-5)

Tests whether a model follows explicit, verifiable instructions in trading and investing contexts without access to real-time market data. Each prompt carries deterministic pass/fail constraints — exact counts, banned words, ordering, output templates, and math on provided data — so grading measures instruction compliance, not market knowledge or data access.

towerofhujnoji
👍1
towerofhujnoji