MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 7541-7560 of 7809

Kırkkilit bitkisini çayının günün hangi saatinde tüketmek da...
👍0
Kırkkilit bitkisini çayının günün hangi saatinde tüketmek da...

Tìm nền tảng tương tự arena.ai và zenmux.ai
👍0
Tìm nền tảng tương tự arena.ai và zenmux.ai

deliver profound upgrades to enhance novelty, usefulness to ...
👍0
deliver profound upgrades to enhance novelty, usefulness to ...

saq
👍0
saq

You are a creative coder and an expert in computer graphics,...
👍0
You are a creative coder and an expert in computer graphics,...

模拟我大师赛
👍0
模拟我大师赛

You are a Senior Business Development Strategist with expert...
👍0
You are a Senior Business Development Strategist with expert...

👍0
111

You are the administrative services manager responsible for ...
👍0
You are the administrative services manager responsible for ...

Как да си отворим фирма ЕООД
👍0
Как да си отворим фирма ЕООД

In 2015, he said, a study completed in cooperation with the ...
👍0
In 2015, he said, a study completed in cooperation with the ...

Выступи в роли прагматичного, жесткого карьерного стратега и...
👍0
Выступи в роли прагматичного, жесткого карьерного стратега и...

OscarGolf-3529
👍0
OscarGolf-3529

Generate a fully synthetic, production-ready pre-training da...
👍0
Generate a fully synthetic, production-ready pre-training da...

In 2015, he said, a study completed in cooperation with the ...
👍0
In 2015, he said, a study completed in cooperation with the ...

Runner system-prompt efficiency — improvement task
👍0
Runner system-prompt efficiency — improvement task

How good models are at improving existing systems.

voy a crear una web para una empresa que vende bebidas refre...
👍0
voy a crear una web para una empresa que vende bebidas refre...

female wrestler
👍0
female wrestler

guess fav female wrestle by birth chart. correct answer: Sasha Banks

/**
 * core-ui.css — Techsillica
 *
 * CSS companion to core...
👍0
/** * core-ui.css — Techsillica * * CSS companion to core...

👍0
Aegis-7: Extreme Reasoning, Embedded Systems, and Linguistic Constraint Benchmark

An advanced stress test designed to push frontier Large Language Models (LLMs) to their operational limits under strict, multi-layered constraints. The prompt forces the AI to resolve a critical orbital recovery dilemma by calculating weighted utility matrices step-by-step, writing memory-safe embedded Rust code (#![no_std]), conducting ethical analysis under a severe linguistic constraint (an 80+ word paragraph omitting the letter "e"), and enforcing strict JSON output adherence without conversational filler.