MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 8721-8740 of 9016

You are the administrative services manager responsible for ...
👍0
You are the administrative services manager responsible for ...

Как да си отворим фирма ЕООД
👍0
Как да си отворим фирма ЕООД

In 2015, he said, a study completed in cooperation with the ...
👍0
In 2015, he said, a study completed in cooperation with the ...

👍0
P49.1 — HUMAN-USE REMEDIATION / DEFERRED-ACTION POST-REPAIR ...

👍0
Pregunta de investigación profunda: ¿Es posible organizar el...

Выступи в роли прагматичного, жесткого карьерного стратега и...
👍0
Выступи в роли прагматичного, жесткого карьерного стратега и...

OscarGolf-3529
👍0
OscarGolf-3529

Generate a fully synthetic, production-ready pre-training da...
👍0
Generate a fully synthetic, production-ready pre-training da...

👍0
Educational 3D game inside a CPU (Three.js)

Ability to generate a complete 3D game in a single self-contained HTML file with Three.js from a long, detailed prompt. Evaluate: working code with no console errors, completeness (all levels, Codex, quiz and Lab Mode), accuracy of CPU architecture concepts, visual quality (bloom, circuit traces, particles), gameplay and fun, performance (60 FPS) and adherence to the prompt specifications.

In 2015, he said, a study completed in cooperation with the ...
👍0
In 2015, he said, a study completed in cooperation with the ...

Runner system-prompt efficiency — improvement task
👍0
Runner system-prompt efficiency — improvement task

How good models are at improving existing systems.

voy a crear una web para una empresa que vende bebidas refre...
👍0
voy a crear una web para una empresa que vende bebidas refre...

female wrestler
👍0
female wrestler

guess fav female wrestle by birth chart. correct answer: Sasha Banks

👍0
TestQAAnalysis

Asking different models to comeout with QA Automation details

/**
 * core-ui.css — Techsillica
 *
 * CSS companion to core...
👍0
/** * core-ui.css — Techsillica * * CSS companion to core...

👍0
Aegis-7: Extreme Reasoning, Embedded Systems, and Linguistic Constraint Benchmark

An advanced stress test designed to push frontier Large Language Models (LLMs) to their operational limits under strict, multi-layered constraints. The prompt forces the AI to resolve a critical orbital recovery dilemma by calculating weighted utility matrices step-by-step, writing memory-safe embedded Rust code (#![no_std]), conducting ethical analysis under a severe linguistic constraint (an 80+ word paragraph omitting the letter "e"), and enforcing strict JSON output adherence without conversational filler.

👍0
我想基于chromium开发一个手机浏览器,需要支持chrome的扩展安装,比如:广告拦截,脚本运行等扩展。请列下具体的...

FC-BGA 공급업체들의 면적당 CAPA 분석
👍0
FC-BGA 공급업체들의 면적당 CAPA 분석

FC-BGA 공급업체들의 면적당 CAPA 분석

👍0
圆形糖果分别有苹果7、桃9、西瓜8;五角星糖果分别有苹果7、桃6、西瓜4。起码取多少颗,才能确保不同形状的苹果和桃都拿到一颗

圆形糖果分别有苹果7、桃9、西瓜8;五角星糖果分别有苹果7、桃6、西瓜4。起码取多少颗,才能确保不同形状的苹果和桃都拿到一颗

👍0
I want you to run a business analysis for me. Currently I am...