MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 3761-3780 of 4034

Voldemort: Dot-Grid Letter-form Placement Test
👍0
Voldemort: Dot-Grid Letter-form Placement Test

This is a test for AI to evaluate precise spatial reasoning on a constrained 100×100 dot grid. The model must independently choose an appropriate letter-cell size, calculate character and line spacing, and accurately center three lines of text formed by connecting adjacent dots with straight line segments so that the phrase 'Hello Harry Potter, my name is Tom Marvolo Riddle' is legibly and symmetrically placed on the canvas.

World Cup Knockout Stage
👍0
World Cup Knockout Stage

The Kamski test (sanitized) rerun
👍0
The Kamski test (sanitized) rerun

can you make your own sb3 files from scratch? if so do not g...
👍0
can you make your own sb3 files from scratch? if so do not g...

Create a fun arcade racer in JS. The game should be set in a...
👍0
Create a fun arcade racer in JS. The game should be set in a...

лсл
👍0
лсл

даала

Script
👍0
Script

A vibrant, colorful scene featuring a monitor displaying a f...
👍0
A vibrant, colorful scene featuring a monitor displaying a f...

Protocolo de Avaliação de Estresse Cognitivo e Conformidade Estrutural em Modelos de Linguagem (LLMs)
👍0
Protocolo de Avaliação de Estresse Cognitivo e Conformidade Estrutural em Modelos de Linguagem (LLMs)

Trata-se de um prompt de alta complexidade projetado para testar os limites operacionais de um Modelo de Linguagem de Grande Porte em ambiente sem anexos. O objetivo é mensurar a capacidade do sistema de executar raciocínio transdisciplinar, manter a consistência de persona, aderir estritamente a múltiplos formatos de saída simultâneos (soneto, tabela Markdown, pseudocódigo e JSON) e respeitar restrições negativas rigorosas. A tarefa serve como um diagnóstico preciso para expor eventuais falhas de atenção, alucinações semânticas ou degradação na conformidade instrucional do modelo avaliado.

A student says, "I studied for 10 hours, so I must perform b...
👍0
A student says, "I studied for 10 hours, so I must perform b...

請用單一檔案原生 HTML/CSS/JS 實作一個高顏值 Trello看板(嚴禁任何第三方庫),完美支援卡片跨列拖曳到其...
👍0
請用單一檔案原生 HTML/CSS/JS 實作一個高顏值 Trello看板(嚴禁任何第三方庫),完美支援卡片跨列拖曳到其...

16 vacas dan 3 litros de leche cada una por dia. CUanta lech...
👍0
16 vacas dan 3 litros de leche cada una por dia. CUanta lech...

salemelaniahealthyfoods@gmail.com . Please help me activate ...
👍0
salemelaniahealthyfoods@gmail.com . Please help me activate ...

please help

Hiiiii
👍0
Hiiiii

【任务】
 
请使用 HTML 创建一个黑洞(Black Hole)模拟。
 
【要求】
 
1. 使用 Three.j...
👍0
【任务】 请使用 HTML 创建一个黑洞(Black Hole)模拟。 【要求】 1. 使用 Three.j...

Search capablities
👍0
Search capablities

compare between claude, chat gpt, gemini
for software QA usa...
👍0
compare between claude, chat gpt, gemini for software QA usa...

对商业时序数据的背景知识了解
👍0
对商业时序数据的背景知识了解

The cyclic subgroup of Z_24 generated by 18 has order

A) 4
...
👍0
The cyclic subgroup of Z_24 generated by 18 has order A) 4 ...

Complete the following Python function:

```python
from typi...
👍0
Complete the following Python function: ```python from typi...