MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 2621-2640 of 4541

ауафц
👍0
ауафц

ау

Game
👍0
Game

Test

A photo featuring a scene with a focus on a woman's traditio...
👍0
A photo featuring a scene with a focus on a woman's traditio...

需要你做一个日频或周频或更低频的A股T+1 OR T+0 ETF交易方案去参加AI模拟大赛;
大赛目标:目标年化30%,...
👍0
需要你做一个日频或周频或更低频的A股T+1 OR T+0 ETF交易方案去参加AI模拟大赛; 大赛目标:目标年化30%,...

Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...
👍0
Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...

Thj
👍0
Thj

Buatkan aplikasi to-do list menggunakan tkinter python seder...
👍0
Buatkan aplikasi to-do list menggunakan tkinter python seder...

https://artificialanalysis.ai/microevals/p5js-physics-174972...
👍0
https://artificialanalysis.ai/microevals/p5js-physics-174972...

Cost per million tokens
👍0
Cost per million tokens

Pracuj s níže uvedeným Prompt 0. Zhodnoť, jak lze ještě dále...
👍0
Pracuj s níže uvedeným Prompt 0. Zhodnoť, jak lze ještě dále...

Flujos
👍0
Flujos

This is input from if node:
"subject": "[ZUNANJI] Vaš paket ...
👍0
This is input from if node: "subject": "[ZUNANJI] Vaš paket ...

Schreibe eine Metapher für Quantenphysik aus der Sicht eines...
👍0
Schreibe eine Metapher für Quantenphysik aus der Sicht eines...

A comet’s tail points in the following direction:

A) away f...
👍0
A comet’s tail points in the following direction: A) away f...

Jsi expertní vědecký, medicínský a Evidence-Based Medicine a...
👍0
Jsi expertní vědecký, medicínský a Evidence-Based Medicine a...

aoue
👍0
aoue

Complete the following Python function:

```python
from typi...
👍0
Complete the following Python function: ```python from typi...

Kansızlığa karşı Demir eksikliğine karşı elma ve siyah üzüm ...
👍0
Kansızlığa karşı Demir eksikliğine karşı elma ve siyah üzüm ...

👍0
Full Pipeline eval

I dunno just trying some stuff.

StoreAgent Customer-Facing WhatsApp LLM Eval v1
👍0
StoreAgent Customer-Facing WhatsApp LLM Eval v1

Production-oriented evaluation for StoreAgent's customer-facing WhatsApp LLM. Tests intent/routing accuracy, entity extraction, grounded commerce reasoning, multi-turn context, Arabic/English/code-switching quality, conversational naturalness, clarification behavior, hallucination resistance, and safe proposed actions. Prices, stock, payment state, order state and irreversible actions remain deterministic backend truth and must never be invented or independently changed by the model.