MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2761-2780 of 3494

👍0
Escreva o algorítimo da arvore rubro negra que tenha uma per...

👍0
Basic Test

👍0
Character reference sheet, 2D semi-realistic Indian animatio...

👍0
Janet’s ducks lay 16 eggs per day. She eats three for breakf...

👍0
homa

👍0
Write a psychological horror story about a man and a woman w...

👍0
title

👍0
You need to answer only one word. Think thoroughly, aggregat...

👍0
Сделай очень подробный, детальный анализ фильма Стэнли Кубри...

👍0
Storefront

👍0
Top-down luxury cosmetic photography featuring Fino Premium ...

👍0
I want you to design a prototype of a windows 11 alternative...

👍0
The Impossible Object Sculpture Park
This benchmark tests an LLM's ability to translate paradoxical, metaphorical, and logically inconsistent concepts into a functional and interactive 3D scene using Three.js. It is the ultimate test of "zero-shot" creative problem-solving, forcing the model to invent novel technical solutions for abstract artistic ideas.

👍0
合成生物學

👍0
Define
\[p = \sum_{k = 1}^\infty \frac{1}{k^2} \quad \text{a...

👍0
Game
Test

👍0
Prosper

👍0
Spx
Spx changing

👍0
REAL USAGE
Only real usage

👍0
math reasoning