MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 5341-5360 of 7035

👍0
salatalık besin değerleri

👍0
The AetherStore Challenge: A Long-Horizon Systems Engineering Benchmark for Frontier AI
This prompt tests long-horizon reasoning, cross-layer systems engineering (kernel to distributed consensus), formal invariant enforcement, and zero-shot low-level code correctness, while actively penalizing the hand-waving and boilerplate typical of shallower models.

👍0
深度研究,逻辑推演以下选股条件层层叠加暗示了什么?详细描绘未来股价走势的画像?深度研究有没有实盘参考价值?盈亏比怎么样?...

👍0
Реши ребус, найдя закономерность. Любую.
Ребус:
1+1=0
2+2=2...

👍0
ne írj semmi mást csak a teljes fájlokat es kommentek nem le...

👍0
Estimate the level of recoverable easily extractable commerc...

👍0
Clifford vs Tensor

👍0
Yoshaa
Sksk

👍0
What are best months to go mushroom picking in Poland?

👍0
Your task is to take the provided technical description and ...

👍0
write an essay about gen ai

👍0
"""Public MCP server for reading and searching Markdown file...

👍0
Testing headshot image editing not generation from scratch

👍0
EXTERNAL METHODOLOGY-ARCHITECTURE AUDIT — ITERATION 23
You a...
👍0
帮我写一篇中考前一封信,作文 先写极其完整的思路,然后再写作文

👍0
In 2015, he said, a study completed in cooperation with the ...

👍0
Create a p5.js animation that is a cool interactive thing ba...

👍0
P20.2 — EVALUATOR / EC / BLINDING / EVIDENCE REGRESSION
1. ...

👍0
BS detection

👍0
Every day, Wendi feeds each of her chickens three cups of mi...