MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 1721-1740 of 3640

👍0
docker-mcp — summary for oracle review
What
Building a TypeS...

👍0
가성비관점에서 비교해줘

👍0
what you want?

👍0
Create an agent which gives me CPU & GPU utilization

👍0
GOV TEST
THE AI CAN MAKE AN INFOGRAPH ABOUT THE COMPOSITION OF THE GOV.

👍0
顶级模型的评测

👍0
Write a 300+ word summary of the wikipedia page "https://en....

👍0
Recommended AI agent

👍0
create a .pdf on how SSRI works and inhibit presynpatic bloc...

👍0
cost vs. intelegence

👍0
write a hello world script in ocaml

👍0
### General Instructions
- Planning: Step-back and take a b...

👍0
Ich habe hier einen 4,7 uF Elektrolytkondensator mit 50 V Ra...

👍0
You are a senior ML research scientist and technical communi...

👍0
- 内容电商创业者,运营抖音账号「水星的审美书单」
- 工作流覆盖:选品 → AI 文案 → 配音 → 视频制作 →...

👍0
证明无理数

👍0
You are an elite, highly influential Persian political comme...

👍0
Make a Russian landing page for a company producing cucumber...

👍0
LLM Ultimate Challenge: Interactive GLSL Shader Art
This benchmark tests an LLM's ability to handle a multi-language, algorithmically complex task. It requires generating a single HTML file with JavaScript (using Three.js) to manage the scene, and GLSL shader code to render a dynamic, interactive fractal. This evaluates advanced knowledge of mathematics, GPU programming, and system integration.

👍0
test