MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 8701-8720 of 9005

👍0
Tìm nền tảng tương tự arena.ai và zenmux.ai

👍0
deliver profound upgrades to enhance novelty, usefulness to ...

👍0
saq
👍0
Create a self-study roadmap with the best books for learning...
👍0
Finance

👍0
You are a creative coder and an expert in computer graphics,...

👍0
模拟我大师赛

👍0
You are a Senior Business Development Strategist with expert...
👍0
111

👍0
You are the administrative services manager responsible for ...

👍0
Как да си отворим фирма ЕООД

👍0
In 2015, he said, a study completed in cooperation with the ...
👍0
P49.1 — HUMAN-USE REMEDIATION / DEFERRED-ACTION POST-REPAIR ...
👍0
Pregunta de investigación profunda: ¿Es posible organizar el...

👍0
Выступи в роли прагматичного, жесткого карьерного стратега и...

👍0
OscarGolf-3529

👍0
Generate a fully synthetic, production-ready pre-training da...
👍0
Educational 3D game inside a CPU (Three.js)
Ability to generate a complete 3D game in a single self-contained HTML file with Three.js from a long, detailed prompt. Evaluate: working code with no console errors, completeness (all levels, Codex, quiz and Lab Mode), accuracy of CPU architecture concepts, visual quality (bloom, circuit traces, particles), gameplay and fun, performance (60 FPS) and adherence to the prompt specifications.

👍0
In 2015, he said, a study completed in cooperation with the ...

👍0
Runner system-prompt efficiency — improvement task
How good models are at improving existing systems.