MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3021-3040 of 10195
👍0
c koi la serendipité
👍0
hello plase you output Chinse if you can output Chinese
👍0
comparacion

👍0
Jsi expertní vědecký, medicínský a Evidence-Based Medicine a...
👍0
Create a single-file index.html using Three.js with OrbitCon...
👍0
Creality Ender-3 Neo Interactive 3D Simulator Benchmark
A technical benchmark in which multiple AI models develop an interactive, realistically scaled 3D simulation of the Creality Ender-3 Neo. The benchmark evaluates accurate bed dimensions, mechanically correct axis movement, manual controls, G-code uploading and playback, toolpath visualization, homing, movement constraints, performance, and progressive filament deposition. Every model receives the exact same prompt and is evaluated using the same criteria.

👍0
a man scared
👍0
svg

👍0
Statement 1 | A permutation that is a product of m even perm...

👍0
Haz un programa en python, una calculadora

👍0
Jsi nezávislý red-team evaluátor řídicího promptu LLM. Odpov...

👍0
هذه محاكاة فقط.
أنت تعمل داخل Hermes Agent.
الأدوات المتاح...

👍0
HOW YOU WORK TELL ME

👍0
Jsi expertní vědecký, medicínský a Evidence-Based Medicine a...

👍0
Solve the hydrogen atom using a melin transform

👍0
year 2000 problem

👍0
search thoroughly for painting contractors because I’m renov...

👍0
Может ли поместиться всё население китая на полуострове Камч...

👍0
Build a polished, playable browser game: Minecraft is a san...

👍0
test