MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 201-220 of 4048

👍1
Crée un mini-jeu 3D jouable dans le navigateur en un seul fi...

👍1
Medical Quiz 3
Cardiac phisiology quiz.

👍1
CloudOps Architect Benchmark
A five-scenario benchmark designed to evaluate an LLM’s ability to perform as a senior Cloud Architect, DevOps Engineer, and SRE. The test covers production Kubernetes architecture, incident response, cloud security, legacy-to-cloud migration, and FinOps. It focuses on technical correctness, practical decision-making, trade-off reasoning, scalability, reliability, and depth of Cloud and Kubernetes expertise.

👍1
Guess a number

👍1
uyftu
API Key for "diqddi88@gmail.com"

👍1
Grok 4.3 high

👍1
Double Pendulum Simulation

👍1
Test 1
Test 1

👍1
Two Math challenges
Persian prompts but translated in output

👍1
AtttTention 't' count
More 't' s and a 'T' is added to confuse AI

👍1
Привет!

👍0
I've planning to create a story telling YouTube channel wher...

👍0
Three.js Games

👍0
pikng

👍0
what is the capital of canada

👍0
hi事实上

👍0
from pptx import Presentation
from pptx.util import Inches, ...

👍0
Tell about gandhiji

👍0
Principal Security Engineer
You are reviewing this as if ...

👍0
Posicionamiento Alegra
Evaluación de prompts según modelos