MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 341-360 of 8798

👍1
Extensively give me an Islamic Table of Contents in 80 parts...

👍1
CloudOps Architect Benchmark
A five-scenario benchmark designed to evaluate an LLM’s ability to perform as a senior Cloud Architect, DevOps Engineer, and SRE. The test covers production Kubernetes architecture, incident response, cloud security, legacy-to-cloud migration, and FinOps. It focuses on technical correctness, practical decision-making, trade-off reasoning, scalability, reliability, and depth of Cloud and Kubernetes expertise.

👍1
Guess a number

👍1
# Prompt de investigación: Nueva arquitectura/algoritmo para...

👍1
uyftu
API Key for "diqddi88@gmail.com"

👍1
1111

👍1
Grok 4.3 high

👍1
Double Pendulum Simulation

👍1
Integral

👍1
Test 1
Test 1

👍1
Two Math challenges
Persian prompts but translated in output
👍1
yooooo
👍1
Crea una pagina web para portfolio

👍1
Extensively give me an Islamic Table of Contents in 80 parts...

👍1
AtttTention 't' count
More 't' s and a 'T' is added to confuse AI

👍1
КОЛЬБА
це кольба
👍1
In life, learning what skills or knowledges/subjects is the ...

👍1
Create a fun arcade racer in JS. The game should be set in a...

👍1
Async python implementation
Models are tasked with implementing a specific type of less commmon rate limiter which must be compatible with a third party library aiohttp
👍1
whats ur name