MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 7541-7560 of 7809

👍0
Kırkkilit bitkisini çayının günün hangi saatinde tüketmek da...

👍0
Tìm nền tảng tương tự arena.ai và zenmux.ai

👍0
deliver profound upgrades to enhance novelty, usefulness to ...

👍0
saq

👍0
You are a creative coder and an expert in computer graphics,...

👍0
模拟我大师赛

👍0
You are a Senior Business Development Strategist with expert...
👍0
111

👍0
You are the administrative services manager responsible for ...

👍0
Как да си отворим фирма ЕООД

👍0
In 2015, he said, a study completed in cooperation with the ...

👍0
Выступи в роли прагматичного, жесткого карьерного стратега и...

👍0
OscarGolf-3529

👍0
Generate a fully synthetic, production-ready pre-training da...

👍0
In 2015, he said, a study completed in cooperation with the ...

👍0
Runner system-prompt efficiency — improvement task
How good models are at improving existing systems.

👍0
voy a crear una web para una empresa que vende bebidas refre...

👍0
female wrestler
guess fav female wrestle by birth chart. correct answer: Sasha Banks

👍0
/**
* core-ui.css — Techsillica
*
* CSS companion to core...
👍0
Aegis-7: Extreme Reasoning, Embedded Systems, and Linguistic Constraint Benchmark
An advanced stress test designed to push frontier Large Language Models (LLMs) to their operational limits under strict, multi-layered constraints. The prompt forces the AI to resolve a critical orbital recovery dilemma by calculating weighted utility matrices step-by-step, writing memory-safe embedded Rust code (#![no_std]), conducting ethical analysis under a severe linguistic constraint (an 80+ word paragraph omitting the letter "e"), and enforcing strict JSON output adherence without conversational filler.