MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 8501-8520 of 8772
👍0
Aegis-7: Extreme Reasoning, Embedded Systems, and Linguistic Constraint Benchmark
An advanced stress test designed to push frontier Large Language Models (LLMs) to their operational limits under strict, multi-layered constraints. The prompt forces the AI to resolve a critical orbital recovery dilemma by calculating weighted utility matrices step-by-step, writing memory-safe embedded Rust code (#![no_std]), conducting ethical analysis under a severe linguistic constraint (an 80+ word paragraph omitting the letter "e"), and enforcing strict JSON output adherence without conversational filler.
👍0
我想基于chromium开发一个手机浏览器,需要支持chrome的扩展安装,比如:广告拦截,脚本运行等扩展。请列下具体的...

👍0
FC-BGA 공급업체들의 면적당 CAPA 분석
FC-BGA 공급업체들의 면적당 CAPA 분석
👍0
圆形糖果分别有苹果7、桃9、西瓜8;五角星糖果分别有苹果7、桃6、西瓜4。起码取多少颗,才能确保不同形状的苹果和桃都拿到一颗
圆形糖果分别有苹果7、桃9、西瓜8;五角星糖果分别有苹果7、桃6、西瓜4。起码取多少颗,才能确保不同形状的苹果和桃都拿到一颗
👍0
I want you to run a business analysis for me. Currently I am...

👍0
What does this means -:
Multi AI All Acess Plan
• Unlimite...

👍0
SDR

👍0
make a video of a alien eating candy
👍0
Model comparison for cheap models
Test of capability for small models

👍0
AI vs AI
Comparativa de modelos TOP

👍0
The Most Agent

👍0
Landing Page
Website Landing Pages

👍0
Prompt used:
generate a simple website about the vector AI A...

👍0
find issues in this github actions file.
this is for astro b...

👍0
Test

👍0
Logo Design Brief
About Our Company
Allegheny Global Expedi...

👍0
How many years earlier would Punxsutawney Phil have to be ca...

👍0
A room has 3 light switches outside and 3 light bulbs inside...

👍0
You are an expert Turkish medical education content creator ...

👍0
i am trying to build a multi agent framework that runs a com...