MicroEvals
Public evaluations
Showing 7561-7580 of 7819






How good models are at improving existing systems.


guess fav female wrestle by birth chart. correct answer: Sasha Banks

An advanced stress test designed to push frontier Large Language Models (LLMs) to their operational limits under strict, multi-layered constraints. The prompt forces the AI to resolve a critical orbital recovery dilemma by calculating weighted utility matrices step-by-step, writing memory-safe embedded Rust code (#![no_std]), conducting ethical analysis under a severe linguistic constraint (an 80+ word paragraph omitting the letter "e"), and enforcing strict JSON output adherence without conversational filler.

FC-BGA 공급업체들의 면적당 CAPA 분석



Test of capability for small models

Comparativa de modelos TOP


Website Landing Pages
