MicroEvals
Public evaluations
Showing 8701-8720 of 8996






Ability to generate a complete 3D game in a single self-contained HTML file with Three.js from a long, detailed prompt. Evaluate: working code with no console errors, completeness (all levels, Codex, quiz and Lab Mode), accuracy of CPU architecture concepts, visual quality (bloom, circuit traces, particles), gameplay and fun, performance (60 FPS) and adherence to the prompt specifications.


How good models are at improving existing systems.


guess fav female wrestle by birth chart. correct answer: Sasha Banks
Asking different models to comeout with QA Automation details

An advanced stress test designed to push frontier Large Language Models (LLMs) to their operational limits under strict, multi-layered constraints. The prompt forces the AI to resolve a critical orbital recovery dilemma by calculating weighted utility matrices step-by-step, writing memory-safe embedded Rust code (#![no_std]), conducting ethical analysis under a severe linguistic constraint (an 80+ word paragraph omitting the letter "e"), and enforcing strict JSON output adherence without conversational filler.

FC-BGA 공급업체들의 면적당 CAPA 분석
圆形糖果分别有苹果7、桃9、西瓜8;五角星糖果分别有苹果7、桃6、西瓜4。起码取多少颗,才能确保不同形状的苹果和桃都拿到一颗