MicroEvals
Public evaluations
Showing 181-200 of 5354


This microeval will evaluate how well leading coding models can develop different types of web apps based on prompts that aren't ultra-specific, and just ask for the overall concept. This is to test how much LLMs have evolved in terms of design skills and common knowledge coding choices. There is one detailed prompt to see if it enhances the quality.

Tests whether a model follows explicit, verifiable instructions in trading and investing contexts without access to real-time market data. Each prompt carries deterministic pass/fail constraints — exact counts, banned words, ordering, output templates, and math on provided data — so grading measures instruction compliance, not market knowledge or data access.



Generator

![Strawrbrerrry [sic] eval](/_next/image?url=https%3A%2F%2Fartificialanalysiscdn.com%2Fmicro-evals%2F45d80a29bd4f4c68b3a92569454473d4.jpg&w=3840&q=75)
Reasoning should include the ability to generalize to unfamiliar words instead of memorizing answers. Let's see if models can detect the number of 'r's in the word "strawrbrerrry."

Visual perception of the 5 most important unsolved concepts in mathematics!







Perplexity AI:


