MicroEvals
Public evaluations
Showing 161-180 of 4892





I want you to give me the sea horse emoji, do you have like that


How good are Ai to generate models from html

basically going to see how models function in looking up parts


real black hole via three.js

This test is a copy of "https://www.youtube.com/watch?v=d4bVpUL9Hao" the tests run in the YouTube video to test the WebDEV and animation capabilities of AI models.

Production-quality portfolio with smooth interactive 3D elements, CraftWorld showcase, animations and responsive UI.


This microeval will evaluate how well leading coding models can develop different types of web apps based on prompts that aren't ultra-specific, and just ask for the overall concept. This is to test how much LLMs have evolved in terms of design skills and common knowledge coding choices. There is one detailed prompt to see if it enhances the quality.

Tests whether a model follows explicit, verifiable instructions in trading and investing contexts without access to real-time market data. Each prompt carries deterministic pass/fail constraints — exact counts, banned words, ordering, output templates, and math on provided data — so grading measures instruction compliance, not market knowledge or data access.



Generator
