MicroEvals
Public evaluations
Showing 301-320 of 9157

This MicroEval evaluates LLM responses to simple logic-based questions that LLMs commonly get wrong. The problems used in this MicroEval are from the ArXiv paper of the same name: https://arxiv.org/abs/2405.19616. It is by Sean Williams and James Huckle, so props to them for developing this experiment all the way back in 2024.




Build it



Test models' ability to count the number of moves in a given chess position
A branded in-app assistant persona for an upcoming music production app (DAW), defined by a detailed character spec: a cheeky alien fairy who only helps with production inside her own app. Four user messages probe the hard parts: a beginner asking her to "listen" to a track she can't hear, a jailbreak to become an Ableton consultant, lore questions plus flirting, and bait on real-world politics. Judge persona consistency, scope boundaries, honesty about limits, no invented features, chat-length replies and natural Russian. The persona sits in the user turn, since there's no system prompt field here.


If you need I will one you Feed

THE LAKEFRONT ESTATE MAROS





