MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2401-2420 of 4093

👍0
You are an administrative operations lead in a government de...

👍0
Create duolingo for music theory

👍0
generame una pagina en html para una empresa de desarrollo d...

👍0
sap abap alv report
simple sap abap program with minimal details

👍0
Act as a Staff AI Research Engineer and Senior Backend Archi...

👍0
Construct a complete, production-ready software system that ...

👍0
test1 of real world prompts for small businesses

👍0
你现在是一名具有高度原创能力、电影思维和文学表达能力的编剧。请为我进行一次15–20分钟电影短片的故事创意开发。 我希望...

👍0
Create a p5.js animation that is a cool interactive thing ba...

👍0
hiibb

👍0
Minecraft

👍0
genara la imagen realista de un perro caminando por la luna

👍0
Statement 1 | A permutation that is a product of m even perm...

👍0
from typing import Any
from . import exa_search_agent
from ...

👍0
LLM as Designer: Self-Evolving OS Visual Generation
This benchmark tests an LLM's ability to create a dynamic visual narrative where an AI agent progressively "builds" and "designs" an operating system interface directly in the browser. It combines creative storytelling, dynamic HTML/CSS/JS generation, and animated visualization of a conceptual AI design process. The goal is to make the user feel like they are watching an AI create its own visual environment.

👍0
eval

👍0
order

👍0
Aungko1990
Arfu

👍0
UniformComet-5447

👍0
תיצור לי את כל הקוד לאפליקציה מוכנה שמה שהיא תעשה זה שליטה מ...