MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2381-2400 of 4048

👍0
Act as a Staff AI Research Engineer and Senior Backend Archi...

👍0
Construct a complete, production-ready software system that ...

👍0
test1 of real world prompts for small businesses

👍0
你现在是一名具有高度原创能力、电影思维和文学表达能力的编剧。请为我进行一次15–20分钟电影短片的故事创意开发。 我希望...

👍0
Create a p5.js animation that is a cool interactive thing ba...

👍0
hiibb

👍0
Minecraft

👍0
genara la imagen realista de un perro caminando por la luna

👍0
Statement 1 | A permutation that is a product of m even perm...

👍0
from typing import Any
from . import exa_search_agent
from ...

👍0
LLM as Designer: Self-Evolving OS Visual Generation
This benchmark tests an LLM's ability to create a dynamic visual narrative where an AI agent progressively "builds" and "designs" an operating system interface directly in the browser. It combines creative storytelling, dynamic HTML/CSS/JS generation, and animated visualization of a conceptual AI design process. The goal is to make the user feel like they are watching an AI create its own visual environment.

👍0
eval

👍0
order

👍0
Aungko1990
Arfu

👍0
UniformComet-5447

👍0
תיצור לי את כל הקוד לאפליקציה מוכנה שמה שהיא תעשה זה שליטה מ...

👍0
Jarvis UI website test
How does Hallucination affect website development?

👍0
help me learn how EU works from JTI corporate affairs and communication manager perspective, working in italy

👍0
generate me an snake game in html

👍0
teting
lets see what happen