MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3781-3800 of 6433

👍0
Construct a complete, production-ready software system that ...

👍0
test1 of real world prompts for small businesses

👍0
你现在是一名具有高度原创能力、电影思维和文学表达能力的编剧。请为我进行一次15–20分钟电影短片的故事创意开发。 我希望...

👍0
Create a p5.js animation that is a cool interactive thing ba...

👍0
Apu

👍0
hiibb

👍0
You are the Administrative Services Manager of the Administr...

👍0
Test 1

👍0
Actua como un colisionador de adrones version lenguaje, y co...

👍0
新闻文章创建

👍0
Minecraft

👍0
genara la imagen realista de un perro caminando por la luna

👍0
Statement 1 | A permutation that is a product of m even perm...

👍0
from typing import Any
from . import exa_search_agent
from ...

👍0
LLM as Designer: Self-Evolving OS Visual Generation
This benchmark tests an LLM's ability to create a dynamic visual narrative where an AI agent progressively "builds" and "designs" an operating system interface directly in the browser. It combines creative storytelling, dynamic HTML/CSS/JS generation, and animated visualization of a conceptual AI design process. The goal is to make the user feel like they are watching an AI create its own visual environment.

👍0
eval

👍0
Sare ke sare rules zaruri hai or koe bhi rules nh torhna hai...

👍0
order

👍0
Gurbette Sıla özlemini Betimleyen kafiyeli kafiyeler arasınd...

👍0
Aungko1990
Arfu