MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2121-2140 of 3613

👍0
BAHAR MAH. 3.BAHÇE ÇIKMA SK. NO: 6 İÇ KApI NO:1 OSMANGAZİ Bu...
👍0
French correction
Which Gemini Free plan is best at correcting mistakes in French text

👍0
prof
xxx

👍0
Eval chat support
MoxChatModel

👍0
Find the order of the factor group (Z_4 x Z_12)/(<2> x <2>)
...

👍0
به نظرت دانشگاه درسدن و دانشگاه اشتوتگارت کدومشون برای علوم...

👍0
You are an administrative operations lead in a government de...

👍0
generame una pagina en html para una empresa de desarrollo d...

👍0
sap abap alv report
simple sap abap program with minimal details

👍0
Act as a Staff AI Research Engineer and Senior Backend Archi...

👍0
Construct a complete, production-ready software system that ...

👍0
test1 of real world prompts for small businesses

👍0
Create a p5.js animation that is a cool interactive thing ba...

👍0
hiibb

👍0
genara la imagen realista de un perro caminando por la luna

👍0
Statement 1 | A permutation that is a product of m even perm...

👍0
LLM as Designer: Self-Evolving OS Visual Generation
This benchmark tests an LLM's ability to create a dynamic visual narrative where an AI agent progressively "builds" and "designs" an operating system interface directly in the browser. It combines creative storytelling, dynamic HTML/CSS/JS generation, and animated visualization of a conceptual AI design process. The goal is to make the user feel like they are watching an AI create its own visual environment.

👍0
eval

👍0
order

👍0
Aungko1990
Arfu