MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 1541-1560 of 3159

👍0
LLM Ultimate Challenge: Interactive GLSL Shader Art
This benchmark tests an LLM's ability to handle a multi-language, algorithmically complex task. It requires generating a single HTML file with JavaScript (using Three.js) to manage the scene, and GLSL shader code to render a dynamic, interactive fractal. This evaluates advanced knowledge of mathematics, GPU programming, and system integration.

👍0
test

👍0
My first comparison

👍0
I am planning a trip to Japan, and I would like thee to writ...

👍0
临床上下文增强测试

👍0
test

👍0
The Cool Professor: Which MBTI or Cognitive Function Combo (ie. Dom+Aux functions) represent the "cool professor" archetype that appeals to the modern young adult generation (Gen-Z and Millennials)?
Cool Professor Archetype - MBTI/Cognitive Functions

👍0
Landscaping Landing Page

👍0
Wordle Clone

👍0
משחק

👍0
test

👍0
test2

👍0
hi can you generate a video

👍0
Написать такой код вба для эксель 2013 который должен:
- У ...

👍0
Что нужно для канонизации святого в православной церкви?

👍0
what is this doing, please analyse following code snippet
...

👍0
Let me clarify Ok suppose 3.8 gpa at math cs Econ hard cours...

👍0
Academic Text Logic and Language Quality Test
This test evaluates how well different models revise academic text. It focuses on logical completeness, theoretical reasoning, linguistic precision, coherence, and readability. Models must preserve the original meaning and citations while correcting weak reasoning, conceptual ambiguity, and redundant expression.

👍0
Hibakfi

👍0
Riddle