MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2421-2440 of 4081
![[SYSTEM] Convert a natural-language user idea + target aspec...](/_next/image?url=https%3A%2F%2Fartificialanalysiscdn.com%2Fmicro-evals%2F81e2375a0ebb4729b2d1b27869f1bb82.jpg&w=3840&q=75)
👍0
[SYSTEM] Convert a natural-language user idea + target aspec...

👍0
Create a fun arcade racer in JS. The game should be set in a...

👍0
roma

👍0
# 全面实施提示:50亿参数混合稀疏MoE语言模型训练系统
构建一个完整的生产级训练系统,用于训练一个50亿参数混合稀...

👍0
Book Segmentation
Reasoning ability to see if the model can reason which sentence/text chunk is spoken by which character

👍0
Vì sao nhiều người bảo thủ tính bừa bãi của mình?

👍0
In life, learning what skills or knowledges/subjects is the ...

👍0
tipos de nubes

👍0
Guessing output with built-in usage.
This code is trivial to run and see the actual output, but most of the models so far fails spectacularly if they can only guessing and not allows code eval. This test is good to determine if model is actually not bluffing.

👍0
==================================================
QISSAH — ...

👍0
compara estos dos modelos

👍0
Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...

👍0
necesito extraer los assets y codigos de un apk, analiza com...

👍0
hffgh

👍0
Recommend Most AI Agent

👍0
First person who discovered that Earth is round.

👍0
How is the optimal automatic bot trading setup that have the...

👍0
test

👍0
Toppers Game Developing
Game Developing

👍0
You’ve been hired as an In Ear Monitor (IEM) Tech for a tour...