MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 5081-5100 of 8492
👍0
game
👍0
SVG de Nave Espacial

👍0
kimi max3
👍0
Test

👍0
# 全面实施提示:50亿参数混合稀疏MoE语言模型训练系统
构建一个完整的生产级训练系统,用于训练一个50亿参数混合稀...

👍0
EXTERNAL METHODOLOGY-ARCHITECTURE AUDIT — ITERATION 23
You a...

👍0
Book Segmentation
Reasoning ability to see if the model can reason which sentence/text chunk is spoken by which character
👍0
whats the single best class in world of warcraft forever

👍0
HumanEval Question 32
HumanEval Question 32

👍0
设计飞机空调
空调设计计划

👍0
Vì sao nhiều người bảo thủ tính bừa bãi của mình?

👍0
In life, learning what skills or knowledges/subjects is the ...

👍0
tipos de nubes

👍0
SCB SCB-deploy_verify-T1 MUSE

👍0
Guessing output with built-in usage.
This code is trivial to run and see the actual output, but most of the models so far fails spectacularly if they can only guessing and not allows code eval. This test is good to determine if model is actually not bluffing.

👍0
Hazme un generador de tablero de catan
👍0
Task: Implement a spreadsheet engine ("sheet.py")
Write a s...

👍0
==================================================
QISSAH — ...
👍0
Estimate the 85th percentile household-networth of the follo...

👍0
fastapi test
testing a slm agaist a sota gpt model