MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 7781-7800 of 8801

👍0
t3-01-repair-completion-ordering
T3-complex

👍0
hi hello

👍0
A.P.6 — RED-TEAM AUDIT LMC-40 + WEB CROSS-CHECK + AUTONOMOUS...

👍0
Start with the word MAGIC. Create a word ladder with as many...

👍0
现在我想做一个视觉算法 用来监测家里宠物的动态 猫猫狗狗的
比如识别动物在监控中的位置 分析是在做什么等等
帮我找找有没...

👍0
Regarding your question, Bain and McKinsey do not really do ...

👍0
Comp Opt Qs
👍0
---
name: architect
description: Software architecture speci...

👍0
传统制造业如何利用SEO、GEO、ASO
谷歌和AI怎样最大化助力传统制造业打开拓客
👍0
You are an administrative operations lead in a government de...

👍0
Professeur universitaire en Chirurgie cardio-vasculaire
réponse a 27 qcms en Chirurgie cardio-vasculaire

👍0
Classification accord
👍0
The Apex LLM Capability Benchmark: 20 Single-File Interactive Applications
A comprehensive suite of 20 stress-test prompts designed to evaluate the limits of modern frontier Large Language Models (LLMs) across real-world software engineering, mathematical modeling, interactive UI/UX design, physics engines, and algorithmic efficiency. Each prompt demands a production-grade, fully functional, self-contained single-file HTML/JS/CSS application built without external dependencies or placeholders.

👍0
Jsi nezávislý red-team evaluátor řídicího promptu LLM. Odpov...

👍0
Jsi nezávislý red-team evaluátor řídicího promptu LLM. Odpov...
👍0
You are the senior engineer reviewing a proposed feature for...

👍0
Give me a joke

👍0
Jsi expertní vědecký, medicínský a Evidence-Based Medicine a...
👍0
Find info

👍0
The cyclic subgroup of Z_24 generated by 18 has order
A) 4
...