MicroEvals
Public evaluations
Showing 321-340 of 2530






r we cooked



A focused set of three coding tasks designed to test practical reasoning and implementation skills in real-world scenarios: robust string parsing with edge cases, async concurrency control with ordered results, and immutable nested state updates. Each prompt omits the explicit answer, making it suitable for evaluating both problem understanding and solution correctness without leakage of target outputs.


hey chat r we cooked



Laboratorio donde poder crear tus propios arquitecturas de Ia y redes neuronales


Chatbot with pdf RAG


This benchmark tests an LLM's ability to interpret and implement abstract, philosophical, and artistic theories. It requires the model to translate Wassily Kandinsky's theories on synesthesia (the connection between sound, color, and shape) into an interactive audio-visual experience. Success is judged on the creative fidelity to the artistic concept, not just technical execution.

