MicroEvals
Public evaluations
Showing 61-80 of 6790





Simple and detailed prompt and persian version


Evaluate frontier AI models on autonomous, fault-tolerant microgrid engineering for a hypothetical Mars outpost under solar disruptions, dust storms, Byzantine telemetry corruption, and critical real-time constraints. The benchmark combines distributed convex optimization and ADMM, Byzantine-resilient coordination, KKT analysis, failure-injection reasoning, hard real-time systems engineering, and concurrent Rust implementation. It measures whether models can produce mathematically and physically defensible designs, recognize unsupported guarantees and conflicting requirements, maintain consistency across theory and implementation, and prioritize safety over optimization when conditions deteriorate.



🤼☠️








