
Main
Prompt
Assess this conspect over SciCode benchmark **SciCode** is a benchmark that tests whether an AI model can write code to solve **realistic scientific research problems**, rather than ordinary programming puzzles. It contains: - 80 main problems decomposed into 338 smaller subproblems - 16 scientific subfields Tests model's - **Scientific knowledge** of model: concepts and formulas - **Reasoning** converting the scientific description into an algorithm - **Code synthesis** > In the original 2024 paper, the best tested model, Claude 3.5 Sonnet, solved only **4.6% of complete problems**, showing how complex the benchmark initially was
Drag to resize
Drag to resize
Drag to resize
Drag to resize