Artificial Analysis
K
Artificial Analysis
Models
Coding Agents
Image, Speech, Video
Inference
Leaderboards
About
AI Trends
Arenas
K
All MicroEvals
👍
0
LLM缺陷
Create MicroEval
LLM缺陷
👍
0
Prompt
分析为什么现在最先进的LLM依然会偷懒,并且在需求笼统的情况下会降低执行标准
OpenAI
GPT-6 Astra (max)
👍
0
Drag to resize
Anthropic
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
👍
0
Drag to resize