Comparisons of Large Open Source AI Models (>150B)
Open source AI models with over 150B parameters.
Models are considered open source (also commonly referred to as open weights) where their weights are accessible to download. This allows self-hosting on your own infrastructure and enables customizing the model such as through fine-tuning.
For more details including relating to our methodology, see our FAQs.
Highlights
Openness
Artificial Analysis Openness Index: Score
Intelligence
Artificial Analysis Intelligence Index
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
Size
Model Size: Total and Active Parameters
Intelligence Index vs. Active Parameters
Intelligence Index vs. Total Parameters
Context Window
Context Window
Further details
Weights | Provider Benchmarks | ||||||||
|---|---|---|---|---|---|---|---|---|---|
Kimi K3 (max) | 60 | 2.8T 104B active at inference time | 1M | $2.3 | 36 | +12 | |||
Qwen3.8 2.4T A95B | 58 | 2.4T 95B active at inference time | 984k | $1.2 | 25 | +3 | |||
GLM-5.3-Flash | 57 | 320B 18B active at inference time | 1M | $0.1 | 50 | +8 | |||
Qwen3.8-Flash-Next | 56 | 180B 6B active at inference time | 256k | $0.1 | 73 | ||||
DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | 53 | 1.6T 49B active at inference time | 1M | $0.7 | 66 | +5 | |||
GLM-5.2 (max) | 53 | 753B 40B active at inference time | 1M | $0.9 | 69 | +20 | |||
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | 52 | 284B 13B active at inference time | 1M | $0.2 | 119 | +13 | |||
Kimi K3 (low) | 48 | 2.8T 104B active at inference time | 1M | $2.3 | 36 | ||||
Motif 3 | 47 | 314B 13.2B active at inference time | 262k | - | - | Not available | - | ||
MiniMax-M3 | 45 | 428B 23B active at inference time | 1M | $0.2 | 118 | +9 | |||
DeepSeek V4 Pro (Reasoning, Max Effort) | 45 | 1.6T 49B active at inference time | 1M | $0.2 | 66 | +9 | |||
DeepSeek V4 Pro (Reasoning, High Effort) | 44 | 1.6T 49B active at inference time | 1M | $0.2 | 63 | +6 |