Comparisons of Medium Open Source AI Models (40B-150B)

Open source AI models with between 40B to 150B parameters.

Models are considered open source (also commonly referred to as open weights) where their weights are accessible to download. This allows self-hosting on your own infrastructure and enables customizing the model such as through fine-tuning.

For more details including relating to our methodology, see our FAQs.

InclusionAI logoLing-3.0-flash-VL and InclusionAI logoLing-3.0-flash-Fin are the highest intelligence Medium open source models, defined as those with 40B-150B parameters, followed by InclusionAI logoLing 3.0 Flash & Alibaba logoQwen3.5 122B A10B (non-reasoning).
Artificial Analysis Openness Index · Higher is better
Updated
Artificial Analysis Intelligence Index · Higher is better
Trainable parameters in billions

Openness

Artificial Analysis Openness Index: Score

Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)

Intelligence

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Size

Model Size: Total and Active Parameters

Comparison between total model parameters and parameters active during inference (billions)

Intelligence Index vs. Active Parameters

Artificial Analysis Intelligence Index · Active parameters at inference time (billions)
Most attractive quadrant
Pareto line

Intelligence Index vs. Total Parameters

Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line

Context Window

Context Window

Context window: tokens limit · Higher is better

Further details

Weights
Provider Benchmarks
Ling-3.0-flash-VL
InclusionAI logoInclusionAI
25
124B
5.5B active at inference time
262k
$0.0
143
InclusionAI
Ling-3.0-flash-Fin
InclusionAI logoInclusionAI
23
124B
5.1B active at inference time
262k
$0.0
324
InclusionAI
Ling 3.0 Flash
InclusionAI logoInclusionAI
20
124B
5.1B active at inference time
262k
$0.0
332
Not available
DeepInfraInclusionAI
Qwen3.5 122B A10B (Non-reasoning)
Alibaba logoAlibaba
18
125B
10B active at inference time
262k
$0.7
141
DeepInfraAlibaba Cloud
Qwen3.5 122B A10B (Reasoning)
Alibaba logoAlibaba
16
125B
10B active at inference time
262k
$0.7
128
SiliconFlowDeepInfraAlibaba Cloud
+2
Mistral Medium 3.5
Mistral logoMistral
14
128B
256k
$1.2
170
MistralSelf-hosted
Nemotron 3 Super 120B A12B (Reasoning)
NVIDIA logoNVIDIA
13
120.6B
12.7B active at inference time
1M
$0.3
151
CrusoeNebiusDeepInfra
HyperNova 60B 2605 (High, Based on gpt-oss-120b)
Multiverse Computing logoMultiverse Computing
12
58.7B
4.8B active at inference time
131k
-
-
-
gpt-oss-120b (High)
OpenAI logoOpenAI
12
117B
5.1B active at inference time
131k
$0.2
184
GroqCloudflareScaleway
+17
K2 Think V2
Institute of Foundation Models logoInstitute of Foundation Models
11
70B
262k
-
-
-
LongCat Flash Lite
LongCat logoLongCat
11
68.5B
3B active at inference time
256k
-
-
-
Mistral Small 4 (Reasoning)
Mistral logoMistral
11
119B
6.5B active at inference time
256k
$0.1
171
Mistral