KI-Trends von Artificial Analysis

Analyse des aktuellen Stands der KI und der wichtigsten Trends, die den Fortschritt vorantreiben. Dazu gehören Analysen der Intelligenz, Effizienz und Architektur von Modellen, ihrer Inferenzgeschwindigkeit und -kosten sowie von Trainingstrends.

Fortschritt der KI

Verfolgung der fortlaufenden Weiterentwicklung von KI und der Position jedes führenden KI-Unternehmens.

Frontier Language Model Intelligence, Over Time

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Investitionsausgaben großer Technologieunternehmen im Zeitverlauf

Investitionsausgaben je Quartal (in Milliarden USD)

Represents major investments by tech companies in infrastructure, including AI hardware (like GPUs and data centers). Capex is a strong indicator of a company's commitment to AI development, as training and running frontier models requires significant computing resources.

Note: Capex data is sourced from publicly available financial reports, news articles, and primarily from the SEC.

Intelligence Index vs. Veröffentlichungsdatum

Artificial Analysis Intelligence Index · Release date
Most attractive region

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Führende Modelle nach KI-Labor

Höchster von jedem KI-Labor erzielter Artificial Analysis Intelligence Index
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Effizienz

Analyse der Entwicklung der KI-Effizienz. Dazu betrachten wir die Kosten für das Erreichen bestimmter Intelligenzniveaus und die Geschwindigkeit, mit der diese Intelligenz verfügbar ist.

Inferenzpreis von Sprachmodellen nach Intelligence-Index-Band im Zeitverlauf

Preis in USD pro 1 Mio. Tokens (Mischung der Preise für Cache-, Eingabe- und Ausgabetokens im Verhältnis 7:2:1). Die Bänder verwenden den Artificial Analysis Intelligence Index v4.1.
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Ausgabegeschwindigkeit von Sprachmodellen nach Intelligence-Index-Band im Zeitverlauf

Ausgabetokens pro Sekunde. Die Bänder verwenden den Artificial Analysis Intelligence Index v4.1.
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Länderanalyse

KI ist ein globales Phänomen. Wir zeigen, wo KI-Fortschritt stattfindet und wie sich die führenden Modelle verschiedener Länder vergleichen lassen. Die USA und China als zwei führende Zentren der KI-Entwicklung analysieren wir besonders eingehend.

Intelligenz von Frontier-Sprachmodellen nach Land im Zeitverlauf

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Offene Gewichte: Intelligenz von Frontier-Sprachmodellen nach Land im Zeitverlauf

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Leading Models by Country

Artificial Analysis Intelligence Index · Leading models
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Modelle mit offenen Gewichten

Modelle mit offenen Gewichten ermöglichen flexible Bereitstellungen und die Feinabstimmung auf bestimmte Anwendungsfälle. Wir analysieren die führenden Modelle mit offenen Gewichten und vergleichen ihre Intelligenz mit proprietären Modellen.

Progress in Open Weights vs. Proprietary Intelligence

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).

Modellarchitektur

Die Architektur von KI-Modellen beeinflusst ihre Leistung und Effizienz. Dieser Abschnitt untersucht Architekturtrends wie die zunehmende Verbreitung der Mixture-of-Experts-Architektur (MoE) und deren Zusammenhang mit Modellfähigkeiten.

Intelligence Index vs. Veröffentlichungsdatum nach Modellarchitektur

Artificial Analysis Intelligence Index, nur Modelle mit offenen Gewichten
Most attractive region

A model where only a subset of parameters ("experts") are active per input. Routing mechanisms select a few experts per forward pass, reducing computation while allowing the model to scale to many more parameters overall.

A model where all parameters are active for every input. Every forward pass involves the full network, making it computationally intensive but straightforward to train and deploy.

Model Size: Total and Active Parameters

Comparison between total model parameters and parameters active during inference
Reasoning models are indicated by a lightbulb icon

The total number of trainable weights and biases in the model, expressed in billions. These parameters are learned during training and determine the model's ability to process and generate responses.

The number of parameters actually executed during each inference forward pass, expressed in billions. For Mixture of Experts (MoE) models, a routing mechanism selects a subset of experts per token, resulting in fewer active than total parameters. Dense models use all parameters, so active equals total.

Intelligence Index vs. Active Parameters

Artificial Analysis Intelligence Index · Active parameters at inference time
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

The number of parameters actually executed during each inference forward pass, expressed in billions. For Mixture of Experts (MoE) models, a routing mechanism selects a subset of experts per token, resulting in fewer active than total parameters. Dense models use all parameters, so active equals total.

Intelligence Index vs. Gesamtparameter

Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

The total number of trainable weights and biases in the model, expressed in billions. These parameters are learned during training and determine the model's ability to process and generate responses.

Kontextlänge (Tokens), Median je Quartal

Mediane Kontextlänge (Tausend Tokens)

Larger context windows are relevant to RAG (Retrieval Augmented Generation) LLM workflows which typically involve reasoning and information retrieval of large amounts of data.

Maximum number of combined input & output tokens. Output tokens commonly have a significantly lower limit (varied by model).

Trainingsanalyse

Unsere Analyse von Trends beim Training von KI-Modellen betrachtet sowohl die Entwicklung der Größe von Trainingsläufen als auch den Zusammenhang zwischen der Größe eines Trainingslaufs und der Intelligenz des Modells.

Trainingstokens nach Modell

Trainingstokens in Billionen
No data available
Reasoning models are indicated by a lightbulb icon

The number of tokens used to train the model, represented in trillions.

Intelligence Index vs. Trainingstokens

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR · Trainingstokens in Billionen
Most attractive quadrant
No data available
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

The number of tokens used to train the model, represented in trillions.