Vision-Modelle: LLMs mit Bildeingabe

Vergleichen Sie multimodale LLMs, die Bild- und Texteingaben unterstützen, mit dem Artificial Analysis Visual Reasoning Index. Vergleichen Sie Leistung, Preise und Latenz verschiedener Anbieter, um das beste bildfähige LLM für visuelle Workloads auszuwählen. Weitere Details finden Sie auf der Methodikseite.

Wichtigste Ergebnisse

MMMU Pro (multimodal reasoning intelligence benchmark) · Higher is better
Output tokens per second · Higher is better
USD per 1k images at 1MP (1024x1024) · Lower is better

Zusammenfassende Analyse

Visuelles Schlussfolgern vs. Preis der Bildeingabe

Visual reasoning intelligence: MMMU Pro evaluation · Image input price: USD per 1k images at 1MP (1024x1024)
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

Visuelles Schlussfolgern vs. Latenz (Eingabe: ein Bild und 1.000 Texttokens)

Visual reasoning intelligence: MMMU Pro evaluation · Seconds to first token received
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Intelligenz

Intelligenz beim visuellen Schlussfolgern (Evaluation MMMU Pro)

Visual reasoning intelligence: MMMU Pro evaluation
Reasoning models are indicated by a lightbulb icon

Multimodal reasoning quality evaluation based on 1.7k questions which require interpreting and reasoning over images.

Preise

Pricing: Image Input Pricing

Image input price: USD per 1k images at 1MP (1024x1024)
Reasoning models are indicated by a lightbulb icon

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Preise: Texteingabe, Bildeingabe und Textausgabe

Price (USD per M Tokens) · Image input price: USD per 1k images at 1MP (1024x1024) · Lower is better
Reasoning models are indicated by a lightbulb icon

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

Latenz und Geschwindigkeit

Latenz (Eingabe: ein Bild und 1.000 Texttokens)

Seconds to first token received · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Latenzvarianz (Eingabe: ein Bild und 1.000 Texttokens)

Seconds to first token received · Results by percentile · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Picture of the author

Ausgabegeschwindigkeit (Eingabe: ein Bild und 1.000 Texttokens)

Output tokens per second · Higher is better
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).