ビジョンモデル:画像入力に対応するLLM

Artificial Analysis Visual Reasoning Indexを使って、画像とテキスト入力に対応するマルチモーダルLLMを比較します。プロバイダー間の性能、料金、遅延を比較し、ビジョンワークロードに最適な画像対応LLMを選べます。詳細は方法論のページをご覧ください。

ハイライト

MMMU Pro (multimodal reasoning intelligence benchmark) · Higher is better
Output tokens per second · Higher is better
USD per 1k images at 1MP (1024x1024) · Lower is better

分析の概要

視覚推論と画像入力料金

Visual reasoning intelligence: MMMU Pro evaluation · Image input price: USD per 1k images at 1MP (1024x1024)
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

視覚推論と遅延(画像1枚と1,000言語トークンの入力)

Visual reasoning intelligence: MMMU Pro evaluation · Seconds to first token received
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

知能

視覚推論の知能(MMMU Pro評価)

Visual reasoning intelligence: MMMU Pro evaluation
Reasoning models are indicated by a lightbulb icon

Multimodal reasoning quality evaluation based on 1.7k questions which require interpreting and reasoning over images.

料金

Pricing: Image Input Pricing

Image input price: USD per 1k images at 1MP (1024x1024)
Reasoning models are indicated by a lightbulb icon

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

料金:言語入力、画像入力、言語出力

Price (USD per M Tokens) · Image input price: USD per 1k images at 1MP (1024x1024) · Lower is better
Reasoning models are indicated by a lightbulb icon

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

遅延と速度

遅延(画像1枚と1,000言語トークンの入力)

Seconds to first token received · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

遅延のばらつき(画像1枚と1,000言語トークンの入力)

Seconds to first token received · Results by percentile · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Picture of the author

出力速度(画像1枚と1,000言語トークンの入力)

Output tokens per second · Higher is better
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).