视觉模型:支持图像输入的 LLM

使用 Artificial Analysis Visual Reasoning Index 比较支持图像和文本输入的多模态 LLM。比较各服务商的表现、价格和延迟,为视觉工作负载选择最佳的图像输入 LLM。如需了解更多详情,请参阅方法论页面

亮点

MMMU Pro (multimodal reasoning intelligence benchmark) · Higher is better
Output tokens per second · Higher is better
USD per 1k images at 1MP (1024x1024) · Lower is better

分析摘要

视觉推理与图像输入价格

Visual reasoning intelligence: MMMU Pro evaluation · Image input price: USD per 1k images at 1MP (1024x1024)
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

视觉推理与延迟(单张图像及 1,000 个语言输入 token)

Visual reasoning intelligence: MMMU Pro evaluation · Seconds to first token received
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

智能

视觉推理智能(MMMU Pro 评测)

Visual reasoning intelligence: MMMU Pro evaluation
Reasoning models are indicated by a lightbulb icon

Multimodal reasoning quality evaluation based on 1.7k questions which require interpreting and reasoning over images.

价格

Pricing: Image Input Pricing

Image input price: USD per 1k images at 1MP (1024x1024)
Reasoning models are indicated by a lightbulb icon

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

价格:语言输入、图像输入与语言输出

Price (USD per M Tokens) · Image input price: USD per 1k images at 1MP (1024x1024) · Lower is better
Reasoning models are indicated by a lightbulb icon

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

延迟与速度

延迟(单张图像及 1,000 个语言输入 token)

Seconds to first token received · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

延迟波动(单张图像及 1,000 个语言输入 token)

Seconds to first token received · Results by percentile · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Picture of the author

输出速度(单张图像及 1,000 个语言输入 token)

Output tokens per second · Higher is better
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).