视觉模型:支持图像输入的 LLM

使用 Artificial Analysis Visual Reasoning Index 比较支持图像和文本输入的多模态 LLM。比较各服务商的表现、价格和延迟,为视觉工作负载选择最佳的图像输入 LLM。如需了解更多详情,请参阅方法论页面

亮点

MMMU Pro (multimodal reasoning intelligence benchmark) · Higher is better
Output tokens per second · Higher is better
USD per 1k images at 1MP (1024x1024) · Lower is better

分析摘要

视觉推理与图像输入价格

Visual reasoning intelligence: MMMU Pro evaluation · Image input price: USD per 1k images at 1MP (1024x1024)
Most attractive quadrant

视觉推理与延迟(单张图像及 1,000 个语言输入 token)

Visual reasoning intelligence: MMMU Pro evaluation · Seconds to first token received
Most attractive quadrant

智能

视觉推理智能(MMMU Pro 评测)

Visual reasoning intelligence: MMMU Pro evaluation

价格

Pricing: Image Input Pricing

Image input price: USD per 1k images at 1MP (1024x1024)

价格:语言输入、图像输入与语言输出

Price (USD per M Tokens) · Image input price: USD per 1k images at 1MP (1024x1024) · Lower is better

延迟与速度

延迟(单张图像及 1,000 个语言输入 token)

Seconds to first token received · Lower is better

延迟波动(单张图像及 1,000 个语言输入 token)

Seconds to first token received · Results by percentile · Lower is better

输出速度(单张图像及 1,000 个语言输入 token)

Output tokens per second · Higher is better