Modelos de visão: LLMs com entrada de imagens

Compare LLMs multimodais que aceitam imagens e texto como entrada usando o Artificial Analysis Visual Reasoning Index. Compare o desempenho, os preços e a latência entre os provedores para escolher o melhor LLM com entrada de imagens para cargas de trabalho de visão. Para mais detalhes, consulte a página de metodologia.

Destaques

MMMU Pro (multimodal reasoning intelligence benchmark) · Higher is better
Output tokens per second · Higher is better
USD per 1k images at 1MP (1024x1024) · Lower is better

Resumo da análise

Raciocínio visual vs. preço da entrada de imagens

Visual reasoning intelligence: MMMU Pro evaluation · Image input price: USD per 1k images at 1MP (1024x1024)
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

Raciocínio visual vs. latência (entrada de uma imagem e 1.000 tokens de linguagem)

Visual reasoning intelligence: MMMU Pro evaluation · Seconds to first token received
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Inteligência

Inteligência de raciocínio visual (avaliação MMMU Pro)

Visual reasoning intelligence: MMMU Pro evaluation
Reasoning models are indicated by a lightbulb icon

Multimodal reasoning quality evaluation based on 1.7k questions which require interpreting and reasoning over images.

Preços

Pricing: Image Input Pricing

Image input price: USD per 1k images at 1MP (1024x1024)
Reasoning models are indicated by a lightbulb icon

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Preços: entrada de linguagem, entrada de imagens e saída de linguagem

Price (USD per M Tokens) · Image input price: USD per 1k images at 1MP (1024x1024) · Lower is better
Reasoning models are indicated by a lightbulb icon

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

Latência e velocidade

Latência (entrada de uma imagem e 1.000 tokens de linguagem)

Seconds to first token received · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Variação da latência (entrada de uma imagem e 1.000 tokens de linguagem)

Seconds to first token received · Results by percentile · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Picture of the author

Velocidade de saída (entrada de uma imagem e 1.000 tokens de linguagem)

Output tokens per second · Higher is better
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).