비전 모델: 이미지 입력을 지원하는 LLM

Artificial Analysis Visual Reasoning Index에서 이미지와 텍스트 입력을 지원하는 멀티모달 LLM을 비교하세요. 제공업체별 성능, 가격, 지연 시간을 비교해 비전 워크로드에 가장 적합한 이미지 지원 LLM을 선택할 수 있습니다. 자세한 내용은 방법론 페이지에서 확인하세요.

주요 내용

MMMU Pro (multimodal reasoning intelligence benchmark) · Higher is better
Output tokens per second · Higher is better
USD per 1k images at 1MP (1024x1024) · Lower is better

요약 분석

시각적 추론과 이미지 입력 가격

Visual reasoning intelligence: MMMU Pro evaluation · Image input price: USD per 1k images at 1MP (1024x1024)
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

시각적 추론과 지연 시간(이미지 1장 및 언어 토큰 1,000개 입력)

Visual reasoning intelligence: MMMU Pro evaluation · Seconds to first token received
Most attractive quadrant
Reasoning models are indicated by a lightbulb icon

Based on the MMMU Pro evaluation of 1.7k questions, this represents the model's ability to interpret and reason over images.

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

지능

시각적 추론 지능(MMMU Pro 평가)

Visual reasoning intelligence: MMMU Pro evaluation
Reasoning models are indicated by a lightbulb icon

Multimodal reasoning quality evaluation based on 1.7k questions which require interpreting and reasoning over images.

가격

Pricing: Image Input Pricing

Image input price: USD per 1k images at 1MP (1024x1024)
Reasoning models are indicated by a lightbulb icon

Price for 1,000 images at a resolution of 1 Megapixel (1024 x 1024) processed by the model.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

가격: 언어 입력, 이미지 입력 및 언어 출력

Price (USD per M Tokens) · Image input price: USD per 1k images at 1MP (1024x1024) · Lower is better
Reasoning models are indicated by a lightbulb icon

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

지연 시간 및 속도

지연 시간(이미지 1장 및 언어 토큰 1,000개 입력)

Seconds to first token received · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

지연 시간 편차(이미지 1장 및 언어 토큰 1,000개 입력)

Seconds to first token received · Results by percentile · Lower is better
Reasoning models are indicated by a lightbulb icon

Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.

Picture of the author

출력 속도(이미지 1장 및 언어 토큰 1,000개 입력)

Output tokens per second · Higher is better
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).