AI 聊天机器人比较

嵌入网站、应用或消息渠道中的对话式智能体,可在聊天界面中回答问题或执行范围明确的任务,通常围绕常见问题、简单工作流或特定领域知识,而非长期负责一项任务。

如需比较语言模型,请参阅我们的模型基准测试

我们使用 AI 收集部分结果

亮点

Intelligence Index · Higher is better
Feature score · Higher is better · 6 categories tracked
Monthly price (USD) · Standard & premium plans

比较套餐选项

产品等级价格模型智能输入媒体生成工具记忆MCP连接器功能得分应用隐私
ChatGPT Plus
OpenAIOpenAI
标准
$20/月GPT-5.6 Sol (max)61
图像PDFExcel语音
图像视频语音语音到语音
网页代码数据研究
记忆历史
9/104.7/6
iOSAndroidmacOSWindows
Claude Pro
AnthropicAnthropic
标准
$20/月Claude Opus 5 (max)63
图像PDFExcel语音
语音
网页代码数据研究
记忆历史
3/104.3/6
iOSAndroidmacOSWindows
Google AI Pro
GoogleGoogle
标准
$20/月Gemini 3.6 Flash52
图像PDFExcel视频语音
图像视频语音语音到语音
网页代码数据研究
记忆历史
5/104.5/6
iOSAndroid
Poe Pro
PoePoe
标准
$20/月
Claude Opus 4.6 (max)
45
图像PDFExcel语音
图像视频
网页
历史
0/102.0/6
iOSAndroidmacOSWindows
Perplexity Pro
PerplexityPerplexity
标准
$20/月
Gemini 3.1 Pro Preview
48
图像PDFExcel语音
图像视频语音
网页代码数据研究
历史
2/103.3/6
iOSAndroidmacOSWindows
SuperGrok
SpaceXAISpaceXAI
标准
$30/月Grok 4.5 (high)56
图像PDFExcel语音
图像视频语音
网页代码数据研究
记忆历史
2/103.8/6
iOSAndroid
Mistral Vibe Pro
MistralMistral
标准
$15/月Mistral Medium 3.530
图像PDFExcel语音
图像语音
网页代码数据研究
记忆历史
2/104.5/6
iOSAndroid

市场概览

Claude、ChatGPT、Gemini、Meta AI 等主要产品在推理、回答质量和功能方面各有所长。大多数产品都提供免费方案和付费方案(每月 15–25 美元),付费方案通常包含更大的上下文窗口、网页搜索、文件上传和图像生成功能。市场已集中在这些头部产品周围;在重视数据隐私的场景中,Llama 和 Mistral 等开源选项正获得更多采用。

差异化特点

  • 有些产品擅长编程(CodeStral、o1),另一些则在对话准确性(Claude)或视觉理解(Gemini)方面表现突出。
  • 产品差异主要来自长上下文、实时搜索、语音和多模态等专门能力,而非基础对话质量——后者已经大致趋同。

智能

Artificial Analysis 智能指数:所支持的最高智能模型

Artificial Analysis Intelligence Index · Higher is better

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

功能

智能指数 vs. 功能得分

Feature score · Higher is better · 6 categories tracked
Most attractive quadrant

A summary score counting the features offered by chatbots across six main categories (all listed in the full comparison table above):

  • Media Generation: 1 point (image generation: 0.25, video generation: 0.25, voice conversation: 0.25, native voice-to-voice: 0.25)
  • Tools: 1 point (web search: 0.25, code interpreter: 0.25, data analysis: 0.25, deep research: 0.25)
  • Input Capabilities: 1 point (image input: 0.20, PDF input: 0.20, Excel/CSV input: 0.20, video input: 0.20, voice input: 0.20)
  • Memory: 1 point (memory: 0.5, chat history: 0.5)
  • MCP Support: 1 point (Model Context Protocol integration: 1.0)
  • Connectors: 1 point (available integrations and connectors for external services and tools; each connector weighs 0.1 points; see the 10 connectors evaluated in the comparison table above)

The maximum total value for this metric is 6 points.

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

价格

智能指数 vs. 价格(付费套餐)

Artificial Analysis Intelligence Index · Monthly price (USD) · Standard & premium plans
Most attractive quadrant

These charts reveal the pricing structure and value proposition of different paid chatbot plans.

Key insights: The market shows a clear pricing hierarchy with three distinct tiers. Premium plans (SuperGrok Heavy, Google AI Ultra, Claude Max, Perplexity Max, ChatGPT Pro) are positioned at ~$200-300/month, offering top-tier capabilities. Standard plans cluster at much more affordable pricing ($15-30/month) with most options around $20/month, providing accessible AI capabilities for broader audiences.

Monthly Price: The monthly subscription cost in USD for standard and premium chatbot plans. Free plans are excluded from this analysis.

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

月费:付费聊天机器人套餐

Monthly price (USD) · Standard & premium plans

These charts reveal the pricing structure and value proposition of different paid chatbot plans.

Key insights: The market shows a clear pricing hierarchy with three distinct tiers. Premium plans (SuperGrok Heavy, Google AI Ultra, Claude Max, Perplexity Max, ChatGPT Pro) are positioned at ~$200-300/month, offering top-tier capabilities. Standard plans cluster at much more affordable pricing ($15-30/month) with most options around $20/month, providing accessible AI capabilities for broader audiences.

Monthly Price: The monthly subscription cost in USD for standard and premium chatbot plans. Free plans are excluded from this analysis.

功能得分 vs. 价格(付费套餐)

Feature score · Higher is better · 6 categories tracked · Monthly price (USD) · Standard & premium plans
Most attractive quadrant

This scatter plot reveals the feature value proposition of different paid chatbot plans by comparing their comprehensive feature offerings against monthly pricing.

Key insights: The market shows a clear pricing hierarchy with three distinct tiers. Premium plans (SuperGrok Heavy, Google AI Ultra, Claude Max, Perplexity Max, ChatGPT Pro) are positioned at ~$200-300/month, offering top-tier capabilities. Standard plans cluster at much more affordable pricing ($15-30/month) with most options around $20/month, providing accessible AI capabilities for broader audiences. Notably, some standard plans outperform premium alternatives on features, highlighting diverse value propositions across price points.

Monthly Price: The monthly subscription cost in USD for standard and premium chatbot plans. Free plans are excluded from this analysis.

常见问题

我们的比较使用基准数据,包括智能指数、功能得分、上下文窗口大小和价格。你可以按模型筛选、并排比较套餐,并查看智能与价格、智能与功能的图表,以找到最适合你使用场景的选择。

智能指数是基于 Artificial Analysis 基准测试的综合得分,用于衡量推理、知识和响应质量。得分越高,表示在标准化评估中的表现越强。付费套餐通常比免费等级得分更高。

大多数主流提供商(Claude、ChatGPT、Gemini、Meta AI)都提供有不同限制的免费等级。我们的比较表显示了每个套餐的功能得分、上下文窗口和能力。免费等级通常上下文窗口较小,且高级功能(如网页搜索或文件上传)较少。

关键差异包括长上下文支持、实时网页搜索、语音输入/输出、多模态(图像)理解和编程能力。有些在编程方面出色(CodeStral、o1),有些在对话准确性方面出色(Claude),有些在视觉理解方面出色(Gemini)。

Artificial Analysis 发布详细的 LLM 基准测试,包括延迟、成本和质量指标。我们的聊天机器人比较链接到模型级数据,以便进行更深入的分析。 查看 LLM 基准测试