AI OCR・文書処理ツールの比較
画像、PDF、スキャンに含まれるテキストを機械可読テキストに変換するエージェントです。多くの場合、OCR(光学文字認識)に文書の取り込み、検索、抽出、エクスポートなどのワークフローを組み合わせています。
言語モデルを比較するには、モデルベンチマークをご覧ください。
| 製品 | リリース | タイプ | オープンソース | 構造化抽出 | 手書き | 出力 | 価格 | 説明 | |
|---|---|---|---|---|---|---|---|---|---|
| Amazon Textract | May 2019 | API | JSON | 無料従量課金 $0.0015–0.05/pg | OCR API with specialist tools for expenses, IDs, and mortgage lending packages. Queries API enables natural-language questions about document content. No custom model training — relies entirely on prebuilt models. | ||||
| Azure Document Intelligence | Mar 2020 | APIアプリケーション | JSONMD | 無料従量課金 $0.0015–0.03/pg | Document processing service with prebuilt models for invoices, W-2s, insurance cards, bank statements, tax forms, and more. Custom Neural Models trainable with 5+ labeled samples; Composite Models combine multiple extractors under one endpoint. Available as on-premises containers. | ||||
| Google Document AI | Apr 2021 | API | JSON | 従量課金 $0.0015–0.03/pg | OCR API with handwriting recognition across 50+ languages and math formula detection. ~16 processor types across lending, procurement, and identity categories. Gemini-powered custom extraction trainable on labeled samples. | ||||
| Unstructured.io | Sep 2022 | APIアプリケーション | JSON | 無料従量課金 $0.03/pg | Document processing pipeline that converts 65+ file types into LLM-ready chunks — designed for RAG ingestion, not extracting named fields like invoice numbers. 30+ source and destination connectors (S3, Salesforce, Pinecone, etc.) move data from enterprise sources to vector databases. | ||||
| Mistral OCR | Mar 2025 | APIアプリケーション | MDHTMLJSON | 無料従量課金 $0.002/pg | Vision-language model OCR service, now on its third generation (OCR 3, Dec 2025). Structured extraction via Annotations with Pydantic/JSON schemas. European-hosted; batch mode at half price. | ||||
| ABBYY Vantage | Aug 2021 | APIアプリケーション | JSONXMLCSVPDFDOCXXLSXTXT | サブスクリプションエンタープライズ From ~$5,000/yr | Enterprise document processing platform with 150+ pre-trained extraction skills across finance, healthcare, logistics, and more. RPA integrations with UiPath, Blue Prism, and Automation Anywhere. On-premises deployment available. | ||||
| LlamaParse | Feb 2024 | API | MDTXTJSONXLSXPDF | 無料従量課金 $0.00125–0.06/pg | RAG-native parser with multimodal output — extracts text and image chunks optimized for LLM ingestion. Auto Mode routes pages to the cheapest tier that meets accuracy requirements. Part of the LlamaIndex ecosystem. | ||||
| Mathpix | Apr 2018 | APIアプリケーション | LaTeXMDDOCXHTMLPDF | 無料従量課金 $0.005/pg | STEM-focused OCR tool that extracts math equations, chemical structures, and scientific notation to LaTeX. Handles two-column journal layouts and inline/block equations. Snip app and Overleaf integration for academic workflows. | ||||
| Marker | Dec 2023 | API | MDJSONHTMLChunks | 無料従量課金 Free / $0.004/pg | Self-hostable pipeline built on sub-billion-parameter Surya models supporting 90+ languages. Runs on consumer GPUs; optional LLM hybrid mode (e.g., Gemini) improves accuracy on complex layouts. | ||||
| Nanonets | Jan 2017 | APIアプリケーション | JSONCSVMDTXTHTML | 無料従量課金 $0.02–0.30/run | End-to-end document workflow platform — OCR plus approval loops, ERP sync (NetSuite, SAP, QuickBooks), and AP/AR automation. Template-free extraction adapts to new vendor layouts without configuration. | ||||
| Reducto | Feb 2024 | APIアプリケーション | JSONMDHTMLCSV | 無料従量課金 $0.015/credit | Multi-pass pipeline with agentic self-correction — purpose-built for complex documents with charts, diagrams, and nested tables. SOC 2 Type II and HIPAA compliant with zero-retention processing. | ||||
| Upstage Document Parse | Oct 2024 | API | HTMLMD | 無料従量課金 $0.01–0.03/pg | Document parsing API with CJK language support (Korean-founded). Layout-aware HTML output preserving reading order at ~0.6 sec/page. Information Extract API (2025) adds structured field extraction. | ||||
| Docsumo | Jun 2019 | APIアプリケーション | JSONCSVExcel | サブスクリプションエンタープライズ Custom pricing | Financial services specialist with 100+ pre-trained models for lending, banking, and insurance documents. Auto-classification, completeness checking, and human-in-the-loop validation workflows. Custom model training from as few as 20 labeled samples. | ||||
| Rossum | Jan 2017 | APIアプリケーション | JSONXMLCSVXLSX | サブスクリプションエンタープライズ From $1,500/mo | Document automation platform for invoices, POs, and shipping docs. Powered by Aurora, a proprietary LLM trained on 11M transactional documents. Template-free extraction with three-way matching (PO/invoice/receipt) across 276 languages. |
市場の概要
この表では、OCRツールを、種類(APIまたはアプリケーション)、オープンソースかどうか、フィールド抽出、手書き文字認識、出力形式、料金で比較します。ほとんどのツールはAPIベースで、Azure Document Intelligence、Mistral OCR、ABBYY Vantage、Nanonets、Reducto、Docsumo、Rossumなどはアプリケーションインターフェースも提供しています。オープンソースの選択肢には、完全にオープンソースのMarkerと、一部がオープンソースのUnstructured.io、Nanonets、Reductoがあります。ほとんどのツールがフィールド抽出に対応しています。LlamaParse、Unstructured.io、Upstageは一部対応し、Mathpixは対応していません。手書き文字認識への対応はさまざまです。Amazon Textract、Azure、Google Document AI、Mistral OCR、ABBYY、LlamaParse、Mathpix、Reducto、Docsumo、Rossumは完全対応し、その他は一部対応または非対応です。料金は、無料でセルフホストできるMarkerやページ単位のAPI料金(1ページあたり$0.001~0.06)から、年間契約のエンタープライズサブスクリプション(Rossumは月額$1,500以上、ABBYYは年額$5,000以上)までさまざまです。
よくある質問
AIのOCRツールは、文書・画像・手書きからテキスト、表、構造化データを抽出します。ビジョンAPIでPDFを読み取れる一般的なチャットボットとは異なり、これらのツールはレイアウト検出、表認識、フィールド抽出に特化したモデルを備え、文書処理専用に作られています。
Markerは完全にオープンソースでセルフホスト可能です。Unstructured.io、Nanonets、Reductoは一部がオープンソース(オープンソースのコアまたはモデルの重みと、独自プラットフォームの組み合わせ)です。その他はクローズドソースです。表の「オープンソース」列をご確認ください。
手書きへの完全対応は、Amazon Textract、Azure Document Intelligence、Google Document AI、Mistral OCR、ABBYY Vantage、LlamaParse、Mathpix、Reducto、Docsumo、Rossum で利用できます。Marker、Nanonets、Unstructured.io は一部対応です。Upstage Document Parse は手書きに非対応です。表の「手書き」列をご確認ください。
ほとんどのツールは請求書、領収書、フォームなどの文書のフィールド抽出に対応しています。LlamaParse、Unstructured.io、Upstage は一部対応です。Mathpix は非対応で、STEMコンテンツ(数式、図、科学的表記)に特化しています。表の「フィールド抽出」列をご確認ください。
タイプ(APIかアプリケーションか)、オープンソースの状況、フィールド抽出、手書き認識、出力形式、価格でツールを比較しています。表は定期的に更新されます。 LLMベンチマークを見る