Unstructured

Document parsing and ETL for LLMs: 60+ file types into RAG-ready elements, with an Apache-2.0 core and per-page cloud pricing.

Best for: Teams ingesting many file types from enterprise sources into vector stores, where element semantics beat raw text

Pros

  • 60+ file types through one pipeline, from PDFs to Slack exports
  • Element-typed output (Title, NarrativeText, Table, ListItem) preserves document semantics for chunking
  • Apache-2.0 open-source core, 15,515 GitHub stars, latest release September 27, 2026
  • Simple money: 10,000 free pages, then $0.015 per page with all features included
  • 30+ source and 30+ destination connectors for warehouse-grade ETL

Cons

  • Slowest of three parsers in an independent benchmark: 51 to 141 seconds per document
  • 75% cell accuracy on a complex hierarchical table where a rival hit 97.9%
  • Table-of-contents extraction returned only the title in one independent test
  • Connector and file-type counts differ across the vendor's own pages
  • The vlm strategy is API-only; the open-source library gets fast and hi_res

Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.

Quick verdict: We recommend Unstructured for teams that need one pipeline turning 60+ file types from enterprise sources into RAG-ready, element-typed output. The trade-off: independent tests found it the slowest of three parsers in one and the weakest on complex tables in two, so test table-heavy work on the free tier first. Among document parsing and RAG frameworks, this Unstructured.io review prints the vendor’s benchmark and the two independent tests that contradict it, labeled.

Key Takeaways 🔍

  • Live pricing, September 2026: 10,000 free pages, then $0.015 per page, all features included; older articles print different figures
  • Apache-2.0 open-source core: 15,515 GitHub stars, 60+ file formats, last pushed September 27, 2026
  • Newest product: Foundation, an early-access MCP plugin giving coding agents enterprise file access
  • Con: independent benchmarks put it slowest of three parsers in one test (51 to 141 seconds per document) and third of three on tables in another
  • Con: the vendor’s own pages variously cite 30, 35, or 71 connectors
Unstructured homepage
Unstructured’s homepage. Source: Panoply.

Pros and Cons

Pros

  • 60+ file types through one pipeline, from PDFs to Slack exports
  • Element-typed output (Title, NarrativeText, Table, ListItem) preserves document semantics for chunking
  • Apache-2.0 open-source core, 15,515 GitHub stars, latest release September 27, 2026
  • Simple money: 10,000 free pages, then $0.015 per page with all features included
  • 30+ source and 30+ destination connectors for warehouse-grade ETL

Cons

  • Slowest of three parsers in an independent benchmark: 51 to 141 seconds per document
  • 75% cell accuracy on a complex hierarchical table where a rival hit 97.9%
  • Table-of-contents extraction returned only the title in one independent test
  • Connector and file-type counts differ across the vendor’s own pages
  • The vlm strategy is API-only; the open-source library gets fast and hi_res

How Much Does Unstructured Cost?

Unstructured pricing
Unstructured’s pricing plans. Source: Panoply.

Unstructured’s pricing page lists three tiers as of September 2026, and both self-serve tiers include every feature:

  • Free ($0): 10,000 pages, no card required, all features included
  • Pay-As-You-Go ($0.015/page): kicks in after the free 10,000 pages; all features; no monthly cap and no per-strategy price split published
  • Business (custom): dedicated instance, VPC or multi-tenant SaaS deployment, multi-user accounts, dedicated technical support
TierPriceDeploymentSupport
Free$0 for 10,000 pagesHosted APINo dedicated support tier listed
Pay-As-You-Go$0.015 per page after 10,000Hosted APINo dedicated support tier listed
BusinessCustomDedicated instance, VPC, or multi-tenant SaaSDedicated technical support

An August 2026 comparison article printed different figures: 15,000 free pages per month, $0.03 per page, and a $3,000 monthly cap. The live vendor page in September 2026 shows 10,000 free pages, $0.015 per page, and no cap. Either pricing changed between those dates or the article was wrong, so treat unstructured.io/pricing as the source of truth and re-check it before you commit.

The open-source library is the real $0 column: Apache-2.0 licensed and self-hosted, with the fast and hi_res strategies included. The vlm strategy is exposed only in the paid API. The same partition logic moves from the library to the hosted API or the Business platform without a rewrite, so starting free costs you no rework later.

Is Unstructured Good Value for Money?

  • Twelve times LlamaParse’s fast tier, roughly level with its agentic modes: LlamaParse’s fast tier runs around $0.00125 per page, and its agentic modes cost 10x to 45x that in credits; our LlamaParse review maps that menu
  • The volume math is blunt: 100,000 pages costs $1,350 on Pay-As-You-Go after the free 10,000, against about $125 at LlamaParse’s fast-tier rate
  • Docling is free but self-run: zero license cost; you supply the compute and the ops
  • The real cost variable is strategy choice: hi_res took 51 to 141 seconds per document in an independent benchmark, which is GPU or CPU time on your hardware if you self-host and wait time in your pipeline if you use the API
  • $0.015 buys every feature: nothing parsing-related sits behind a higher tier; Business adds deployment options, multi-user accounts, and support

My plan pick: a RAG team ingesting under 100,000 pages a month belongs on Pay-As-You-Go after the free pages run out. Business only makes sense when your compliance team requires a dedicated instance or a VPC.

Author’s Testing Notes 📝

Burn the free 10,000 pages on your own document mix, with both fast and hi_res, before you believe anyone’s benchmark, ours included. Ten thousand pages is enough to parse your ugliest PDFs several times over and check the table output by hand.

— Panoply reviewer

My Experience With Unstructured

I approached Unstructured the way its documentation steers you: install the library, partition one document, inspect the elements, then widen out to the pipeline.

Partitioning a Document

Setup is a pip install of the unstructured package, and the core of the product is one function. partition() auto-detects the file type; for PDFs I used partition_pdf from unstructured.partition.pdf. On a clean, text-native PDF, partition_pdf(filename=”report.pdf”, strategy=”fast”, extract_images=False) is the documented quick path.

What comes back is the product’s real differentiator: a list of typed elements (Title, NarrativeText, Table, ListItem) rather than a wall of text, each carrying metadata about where it came from. Run a scanned or mixed-layout file with strategy=”hi_res” and infer_table_structure=True and tables are preserved as structured HTML inside the element metadata instead of being flattened into prose.

Choosing a Strategy

Three strategies behave like three different tools:

  • fast: pdfminer-based, seconds per document, built for simple digital PDFs
  • hi_res: Detectron2/Chipper layout models, slower, better on tables and figures, GPU-optional
  • vlm: vision-language models for image-heavy documents and screenshots, exposed only in the paid API

For financial filings, one practitioner’s published configuration is hi_res with table inference on, chunked by title so tables never mix with narrative text; its author self-reported 40% fewer hallucinated KPIs versus a plain pdfminer pipeline, a figure we have not audited. There is also a documented Docker path: pull unstructured-api:latest, run it on port 8000, and POST files with parameters like chunking_strategy=by_title and combine_under_n_chars=500.

The ETL Frame

The ingest CLI is where Unstructured stops looking like a parser and starts looking like a data pipeline. unstructured-ingest connects sources (S3, SharePoint, Google Drive, Slack, Airtable) to destinations (vector databases and warehouses), with flags for parallelism and reprocessing. The documented Airtable example goes from a personal access token to partitioned JSON in one command. Categorize it accordingly: an ingestion pipeline for LLM stacks, not just a PDF parser.

Author’s Testing Notes 📝

The strategy flag is the whole speed-versus-fidelity dial, and undisclosed strategy choices are, in our read, why published Unstructured benchmarks disagree so widely. Before you trust any number in this review or anywhere else, run fast and hi_res on your own corpus and compare the output.

— Panoply reviewer

The Benchmark Contradiction, Printed Whole

Unstructured’s own benchmark page says it wins. Two independent tests put it last: on speed in one, on tables in the other.

The vendor-published case: unstructured.io/benchmarks tests Unstructured against Reducto, LlamaParse, Docling, Snowflake, Databricks, NVIDIA Nemotron-Parse, and off-the-shelf frontier models on a 1,000-plus-page dataset of scanned invoices, complex layouts, nested tables, and handwritten notes. Unstructured’s best configuration, the VLM partitioner paired with GPT-5-mini, scored 0.883 on adjusted CCT, a 0.036 hallucinated-token rate, and roughly 0.82 on table cell accuracy, against competitor top scores of 0.625 to 0.835. The caveat: Unstructured chose the dataset, the metrics, and its own best configuration to feature. These are vendor-claimed wins, not independently reproduced.

One independent benchmark, run on five corporate sustainability PDFs, measured Unstructured at 75% cell accuracy on a three-level hierarchical table where Docling hit 97.9%, and called its table-of-contents output severely deficient (the title came back, no entries, no page numbers). The same test clocked it slowest of the three tools: 51.06 seconds for a one-page document, stretching to 141.02 seconds at 50 pages, against Docling’s 6.28 to 65.12 seconds and LlamaParse’s roughly 6 seconds regardless of length. Its authors did rate Unstructured’s text extraction and simple-table handling adequate.

A second independent test, on multi-column scientific papers, financial statements, and business memos, ranked Unstructured third of three on table completeness, accuracy, and structure, calling it “more like a generic document scanner than a table-first parser.”

Here is our read on how both can be true. The vendor benchmarked its paid VLM configuration; the independent tests most likely ran the default or hi_res paths, and did not always disclose the strategy. Our conclusion: out-of-the-box table extraction lags Docling and LlamaParse, while the paid VLM path may not. That question is exactly what the free 10,000 pages exist to answer on your own documents.

Engineers on Hacker News add a third view: cloud OCR platforms (AWS Textract, Azure, Google) still lead for true structured extraction of forms and tables, and tools like Unstructured shine at segmenting text for RAG rather than defining document structure.

Foundation: The MCP Play

Unstructured’s newest product is not a parser at all. Foundation, now in early access, is an MCP (Model Context Protocol) plugin that lets coding agents (the vendor names Claude, Cursor, and Codex) read 65+ file types across SharePoint, Google Drive, Slack, S3, and databases, returning targeted context instead of flooding the model’s window with whole files.

The headline claims are vendor-stated, and we do not endorse them: 14x higher retrieval accuracy, 85% fewer agent turns, 10x lower costs. No independent test of Foundation is publicly documented yet, so treat those numbers as a pitch, not a result.

Unstructured does not market branded tiers like a Lite or an Enterprise edition. It presents one platform reached through API, UI, and now MCP, deployable bare metal, SaaS, or in-VPC. The scaling figure attached to that platform, over 15 million pages per hour per workflow, is likewise vendor-stated.

Strategically, Foundation points the same ETL engine at the agent market, and if your team runs Claude or Cursor against enterprise document stores, it is the one piece of Unstructured with no direct rival in this review. Practically, it is early access, and early access means evaluate it, not standardize on it.

Adoption, Momentum, and the Counting Problem

The Apache-2.0 repository is actively maintained: 15,515 GitHub stars, 1,349 forks, and 319 open issues, with the last push on September 27, 2026, two days before our check. PyPI version 0.27.10 shipped the same day and supports Python 3.11 through 3.13.

The vendor’s own spec sheets disagree on counts. Its docs cite 35+ connectors, its product page cites 30+ sources plus 30+ destinations, and one of its own articles claims over 71 pre-built connectors; file-type counts drift between 60+ and 65+. Our best explanation is that the biggest number bundles destinations and LLM and vector-database integrations in with source connectors. None of this is disqualifying, but when a vendor’s pages disagree with each other, verify the one connector you actually need before you commit.

A GitHub issue (#2557) reporting that both fast and hi_res ran slower than expected on larger files was closed back in May 2024, yet the 2025 independent benchmark above measured the same slowness. That benchmark’s authors named speed and complex-table parsing as the two areas Unstructured most needs to improve. Treat hi_res latency as a persistent characteristic of the tool, not a bug that will patch itself away.

How Does Unstructured Compare to Competitors?

Every rival in this category wins somewhere specific, and two of them beat Unstructured outright in the independent tests above.

  • Docling had the best overall accuracy and table-structure preservation in both independent benchmarks: 97.9% table-cell accuracy and a 100% sentence match in one test. It is free, has IBM Research lineage, and runs at a moderate, predictable speed that scales linearly with page count. Its noted gap is flat Markdown heading levels, which lose parent-child section hierarchy. Pick it for maximum out-of-the-box table fidelity at zero license cost, if you are willing to run it yourself.
  • LlamaParse was the fastest tool tested, roughly 6 seconds per document regardless of length, with the best row and column structure preservation in one benchmark and a fast tier around $0.00125 per page. Its failure mode is systematic misplacement in complex tables (100% of the data extracted, 0% placed correctly in one test), and it also stumbles on multi-column layouts, word merging, currency symbols, and footnotes. Its credit pricing jumps 10x to 45x across modes; our LlamaParse review covers that menu.
  • Cloud OCR (AWS Textract, Azure Document Intelligence, Google) is the practitioner pick for true structured extraction. Engineers on Hacker News report the cloud platforms outperform both Unstructured and LlamaParse on forms and complex tables, and position Unstructured as a RAG segmentation tool first, not a structure-definition tool.
  • Reducto appears only as a comparator (in its Agentic mode) in Unstructured’s own benchmark. No independent benchmark of it is publicly documented, so if it is on your shortlist, run it on the same documents you put through Unstructured’s free tier.

The decision rule: many file types plus enterprise connectors plus RAG-ready elements means Unstructured; maximum table fidelity out of the box means Docling; speed and a LlamaIndex-native flow means LlamaParse; forms and high-stakes structure mean cloud OCR.

How We Test Document Parsing Tools

We combine independent analysis, data collection, and hands-on testing: we set the tool up ourselves, run a real task end to end, and compare it against its closest rivals. Data collection covers public signals (GitHub, PyPI, Docker Hub, Stack Overflow, G2, and Gartner peer reviews), refreshed monthly, with editorial verdicts revisited quarterly. Vendor pricing pages are hand-checked and dated rather than quoted from third parties.

Aggregate review scores are deliberately small-weighted, because sustained practitioner sentiment tells us more than star averages, and where an area cannot be measured we mark it N/A rather than scoring it zero. Sponsors and affiliates cannot change a score. Prices current as of September 2026.

Unstructured Review: Should You Feed Your RAG Stack With Unstructured?

We recommend Unstructured as the ingestion backbone for teams pulling many file types out of enterprise sources and into vector stores. If your pipeline touches PDFs, Slack exports, SharePoint libraries, and S3 buckets in the same week, the 60+ formats, the 30+ connectors, and element-typed output that keeps tables separate from narrative text are the practical argument. The Apache-2.0 core is real, actively maintained, and a legitimate exit path. Foundation extends the same engine to coding agents, though as an early-access product we treat it as a bonus, not a reason to buy.

Test before you buy if tables are the job. Two independent benchmarks put out-of-the-box table extraction behind Docling, and hi_res took 51 to 141 seconds per document in one of them, a pace that will hurt latency-sensitive pipelines. Skip it for single-format PDF-to-table jobs, where Docling or a cloud OCR service wins on the independent evidence.

The next action costs nothing: run the free 10,000 pages against your ugliest documents, fast versus hi_res, and check the extracted tables by hand. The benchmark that matters is yours.

FAQ

Is Unstructured free?

Partly. The open-source library is Apache-2.0 licensed and free to self-host, with the fast and hi_res strategies included. The hosted API adds 10,000 free pages with all features and no card required, then bills $0.015 per page. The vlm strategy is only available through the paid API.

Is Unstructured accurate for tables?

It depends on whose test you trust and which strategy you run. The vendor’s own benchmark, on a dataset and metrics it chose, reports roughly 0.82 table cell accuracy for its best paid VLM configuration. Two independent benchmarks tell a different story: 75% cell accuracy on a complex hierarchical table versus Docling’s 97.9%, and a last-of-three table ranking. Test both strategies on your own tables.

What is the difference between Unstructured and LlamaParse?

Unstructured is a document ETL platform: connectors, partitioning, chunking, enrichment, and destinations in one pipeline. LlamaParse is a parse-first API with credit-priced quality tiers, built for the LlamaIndex ecosystem. LlamaParse was faster in independent testing (roughly 6 seconds per document), and its fast tier costs around $0.00125 per page; Unstructured covers far more file types and enterprise sources. Our LlamaParse review covers its mode pricing.

What is Unstructured Foundation?

Foundation is an early-access MCP plugin that lets coding agents such as Claude, Cursor, and Codex read 65+ file types across SharePoint, Google Drive, Slack, S3, and databases. The vendor claims 14x retrieval accuracy, 85% fewer agent turns, and 10x lower costs; none of those figures is independently verified.

Why do published Unstructured prices differ?

Either pricing changed or older articles were wrong. An August 2026 comparison printed 15,000 free pages at $0.03 per page with a $3,000 monthly cap; the live vendor page in September 2026 shows 10,000 free pages at $0.015 per page with no stated cap. Check unstructured.io/pricing at decision time.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.

Similar tools

Other tools in the same category, with the same card and the same honest pricing.

Contact Databricks

Tecton

Feature store

Feature platform for production machine learning, with batch, streaming, and request-time features. Acquired by Databricks in 2025.

Visit site

Helicone

Observability

Proxy-based LLM observability gateway with one-line setup; hosted product in maintenance mode since the March 2026 Mintlify acquisition.

Visit site

Arize AX

Observability

Enterprise LLM/ML observability (AX) with Phoenix, its ungated open-source tracing and evals core; deep drift and embeddings analysis.

Visit site

LlamaParse

Document parsing

LlamaIndex's managed parser for complex PDFs, tables, and scans; credit-priced by mode from $0.00125 to $0.056 per page.

Visit site

ClearML

ML platform

End-to-end open-source MLOps: experiment tracking, GPU orchestration with fractional GPUs, dataset versioning, and pipelines at $15/user.

Visit site

Milvus

Dedicated

Open-source, billion-scale vector database under LF AI & Data; 3.0 indexes lakehouse data in place, with DiskANN and GPU CAGRA indexes.

Visit site

Zilliz Cloud

Dedicated

Fully managed Milvus from its commercial steward, with AutoIndex tuning, the Cardinal engine, and post-2026 storage pricing at $0.04/GB/month.

Visit site

Qdrant

Dedicated

Open-source Rust vector database with in-graph payload filtering, three quantization families, and a free forever cloud tier.

Visit site

Haystack

RAG framework

deepset's open-source RAG framework: typed component pipelines, YAML serialization, and enterprise connectors under Apache-2.0.

Visit site