pdfmux

Self-healing PDF extraction that flags the pages it can't read instead of dropping them β and now certifies any extractor's output for silent drops. Open-source LlamaParse alternative for RAG pipelines, MCP server for Claude Desktop, LangChain + LlamaIndex loaders.
pdfmux extracts PDFs and checks its own work β and now certifies any extractor's, telling you which pages it silently dropped. Free, MIT. Patent-pending method. pip install pdfmux.
Two jobs, one tool:
- Self-healing extraction. The only PDF extractor that audits its own output. Catches blank pages, scrambled columns, broken tables β re-extracts them with a stronger backend, and flags what it still can't read instead of silently dropping it. So your LLM gets clean data, not silent garbage. Routes each page to the best of 7 built-in extraction backends + BYOK LLM fallback (Gemini / Claude / GPT-4o / Ollama). One CLI. One API. Zero config.
- Certify Anything β new in v1.8.1.
pdfmux verify audits any extraction engine's output against the source PDF β Reducto, Mistral OCR, LlamaParse, Docling, your in-house parser β and tells you which pages it silently dropped. Free, MIT, patent-clean.
PDF ββ> pdfmux router ββ> best extractor per page ββ> audit ββ> re-extract failures ββ> Markdown / JSON / chunks
|
ββ PyMuPDF (digital text, 0.01s/page)
ββ OpenDataLoader (complex layouts, 0.05s/page)
ββ RapidOCR (scanned pages, CPU-only)
ββ Docling (tables, 97.9% TEDS)
ββ Surya (heavy OCR fallback)
ββ Marker (academic papers, neural)
ββ Mistral OCR ($0.002/page, 96.6% tables)
ββ YOUR LLM (Gemini / Gemma 3 / Claude / GPT-4o / Ollama / Mistral β BYOK via YAML)
Install
That handles digital PDFs. For any real-world batch, install pdfmux[ocr] too β almost every directory of PDFs has at least one scan, and without OCR those pages return empty text:
pip install "pdfmux[ocr]"
Other backends, by document type:
pip install "pdfmux[tables]"
pip install "pdfmux[opendataloader]"
pip install "pdfmux[marker]"
pip install "pdfmux[llm]"
pip install "pdfmux[llm-claude]"
pip install "pdfmux[llm-openai]"
pip install "pdfmux[llm-ollama]"
pip install "pdfmux[llm-mistral]"
pip install "pdfmux[llm-all]"
pip install "pdfmux[watch]"
pip install "pdfmux[all]"
Requires Python 3.11+.
Quick Start
CLI
pdfmux convert invoice.pdf
pdfmux convert report.pdf --chunk --max-tokens 500
pdfmux convert report.pdf --mode economy --budget 0.50
pdfmux convert invoice.pdf --schema invoice
pdfmux convert scan.pdf --llm-provider claude
pdfmux convert invoice.pdf --profile invoices
pdfmux estimate big-report.pdf --llm-provider gemini
pdfmux stream report.pdf --quality high
pdfmux watch ./inbox/ -o ./output/
pdfmux diff old.pdf new.pdf
pdfmux convert ./docs/ -o ./output/
pdfmux convert ./docs/ -o ./output/ --strict --min-confidence 0.20
pdfmux doctor --check ./docs/
pdfmux convert report.pdf --no-cache
pdfmux convert report.pdf --clear-cache
Python
For batch processing, use batch_extract() β not a subprocess.run(['pdfmux', ...]) loop. Same pipeline, no per-file process spawn, handles non-ASCII filenames:
import pdfmux
from pathlib import Path
pdfs = list(Path("./inbox").glob("*.pdf"))
for path, result in pdfmux.batch_extract(pdfs, quality="standard"):
if isinstance(result, Exception):
print(f"FAILED {path.name}: {result}")
continue
if result.confidence < 0.50:
print(f"REVIEW {path.name} ({result.confidence:.2f})")
else:
print(f"OK {path.name} ({result.confidence:.2f})")
text = pdfmux.extract_text("report.pdf")
data = pdfmux.extract_json("report.pdf")
chunks = pdfmux.chunk("report.pdf", max_tokens=500)
Don't wrap pdfmux with your own pypdf/pdfplumber fallback. pdfmux already routes per page through PyMuPDF β RapidOCR β vision LLM. PyMuPDF tolerates malformed PDFs that pypdf rejects ("Stream has ended unexpectedly"), so a downstream pypdf fallback turns recoverable PDFs into failures. Trust the router; check the confidence score on the result.
Certify Anything
pdfmux verify audits any extraction engine's output against the source PDF and tells you which pages it silently dropped β not just pdfmux's own extraction. Point it at the output of Reducto, Mistral OCR, LlamaParse, Docling, or your in-house parser and it re-derives the source text with pdfmux's own audit pass, aligns the extraction to it, and scores every page.
The failure it catches: a page where the source has real text but the engine returned nothing β while reporting success. That "silent drop" is the exact failure that poisons a RAG index without a single error in the logs.
pdfmux verify --source report.pdf --engine pdfmux
pdfmux verify --source report.pdf --extracted reducto.json --engine-name reducto
pdfmux verify --source ./pdfs/ --extracted ./engine-outputs/ -o certification.json
pdfmux verify --source report.pdf --extracted out.json --strict
Every run prints a PASS / REVIEW / FAIL verdict, overall confidence and coverage, and β when it finds them β the silently dropped pages by number:
pdfmux verify β report.pdf Β· engine: reducto
FAIL confidence 71% Β· coverage 68%
reducto: FAIL; 3 page(s) SILENTLY DROPPED (pages 7, 12, 31); overall
confidence 71%, coverage 68% across 40 page(s).
β 3 page(s) SILENTLY DROPPED: 7, 12, 31
Per page you get a verdict (pass / review / fail), confidence, coverage, alignment, hallucination-risk, and table/heading integrity. Batch mode rolls that up into a single "N pages silently dropped across M documents" line β the report you run on 100 of your own PDFs to find the silent failures already in your pipeline.
It works on any engine's output
--extracted accepts JSON, Markdown, or plain text (--extracted-format auto | json | markdown | text). When the extraction exposes real per-page structure, pdfmux compares page-by-page; when it's a single blob, it falls back to content-presence checks so it never fabricates a "silent drop" from a pagination mismatch.
Python API
from pdfmux import verify_extraction, verify_batch
manifest = verify_extraction("report.pdf", "reducto.json", engine="reducto")
print(manifest.verdict)
print(manifest.silent_drops)
print(manifest.coverage)
batch = verify_batch([("a.pdf", "a.json"), ("b.pdf", "b.json")], engine="llamaparse")
print(batch.total_silent_drops, "pages dropped across", batch.doc_count, "docs")
Each manifest carries a tamper-evident SHA-256 content signature over its canonical body and an embedded, honest limitations list: the certifier is lexical, not linguistic β it detects missing and garbled content, not faithful paraphrase or translation.
MCP
verify_extraction is exposed as an MCP tool (the 7th β see MCP Server), so an agent can certify an engine's output in the same session it extracts.
Free, MIT, patent-clean
Certify Anything reuses only pdfmux's shipped MIT audit layer. It does not include, and does not require, the patent-pending decision-trace method β that stays in pdfmux Cloud/Pro. pip install pdfmux gives you the full verify command at no cost.
Full reference: docs/CERTIFY-ANYTHING.md.
When you need to prove it to someone else
A local install can audit an extraction, but it cannot attest to one β anything it signs, anyone could forge. pdfmux Cloud returns an Ed25519-signed manifest over the extraction: your auditor verifies it offline, against a published public key, without an account and without trusting pdfmux.
pdfmux verify-manifest manifest.json
Verification is free and open forever; only generation is paid ($49/mo). That asymmetry is deliberate β you should never need our permission to check our work.
Free tool, no signup: app.pdfmux.com/audit β upload a PDF and see which pages your current extractor silently dropped. Measured accuracy (and its blind spots) published in pdfmux-bench.
Architecture
βββββββββββββββββββββββββββββββ
β Segment Detector β
β text / tables / images / β
β formulas / headers per page β
βββββββββββββββ¬ββββββββββββββββ
β
ββββββββββββββββββββββββββββββββββββββββββ
β Router Engine β
β β
β economy ββ balanced ββ premium β
β (minimize $) (default) (max quality)β
β budget caps: --budget 0.50 β
ββββββββββββββββββββββ¬ββββββββββββββββββββ
β
ββββββββββββ¬βββββββββββ¬βββββββββ΄βββββββββ¬βββββββββββ
β β β β β
PyMuPDF OpenData RapidOCR Docling LLM
digital Loader scanned tables (BYOK)
0.01s/pg complex CPU-only 97.9% any provider
layouts TEDS
β β β β β
ββββββββββββ΄βββββββββββ΄βββββββββ¬βββββββββ΄βββββββββββ
β
ββββββββββββββββββββββββββββββββββββββββββ
β Quality Auditor β
β β
β 4-signal dynamic confidence scoring β
β per-page: good / bad / empty β
β if bad -> re-extract with next backendβ
ββββββββββββββββββββββ¬ββββββββββββββββββββ
β
ββββββββββββββββββββββββββββββββββββββββββ
β Output Pipeline β
β β
β heading injection (font-size analysis)β
β table extraction + normalization β
β text cleanup + merge β
β confidence score (honest, not inflated)β
ββββββββββββββββββββββββββββββββββββββββββ
Key design decisions
- Router, not extractor. pdfmux does not compete with PyMuPDF or Docling. It picks the best one per page.
- Agentic multi-pass. Extract, audit confidence, re-extract failures with a stronger backend. Bad pages get retried automatically.
- Segment-level detection. Each page is classified by content type (text, tables, images, formulas, headers) before routing.
- 4-signal confidence. Dynamic quality scoring from character density, OCR noise ratio, table integrity, and heading structure. Not hardcoded thresholds.
- Document cache. Each PDF is opened once, not once per extractor. Shared across the full pipeline.
- Data flywheel. Local telemetry tracks which extractors win per document type. Routing improves with usage.
Features
| Feature | What it does | Command |
|---|
| Zero-config extraction | Routes to best backend automatically | pdfmux convert file.pdf |
| RAG chunking | Section-aware chunks with token estimates | pdfmux convert file.pdf --chunk --max-tokens 500 |
| Cost modes | economy / balanced / premium with budget caps | pdfmux convert file.pdf --mode economy --budget 0.50 |
| Schema extraction | 5 built-in presets (invoice, receipt, contract, resume, paper) | pdfmux convert file.pdf --schema invoice |
| Profiles | Save and re-use config; built-ins for invoices/receipts/papers/contracts/bulk-rag | pdfmux convert file.pdf --profile invoices |
| BYOK LLM | Gemini, Gemma 3, Claude, GPT-4o, Ollama, Mistral, any OpenAI-compatible API | pdfmux convert file.pdf --llm-provider claude |
| Cost estimate | Predict spend before running | pdfmux estimate file.pdf --llm-provider gemini |
| Streaming output | NDJSON events page-by-page for long docs | pdfmux stream file.pdf |
| Smart cache | Hash-keyed result cache, 30-day TTL, 1 GB LRU | pdfmux convert file.pdf (auto), --no-cache to bypass |
| Watch mode | Auto-convert any PDF added to a folder | pdfmux watch ./inbox/ |
| Diff | Compare two extractions | pdfmux diff a.pdf b.pdf |
| Benchmark | Eval all installed extractors against ground truth | pdfmux benchmark |
| Doctor | Show installed backends, coverage gaps, recommendations | pdfmux doctor |
| MCP server | AI agents read PDFs via stdio or HTTP | pdfmux serve |
| Batch processing | Convert entire directories | pdfmux convert ./docs/ |
| Page-level streaming API | Bounded-memory page iteration for large files | for page in ext.extract("500pg.pdf") |
| Retry with backoff | Every LLM provider auto-retries with exponential backoff + Retry-After | (built-in) |
CLI Reference
pdfmux convert
pdfmux convert <file-or-dir> [options]
Options:
-o, --output PATH Output file or directory
-f, --format FORMAT markdown | json | csv | llm (default: markdown)
-q, --quality QUALITY fast | standard | high (default: standard)
-s, --schema SCHEMA JSON schema file or preset (invoice, receipt, contract, resume, paper)
--chunk Output RAG-ready chunks
--max-tokens N Max tokens per chunk (default: 500)
--mode MODE economy | balanced | premium (default: balanced)
--budget AMOUNT Max spend per document in USD
--llm-provider PROVIDER LLM backend: gemini | claude | openai | ollama
--confidence Include confidence score in output
--stdout Print to stdout instead of file
pdfmux serve
Start the MCP server for AI agent integration.
pdfmux serve
pdfmux serve --http 8080
pdfmux doctor
pdfmux benchmark
pdfmux benchmark report.pdf
pdfmux estimate
Predict spend (and which backends will run) before processing.
pdfmux estimate report.pdf --quality high --llm-provider gemini
pdfmux stream
Emit NDJSON events as pages complete β useful for very long PDFs and live UIs.
pdfmux stream long.pdf --quality high
pdfmux watch
Auto-convert any PDFs that land in a directory. Survives until Ctrl+C.
pdfmux watch ./inbox/ -o ./output/ --profile bulk-rag
pdfmux diff
Side-by-side extraction comparison (quality, content, cost).
pdfmux diff a.pdf b.pdf --quality standard
pdfmux profiles
Saved configs at ~/.config/pdfmux/profiles.yaml. Built-ins ship for the
common shapes; save your own for project defaults.
pdfmux profiles list
pdfmux profiles show invoices
pdfmux profiles save my-default --quality high --format llm --chunk
pdfmux profiles delete my-default
pdfmux convert file.pdf --profile invoices
Python API
import pdfmux
text = pdfmux.extract_text("report.pdf")
text = pdfmux.extract_text("report.pdf", quality="fast")
text = pdfmux.extract_text("report.pdf", quality="high")
data = pdfmux.extract_json("report.pdf")
RAG chunking
chunks = pdfmux.chunk("report.pdf", max_tokens=500)
for c in chunks:
print(f"{c['title']}: {c['tokens']} tokens (pages {c['page_start']}-{c['page_end']})")
data = pdfmux.extract_json("invoice.pdf", schema="invoice")
Streaming (bounded memory)
from pdfmux.extractors import get_extractor
ext = get_extractor("fast")
for page in ext.extract("large-500-pages.pdf"):
process(page.text)
Types and errors
from pdfmux import (
Quality,
OutputFormat,
PageQuality,
PageResult,
DocumentResult,
Chunk,
PdfmuxError,
FileError,
ExtractionError,
ExtractorNotAvailable,
FormatError,
AuditError,
)
Framework Integrations
LangChain
pip install langchain-pdfmux
from langchain_pdfmux import PDFMuxLoader
loader = PDFMuxLoader("report.pdf", quality="standard")
docs = loader.load()
LlamaIndex
pip install llama-index-readers-pdfmux
from llama_index.readers.pdfmux import PDFMuxReader
reader = PDFMuxReader(quality="standard")
docs = reader.load_data("report.pdf")
MCP Server (AI Agents)
Listed on mcpservers.org. One-line setup:
{
"mcpServers": {
"pdfmux": {
"command": "npx",
"args": ["-y", "pdfmux-mcp"]
}
}
}
Or via Claude Code:
claude mcp add pdfmux -- npx -y pdfmux-mcp
Tools exposed: convert_pdf, analyze_pdf, extract_structured,
extract_streaming, get_pdf_metadata, batch_convert.
BYOK LLM Configuration
pdfmux supports any LLM via 5 lines of YAML. Bring your own keys -- nothing leaves your machine unless you configure it to.
provider: claude
model: claude-sonnet-4-20250514
api_key: ${ANTHROPIC_API_KEY}
base_url: https://api.anthropic.com
max_cost_per_page: 0.02
Supported providers:
| Provider | Models | Local? | Cost |
|---|
| Gemini | 2.5 Flash, 2.5 Pro | No | ~$0.01/page |
| Gemma 3 | 27B IT, 12B IT (great for Arabic) | No (via Gemini key) | ~$0.0002/page |
| Claude | Sonnet, Opus | No | ~$0.015/page |
| GPT-4o | GPT-4o, GPT-4o-mini | No | ~$0.01/page |
| Mistral | mistral-ocr-latest | No | $0.002/page |
| Ollama | Any local model | Yes | Free |
| Custom | Any OpenAI-compatible API | Configurable | Varies |
Every provider's extract_page() is wrapped in @with_retry(max_attempts=3, backoff_base=2.0), which honors Retry-After headers on 429s and skips
retries on auth failures so a bad key fails fast.
Arabic & RTL Support
pdfmux ships first-class support for Arabic, Persian, Urdu, and Hebrew.
Out of the box, RTL detection runs on every PDF and PyMuPDF-extracted
pages are passed through the Unicode Bidirectional Algorithm so glyphs
that were stored in left-to-right order render in correct reading order.
pip install pdfmux
pip install "pdfmux[llm-openai]"
export GEMINI_API_KEY=...
What happens automatically:
pdfmux convert detects Arabic content and routes pages with >5%
Arabic characters through the Arabic-aware extractor chain.
- PyMuPDF, RapidOCR, and Docling outputs are post-processed with the
Bidi algorithm β markdown headings (
#) and pipe-table rows preserve
structure, only inner text is reordered.
DocumentResult.has_arabic is set to True whenever any page contains
Arabic script.
What requires opt-in:
- Vision LLM extraction. Set
--llm-provider gemma (or any vision
provider) to route Arabic pages through Gemma instead of PyMuPDF.
- Aggressive normalization (Tatweel removal, Alef/Yeh unification,
Tashkeel stripping) β call
pdfmux.arabic.normalize_arabic(text)
on extracted strings if you need canonicalized output for search or
embedding.
from pdfmux.arabic import (
is_arabic_text,
is_rtl_dominant,
fix_bidi_order,
normalize_arabic,
)
text = "Ω
Ψ±ΨΨ¨Ψ§ Ψ¨Ψ§ΩΨΉΨ§ΩΩ
"
assert is_arabic_text(text)
assert is_rtl_dominant(text)
visual = fix_bidi_order(text)
indexable = normalize_arabic("Ψ£ΩΨΩΩ
ΩΨ―Ω")
Proof: a real customer batch
We measured pdfmux on 433 real customer documents β technical and safety data sheets, mixed digital and scanned, some encoding-corrupted. Run the naive way first (an early pdfmux CLI in a subprocess, pypdf fallback, no OCR), the pipeline silently dropped 16 documents β 11 of them with no log line at all. That was our own tool failing at the exact thing it promises.
Rebuilt with the per-page audit + budgeted OCR cascade: 433 of 433 processed, zero silent failures. Every unrecoverable page is flagged, not dropped.
(A small internal confidence-calibration set also ships under eval/ β it's a regression guard on the confidence gate, not a competitive benchmark; see eval/README.md.)
Benchmark
On opendataloader-bench β 200 real-world PDFs (financial filings, academic papers, legal contracts, government reports) β pdfmux scores 0.903 overall β #2 of the 8 engines measured, behind opendataloader-hybrid (0.909). Re-run 2026-07-16 (reproduction below).
| Rank | Engine | Overall | Reading order | Tables (TEDS) | License | GPU |
|---|
| 1 | opendataloader-hybrid | 0.909 | 0.935 | 0.928 | Apache-2.0 | No |
| 2 | pdfmux | 0.903 | 0.920 | 0.911 | MIT | No |
| 3 | Docling | 0.877 | 0.900 | 0.887 | MIT | Optional |
| 4 | marker | 0.861 | 0.890 | 0.808 | free | GPU |
| 5 | mineru | 0.831 | 0.857 | 0.873 | free | GPU |
Full per-document scores: the 200-PDF head-to-head Β· methodology: best PDF extraction library, benchmarked.
Smart Result Cache
Re-running the same extraction is instant. pdfmux hashes every input PDF
(SHA-256) and keys results on (file_hash, quality, format, schema). Cache
files live under ~/.cache/pdfmux/results/, expire after 30 days, and are
LRU-evicted at 1 GB.
pdfmux convert big-report.pdf
pdfmux convert big-report.pdf
pdfmux convert big-report.pdf --no-cache
pdfmux convert big-report.pdf --clear-cache
The cache also speeds up --profile, --schema, and --format switches β
each combination is keyed independently, so you can flip between Markdown
and JSON for the same document for free after the first extraction.
Confidence Scoring
Every result includes a 4-signal confidence score:
- 95-100% -- clean digital text, fully extractable
- 80-95% -- good extraction, minor OCR noise on some pages
- 50-80% -- partial extraction, some pages unrecoverable
- <50% -- significant content missing, warnings included
When confidence drops below 80%, pdfmux tells you exactly what went wrong and how to fix it:
Page 4: 32% confidence. 0 chars extracted from image-heavy page.
-> Install pdfmux[ocr] for RapidOCR support on 6 image-heavy pages.
Cost Modes
| Mode | Behavior | Typical cost |
|---|
| economy | Rule-based backends only. No LLM calls. | $0/page |
| balanced | LLM only for pages that fail rule-based extraction. | ~$0.002/page avg |
| premium | LLM on every page for maximum quality. | ~$0.01/page |
Set a hard budget cap: --budget 0.50 stops LLM calls when spend reaches $0.50 per document.
Why pdfmux?
pdfmux is not another PDF extractor. It is the orchestration layer that picks the right extractor per page, verifies the result, and retries failures.
| Tool | Good at | Limitation |
|---|
| PyMuPDF | Fast digital text | Cannot handle scans or image layouts |
| Docling | Tables (97.9% accuracy) | Slow on non-table documents |
| Marker | Neural extraction for academic papers | Needs GPU for speed; overkill for digital PDFs |
| Mistral OCR | Tables (96.6% TEDS), $0.002/page | Cloud-only API |
| Unstructured | Enterprise platform | Complex setup, paid tiers |
| LlamaParse | Cloud-native | Requires API keys, not local |
| Reducto | High accuracy | $0.015/page, closed source |
| pdfmux | Orchestrates all of the above | Routes per page, audits, re-extracts |
Open source Reducto alternative: what costs $0.015/page elsewhere is free with pdfmux's rule-based backends, or ~$0.002/page average with BYOK LLM fallback.
Development
git clone https://github.com/NameetP/pdfmux.git
cd pdfmux
python3.12 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest
ruff check src/ tests/
ruff format src/ tests/
Contributing
- Fork the repo
- Create a branch (
git checkout -b feature/your-feature)
- Write tests for new functionality
- Ensure
pytest and ruff check pass
- Open a PR
License
The pdfmux library and MCP server in this repository are MIT licensed β free for any use, and every released version stays MIT.
The confidence-budgeted decision-trace method (the persisted per-page decision trace with retained rejected candidates, and the monotonic repair guard) is patent-pending (US Provisional App No. 64/106,302) and is reserved for pdfmux Cloud/Pro under a separate commercial license β it is not part of the MIT grant. See LICENSING.md and NOTICE.