The io.github.jztan/pdf-mcp MCP server provides agentic RAG over one PDF or a whole folder, supporting hybrid search, selective page reads, and extraction of tables and OCR content. It is positioned for document processing workflows using tools and libraries related to PDF parsing and semantic search.
🛠️ Key Features
Hybrid search across PDF content
Selective page reads
Table extraction
OCR support
Agentic RAG for single PDFs or folders
🚀 Use Cases
Querying a single PDF with semantic and hybrid retrieval
Building RAG over a directory of PDFs
Extracting structured table data and OCR text for downstream LLM use
⚡ Developer Benefits
Supports MCP integration for document retrieval
Python-focused tooling (topics include python, pymupdf, and model-context-protocol)
Keyword-oriented organization for developers (topics: semantic-search, pdf-extraction, table-extraction, agentic-rag)
⚠️ Limitations
Scope is centered on PDF inputs (one file or a folder), including OCR and table handling
Agentic RAG over your PDFs, one file or a whole folder, as a single MCP tool.
The agent decides when to search; pdf-mcp does the retrieval and hands back excerpts. It is an MCP server that lets Claude Code and other AI agents search one PDF or a whole folder by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts, with optional CUDA acceleration for warming large corpora.
Drop in any PDF, or a whole folder of them, and watch an agent triage the corpus, search across every document at once, and read only the pages that matter, using a fraction of the tokens. 100% client-side, no install required.
In Claude Desktop, open Settings > Extensions, drag the file onto
that page, and click Install.
Ask Claude about a PDF by its location, for example "Use pdf-mcp to
summarize C:\Users\me\Downloads\report.pdf", or about a whole folder.
The first start downloads pdf-mcp's components (about 250 MB) and can take a
few minutes; later starts take seconds. OCR for scanned pages is included:
the first scanned page downloads an English-only Tesseract (about 14 MB).
Needs Windows 10 or later, or macOS 13 or later (14 on Apple Silicon), and
works in Claude Desktop's Chat. Updating, uninstalling and other details are
in docs/clients.md.
Claude Code and other MCP clients
Needs Python 3.10 or later. Install the pdf-mcp command with
uv or pipx:
Optional: CUDA embedding on an NVIDIA card warms large folders one to two
orders of magnitude faster; see
docs/configuration.md.
From Python
pdf-mcp's tools are also plain Python functions, so you can import them and
hand a PDF to the Anthropic SDK without running a server. Two runnable
scripts, for a question and for a whole document: examples/.
13 specialized tools rather than one monolithic one. Typical pattern:
pdf_info to plan, pdf_search to locate (its paragraph excerpts often
answer the question outright), pdf_read_pages when you need more. For a
folder, pdf_corpus_overview to triage, then pdf_corpus_search.
Hybrid search (keyword + semantic), page or section granularity, paragraph or context-window excerpts with source coordinates
pdf_read_pages
Read specific pages or ranges, with OCR on demand, tables, and embedded images
pdf_read_all
Read a whole document in one call, byte-capped
pdf_get_toc
Full table of contents for documents with many bookmarks
pdf_render_pages
Render pages as PNG for vision models: diagrams, handwriting, scans
pdf_extract_chart
Chart data as exact (x, y) tables, read from plot geometry
pdf_corpus_warm
Warm a folder of PDFs into the cache within a time budget
pdf_corpus_overview
Per-document triage cards for a folder
pdf_corpus_search
Search across a folder, with document and page provenance; excerpt_style="auto" picks the excerpt unit per query
pdf_cache_stats
Per-document cache breakdown and total size
pdf_cache_clear
Clear expired or all cache entries
server_info
Which optional features and config are active
Text returned by any of these is untrusted content extracted from a PDF.
pdf_info(content_trust=True) reports hidden text a human reader cannot
see, and the read tools flag it per page.
Example prompts:
code
"Read the PDF at /path/to/document.pdf"
"Which pages discuss supply chain risks?"
"Find sections about the training process"
"Show me what page 5 looks like"
"OCR pages 3-5 of the scanned PDF"
For a large document (e.g., a 200-page annual report):
code
User: "Summarize the risk factors in this annual report"
Agent workflow:
1. pdf_info("report.pdf")
→ 200 pages, TOC shows "Risk Factors" on page 89
2. pdf_search("report.pdf", "risk factors")
→ Matches with structural paragraph excerpts: each excerpt
is the bullet, paragraph, or heading that matched, not a
fixed-width window. Often enough to answer directly.
3. If excerpts are sufficient → synthesize answer
4. If more context needed:
pdf_read_pages("report.pdf", "89-95")
→ Full page text for deeper reading
Remote / HTTP transport
STDIO is the default and is what every example above uses. pdf-mcp-http
serves the same tools over HTTP, for clients that cannot spawn a process
(the Anthropic API MCP connector, claude.ai custom connectors) and for a
warm corpus shared by several clients.
Paths resolve on the server, so an HTTP agent reads what is already there:
files under an allow-listed root, or a URL the server fetches. It cannot
hand over a file from its own machine. It is single-tenant and fails
closed: with no auth token and no [paths] allow list, the process exits
rather than serving an open endpoint.
Docker images are published to GHCR for amd64 and arm64, with everything
baked in, so every tool works on the first request:
bash
./deploy.sh # token, image, start, health-checkcp your.pdf documents/ # this folder is the server's /data/pdfs
pdf-mcp works out of the box. To restrict which paths and URL hosts the
server may touch, tune cache and worker settings, or add your own
content-trust phrases, see docs/configuration.md.
Roadmap
See ROADMAP.md for planned features and release history.
Contributing
Contributions are welcome. See docs/contributing.md for setup, checks, the coherence eval harness, and quality-loop guidelines.
Contributors
Thank you to everyone who has helped improve this project through code, reviews, testing, and feature requests:
Per-release contributor credits are listed in the Changelog.
Security
Found a vulnerability? See SECURITY.md for the threat model, reporting channel, and expected response timeline. Please do not open a public GitHub issue for unpatched security reports.
The story behind the releases. Building pdf-mcp keeps surprising me: benchmarks that go the wrong way, formats that break everything, features I had to remove. I write about that thinking in The Dispatch. Come along if that's your kind of thing.
Background, benchmarks, and design notes from building pdf-mcp:
Getting started
How I Built pdf-mcp: The problem with large PDFs in AI agents and a working solution
How to Send PDFs Over 100 Pages to Claude's API: Measuring the real page and token ceilings on the Claude API, and the two ways around them: search the PDF on disk for questions, window the text and render only the picture pages for whole-document tasks
How Amazon Bedrock Helped Me Make My RAG Better: Benchmarking pdf_corpus_search against Bedrock Knowledge Bases at an equal 2,000-token budget for four cents surfaced two bugs eight months of self-testing missed: an excerpt picker discarding answers from pages it had already retrieved, and one embedding per page hiding short answers
How One Search Change Eliminated an Entire Agent Step: Switching pdf_search from fixed-width snippets to paragraph excerpts turned it from a pivot tool into a terminal tool: 97% vs 80% answer containment across a 30-query benchmark
Why Multi-Column PDFs Scramble Reading Order in RAG: Fixing two-column extraction (0.564 → 0.816 fidelity), the title-page author-grid regression it caused, and the aggregate metric that stayed blind to both
How I Fixed Vertical Japanese PDF Extraction: Tategaki pages extract scrambled because reading order is geometric, not stored; rebuilding it from glyph positions (columns right to left, characters top to bottom), with no OCR and no new dependency
Install
Configuration
Environment variables
PDF_MCP_CACHE_DIR
Directory for storing PDF cache (default: ~/.cache/pdf-mcp)