Agent♥︎Age
Catalog

io.github.jztan/pdf-mcp

Official

by jztan · Python

Agentic RAG over one PDF or a whole folder: hybrid search, selective page reads, tables, OCR.

io.github.jztan/pdf-mcp (Model Context Protocol Server)

The io.github.jztan/pdf-mcp MCP server provides agentic RAG over one PDF or a whole folder, supporting hybrid search, selective page reads, and extraction of tables and OCR content. It is positioned for document processing workflows using tools and libraries related to PDF parsing and semantic search.

🛠️ Key Features

  • Hybrid search across PDF content
  • Selective page reads
  • Table extraction
  • OCR support
  • Agentic RAG for single PDFs or folders

🚀 Use Cases

  • Querying a single PDF with semantic and hybrid retrieval
  • Building RAG over a directory of PDFs
  • Extracting structured table data and OCR text for downstream LLM use

⚡ Developer Benefits

  • Supports MCP integration for document retrieval
  • Python-focused tooling (topics include python, pymupdf, and model-context-protocol)
  • Keyword-oriented organization for developers (topics: semantic-search, pdf-extraction, table-extraction, agentic-rag)

⚠️ Limitations

  • Scope is centered on PDF inputs (one file or a folder), including OCR and table handling

Topics

aiclaudedocument-processingllmmcppdfpythoncodex-cliopencodemcp-servercjkocrpdf-extractionpymupdfsemantic-searchtable-extractionagentic-ragclaude-codemodel-context-protocolrag