PDF SPEC MCP Server

ๆฅๆฌ่ช็ README ใฏใใกใ
An MCP (Model Context Protocol) server that provides structured access to ISO 32000 (PDF) specification documents. Enables LLMs to navigate, search, and analyze PDF specifications through well-defined tools.
IMPORTANT
This is a specification reference, not a rule engine.
It retrieves and structures the text of ISO 32000 โ clauses, tables, definitions, and
shall/should/may requirements. It does not examine a PDF file, and it cannot tell
you whether a document conforms to anything. Conformance verdicts come from
pdf-verify-mcp
(validate_conformance / evaluate_policy).
The distinction matters because three different things get conflated:
declaration โ a label the file wrote about itself ("I am PDF/A" in the metadata). Writing it is not evidence /
conformance โ whether the file actually meets the standard. There is no way to prove it in full; you can only find where it breaks the rules /
validation โ what a validator (veraPDF and the like) reports against the checks it implements. A pass means "this inspection did not fail", not "the file conforms to the standard".
Reading a shall here tells you what the standard requires โ not whether your file meets it.
A search that returns nothing means "cannot answer", not "no such requirement."
ISO 19005 (PDF/A) and ETSI PAdES are outside this corpus; see list_specs โ coverage.gaps.
What each PDF family server does โ and does not do
| Server | Does | Does not |
|---|
| pdf-spec-mcp (this) | Search, retrieve and extract requirements from 17 PDF-related documents | Is not a rule engine. Does not define business rules, inspect PDF files, or validate schemas. ISO 19005 (PDF/A) is not part of the corpus |
| pdf-reader-mcp | Extract text / tables / structure tree / fonts / annotations / images / signature fields | Does not verify cryptography. Does not read the incremental-update history, does not map object IDs to coordinates, does not OCR |
| pdf-writer-mcp | Create, page operations, tagging, forms, annotations, metadata, attachments, PDF/A-3b scaffolding | Does not sign. Does not make the file meet the standard โ it can write a label, not conformance |
| pdf-verify-mcp | Conformance validation (delegated to veraPDF), cryptographic signature verification, tamper detection, policy verdicts | Does not prove the file meets the standard (it can only find where it breaks the rules). Does not vouch for the signer's identity. Does not judge whether the content is true |
IMPORTANT
PDF specification files are NOT included in this package.
You must obtain the PDF specification documents separately and place them in a local directory.
Download from: PDF Association โ Sponsored Standards
See "Setup" for details.
Features
- Multi-spec support โ Auto-discovers and manages up to 17 PDF-related documents (ISO 32000-2, PDF/UA, Tagged PDF guides, etc.)
- Structured content extraction โ Headings, paragraphs, lists, tables, and notes from any section
- Full-text search โ Keyword search with section-aware context snippets
- Requirements extraction โ Extracts normative language (shall / must / may) per ISO conventions
- Definitions lookup โ Term definitions from Section 3 (Definitions)
- Table extraction โ Multi-page table detection with header merging
- Version comparison โ Diff PDF 1.7 vs PDF 2.0 section structures
- Bounded-concurrency processing โ Parallel page processing for large documents
- On-disk index cache โ The search index and the full requirements scan are built once per PDF and reused by every later process (ISO 32000-2: ~6 s โ ~0.2 s)
Architecture
graph LR
subgraph Client["MCP Client"]
LLM["LLM<br/>(Claude, etc.)"]
end
subgraph Server["PDF Spec MCP Server"]
direction TB
MCP["MCP Server<br/>index.ts"]
subgraph Tools["Tools Layer"]
direction LR
T1["list_specs"]
T2["get_structure"]
T3["get_section"]
T4["search_spec"]
T5["get_requirements"]
T6["get_definitions"]
T7["get_tables"]
T8["compare_versions"]
end
subgraph Services["Services Layer"]
direction LR
REG["Registry<br/>Auto-discovery"]
LOADER["Loader<br/>LRU Cache"]
SVC["PDFService<br/>Orchestration"]
CMP["CompareService<br/>Version Diff"]
end
subgraph Extractors["Extractors"]
direction LR
OUTLINE["OutlineResolver<br/>TOC & Section Index"]
CONTENT["ContentExtractor<br/>Structured Extraction"]
SEARCH["SearchIndex<br/>Full-text Search"]
REQ["RequirementExtractor"]
DEF["DefinitionExtractor"]
end
subgraph Utils["Utils"]
direction LR
CACHE["LRU Cache"]
CONC["Concurrency"]
VALID["Validation"]
end
end
subgraph PDFs["PDF Spec Files (obtained separately)"]
direction LR
PDF1["ISO 32000-2<br/>(PDF 2.0)"]
PDF2["ISO 32000-1<br/>(PDF 1.7)"]
PDF3["TS 32001โ32005<br/>PDF/UA, etc."]
end
LLM <-->|"stdio / JSON-RPC"| MCP
MCP --> Tools
Tools --> Services
Services --> Extractors
Services --> Utils
LOADER --> PDFs
REG -->|"Filename pattern<br/>auto-discovery"| PDFs
style Client fill:#e8f4f8,stroke:#2196F3
style PDFs fill:#fff3e0,stroke:#FF9800
style Tools fill:#e8f5e9,stroke:#4CAF50
style Services fill:#f3e5f5,stroke:#9C27B0
style Extractors fill:#fce4ec,stroke:#E91E63
style Utils fill:#f5f5f5,stroke:#9E9E9E
Layer Overview
| Layer | Responsibility |
|---|
| Tools | MCP tool schema definitions & handlers (input validation) |
| Services | Business logic (PDF registry, loader, orchestration) |
| Extractors | Information extraction from PDFs (TOC, content, search, requirements, definitions) |
| Utils | Shared utilities (cache, concurrency, validation) |
Setup
1. Obtain PDF Specification Files
WARNING
PDF specifications are copyrighted documents and are not included in this package.
Download them from the sources below and place them in a local directory.
All 17 files below are supported. You do not need all of them โ place only the specs you need (at minimum, ISO 32000-2 is recommended).
pdf-specs/
โ
โ โโ Standards โโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โโโ ISO_32000-2_sponsored_EC3.pdf # iso32000-2 : PDF 2.0 EC3 (recommended; falls back to -ec2.pdf)
โโโ ISO_32000-2-2020_sponsored.pdf # iso32000-2-2020 : PDF 2.0 original
โโโ PDF32000_2008.pdf # pdf17 : PDF 1.7 (for version comparison)
โโโ pdfreference1.7old.pdf # pdf17old : Adobe PDF Reference 1.7
โ
โ โโ Technical Specifications (TS) โโโโโโโโโ
โโโ ISO_TS_32001-2022_sponsored_EC3.pdf # ts32001 : Hash extensions (SHA-3)
โโโ ISO_TS_32002-2022_sponsored_EC3.pdf # ts32002 : Digital signature extensions (ECC/PAdES)
โโโ ISO_TS_32003-2023_sponsored.pdf # ts32003 : AES-GCM encryption
โโโ ISO-TS-32004-2024_sponsored.pdf # ts32004 : Integrity protection
โโโ ISO-TS-32005-2023-sponsored.pdf # ts32005 : Namespace mapping
โ
โ โโ PDF/UA (Accessibility) โโโโโโโโโโโโโโโโ
โโโ ISO-14289-1-2014-sponsored.pdf # pdfua1 : PDF/UA-1
โโโ ISO-14289-2-2024-sponsored.pdf # pdfua2 : PDF/UA-2
โ
โ โโ Guides โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โโโ Tagged-PDF-Best-Practice-Guide.pdf # tagged-bpg : Tagged PDF Best Practice
โโโ Well-Tagged-PDF-WTPDF-1.0.pdf # wtpdf : Well-Tagged PDF
โโโ PDF-Declarations.pdf # declarations: PDF Declarations
โ
โ โโ Application Notes โโโโโโโโโโโโโโโโโโโโโ
โโโ PDF20_AN001-BPC.pdf # an001 : Black Point Compensation
โโโ PDF20_AN002-AF.pdf # an002 : Associated Files
โโโ PDF20_AN003-ObjectMetadataLocations.pdf # an003 : Object Metadata
2. Install
This package ships a CLI binary (pdf-spec-mcp) intended to be launched by an MCP client.
You do not need to install it manually โ just point your MCP client to npx @shuji-bonji/pdf-spec-mcp@latest as shown in the next step.
If you want to run it directly from the shell (e.g. for debugging):
PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest
Or install it globally (optional):
npm install -g @shuji-bonji/pdf-spec-mcp
PDF_SPEC_DIR=/path/to/pdf-specs pdf-spec-mcp
Environment Variable
| Variable | Description | Default |
|---|
PDF_SPEC_DIR | Directory containing PDF specification files | (required) |
PDF_SPEC_CACHE_DIR | Where the on-disk index cache lives (see Index cache) | ${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp |
PDF_SPEC_CACHE | Set to off to neither read nor write the index cache | on |
Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"pdf-spec": {
"command": "npx",
"args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"],
"env": {
"PDF_SPEC_DIR": "/path/to/pdf-specs"
}
}
}
}
IMPORTANT
Use @latest (or pin a version). npx -y <pkg> without a version keeps running whatever
it cached the first time โ -y only skips the install prompt, it does not check for updates.
A bare specifier will happily run a months-old release. @latest makes npx check the registry
on each start; pin @0.4.0 instead if you want reproducibility.
To clear a stale cache: rm -rf ~/.npm/_npx.
Cursor / VS Code
Add to .cursor/mcp.json or VS Code MCP settings:
{
"mcpServers": {
"pdf-spec": {
"command": "npx",
"args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"],
"env": {
"PDF_SPEC_DIR": "/path/to/pdf-specs"
}
}
}
}
Index cache
Two operations walk every page of a specification: the first search_spec on a spec builds
its full-text index (ISO 32000-2, 1023 pages: about 6 s on a laptop), and get_requirements
without a section scans every section (about 11 s). Everything else opens only the pages it
needs and answers in well under a second.
Since 0.5.0 those two results are written to disk after the first build and read back by every
later process โ an MCP client that starts one server per session no longer pays the build each
time. The second process answers the same search_spec in about 0.2 s and the full
requirements scan in about 0.02 s, from byte-for-byte the same index.
- Location:
${PDF_SPEC_CACHE_DIR:-${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp}/v1/<version>/<spec>.<kind>.<sha256[0:16]>.json.
The whole 17-spec corpus is about 18 MB per package version.
- Key: package version,
pdfjs-dist version, spec id, and the SHA-256 of the PDF. A
replaced PDF, an upgraded server, or an upgraded pdfjs all miss and rebuild. Entries of
older versions are left in place (another install may still use them); --clear-cache
removes everything.
- Failure is a miss, never an error: an unreadable, truncated, or foreign file is rebuilt;
an unwritable directory is reported once on stderr and the server carries on without a cache.
- It is derived from your copy of the PDFs and stays on your machine. It is not part of
the package and must not be redistributed โ the specifications are copyrighted.
Nothing about searching changes: the same in-memory structure is searched by the same code.
Only where it comes from (built vs. read) does.
Pre-building the cache
The cache fills lazily, one spec at a time as tools touch it. To warm every spec up front โ after
installing, after upgrading, or from cron โ run the CLI (it uses the same code path as the tools,
processes specs sequentially, and exits):
PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest --build-cache
npx -y @shuji-bonji/pdf-spec-mcp@latest --cache-info
npx -y @shuji-bonji/pdf-spec-mcp@latest --clear-cache
A full build of the 17-spec corpus takes about a minute on a laptop.
All tools accept an optional spec parameter to target a specific specification (default: iso32000-2).
| Tool | Description |
|---|
list_specs | List all discovered PDF specifications with metadata |
get_structure | Get section hierarchy (table of contents) with configurable depth |
get_section | Get structured content of a specific section |
search_spec | Full-text keyword search across a specification |
get_requirements | Extract normative requirements (shall/must/may) |
get_definitions | Lookup term definitions |
get_tables | Extract table structures from a section |
compare_versions | Compare PDF 1.7 and PDF 2.0 section structures |
list_specs โ Discover Specifications
List all available specification documents. Use the returned IDs as the spec parameter in other tools.
{ }
{ "category": "ts" }
{ "category": "pdfua" }
{ "category": "guide" }
get_structure โ Table of Contents
Get the section hierarchy (TOC tree) of a specification.
{ "max_depth": 1 }
{ "max_depth": 2 }
{ "spec": "ts32002" }
{ "spec": "pdfua2", "max_depth": 2 }
get_section โ Section Content
Get structured content (headings, paragraphs, lists, tables, notes) of a specific section.
A parent section returns its entire subtree (its preamble followed by all subsections, in document order). Top-level clauses can be very large โ prefer the most specific section number.
{ "section": "7.3.4.2" }
{ "section": "Annex A" }
{ "spec": "ts32002", "section": "5" }
{ "spec": "pdfua2", "section": "8" }
search_spec โ Full-text Search
Search across a specification with section-aware context snippets. The first call on a spec
builds its index (a few seconds); the index is then cached on disk (see Index cache).
{ "query": "digital signature" }
{ "query": "font", "max_results": 5 }
{ "spec": "ts32002", "query": "CMS" }
get_requirements โ Normative Requirements
Extract normative requirements (shall / must / may) per ISO conventions.
{ "section": "12.8" }
{ "section": "12.8", "level": "shall" }
{ "section": "7.3", "level": "shall not" }
{ "spec": "pdfua2", "section": "8", "level": "shall" }
get_definitions โ Term Definitions
Look up term definitions from Section 3 (Definitions).
{ "term": "font" }
{ }
{ "spec": "pdfua2", "term": "artifact" }
Extract table structures (headers, rows, captions) from a section. Multi-page tables are automatically merged.
{ "section": "7.3.4.2" }
{ "section": "7.3.4.2", "table_index": 0 }
{ "spec": "ts32002", "section": "5" }
compare_versions โ Version Comparison
Compare section structures between PDF 1.7 (ISO 32000-1) and PDF 2.0 (ISO 32000-2). Uses title-based automatic matching to detect matched, added, and removed sections.
NOTE
This tool requires both PDF 1.7 (PDF32000_2008.pdf) and PDF 2.0 files in PDF_SPEC_DIR.
{ "section": "12.8" }
{ }
Supported Specifications
The server auto-discovers PDF files in PDF_SPEC_DIR by filename pattern matching:
| Category | Spec IDs | Documents |
|---|
| Standard | iso32000-2, iso32000-2-2020, pdf17, pdf17old | ISO 32000-2 (PDF 2.0), ISO 32000-1 (PDF 1.7) |
| Technical Spec | ts32001 โ ts32005 | Hash, Digital Signatures, AES-GCM, Integrity, Namespace |
| PDF/UA | pdfua1, pdfua2 | Accessibility (ISO 14289-1, 14289-2) |
| Guide | tagged-bpg, wtpdf, declarations | Tagged PDF, Well-Tagged PDF, Declarations |
| App Note | an001 โ an003 | BPC, Associated Files, Object Metadata |
Directory Structure
src/
โโโ index.ts # Entry point: MCP server on stdio, or the cache CLI
โโโ cli.ts # --build-cache / --clear-cache / --cache-info
โโโ config.ts # Configuration & spec patterns
โโโ errors.ts # Error hierarchy (PDFSpecError โ sub-classes)
โโโ services/
โ โโโ pdf-registry.ts # Auto-discovery of PDF files
โ โโโ pdf-loader.ts # PDF loading with LRU cache
โ โโโ pdf-service.ts # Orchestration layer
โ โโโ index-store.ts # On-disk cache for the search / requirements indexes
โ โโโ compare-service.ts # Version comparison
โ โโโ outline-resolver.ts # Section index builder
โ โโโ content-extractor.ts # Structured content extraction
โ โโโ search-index.ts # Full-text search index
โ โโโ requirement-extractor.ts
โ โโโ definition-extractor.ts
โโโ tools/
โ โโโ definitions.ts # MCP tool schemas
โ โโโ handlers.ts # Tool implementations
โโโ types/
โ โโโ index.ts # Shared type definitions
โโโ utils/
โโโ concurrency.ts # mapConcurrent (bounded Promise.all)
โโโ text.ts # Text normalization
โโโ cache.ts # LRU cache
โโโ file-hash.ts # SHA-256 of a PDF (index cache key)
โโโ validation.ts # Input validation
โโโ logger.ts # Structured logger
Development
git clone https://github.com/shuji-bonji/pdf-spec-mcp.git
cd pdf-spec-mcp
npm install
npm run build
npm run test
npm run test:e2e
npm run lint
npm run format:check
License
MIT