Pseudonymise and restore patient identifiers & PII in text — local, HIPAA Safe Harbor mode.
io.github.nickjlamb/redacta-mcp MCP Server
The io.github.nickjlamb/redacta-mcp Model Context Protocol (MCP) server pseudonymises and restores patient identifiers and other personally identifiable information (PII) in text. It operates locally in “HIPAA Safe Harbor mode,” focusing on anonymization, de-identification, and privacy for healthcare content.
🛠️ Key Features
Pseudonymisation and restoration of patient identifiers & PII in text
Local operation with HIPAA Safe Harbor mode
De-identification/redaction oriented workflow
🚀 Use Cases
Healthcare text processing where patient identifiers must be protected
PII redaction and later restoration for authorized use
Supporting privacy requirements in NHS/healthcare contexts
⚡ Developer Benefits
MCP-compatible service for agent workflows (mcp, agent-skill)
Pseudonymise medical and clinical documents before they're processed by AI or
shared. Redacta replaces patient identifiers with labelled tokens —
[PATIENT_NAME_1], [NHS_NUMBER_1], [DATE_OF_BIRTH_1], … — while leaving the
clinical meaning intact, and returns a redaction report alongside the cleaned
text.
It started as an Agent Skill and is now one engine
shipped across eight surfaces — an iOS app, agent skill, MCP server, a
self-hosted HTTP service with a Kubernetes deployment, two libraries, a CLI,
and a FigJam whiteboard plugin.
Running this in production? Redacta offers a small number of fixed-price
design-partner integrations for teams shipping AI agents on clinical or patient
data — deployment in your environment, one real workflow integrated, and a
data-flow document written for your DPO.
Details →
The detection logic lives in one place — the TypeScript engine
(@pharmatools/redacta, in npm-package/), which the MCP server and the
FigJam plugin consume, and which the iOS app runs on-device via JavaScriptCore.
The Python package mirrors it for pip users; the agent skill adds LLM reasoning
for free-text names on top of the deterministic patterns.
How it works
Two layers:
Patterns (deterministic). A bundled script (scripts/redact_structured.py,
Python standard library only, no network) matches fixed-format identifiers:
NHS numbers (Modulus-11 validated), UK National Insurance numbers, dates of
birth, UK postcodes, phone numbers, emails, and hospital/MRN numbers. US SSN
and ZIP codes are also handled.
Reasoning (judgement). The skill then has the agent handle what patterns
can't: patient names (told apart from the clinicians treating them), relatives
and carers, postal addresses, and identifying ages.
Self-check. A final pass re-reads the output for any identifier that slipped
through before the report is written.
It also works in reverse. Re-identification (scripts/reinstate.py) takes the
token map from an earlier redaction and restores the original values — so you can
redact a document, run it through another AI tool, and put the real details back
locally. Redact → process → re-identify is a complete round trip, and identifiers
only ever exist on your machine.
Safe Harbor mode. Ask for HIPAA Safe Harbor (or "US de-identification") and
Redacta applies a stricter pass: all dates (not just the date of birth), all
specific ages, and the remaining HIPAA identifiers — fax, certificate/licence,
device serial, VIN, and health-plan beneficiary numbers.
Self-hosting on Kubernetes
Organisations that can't let identifiable text leave their environment can
run Redacta inside their own infrastructure: a small HTTP service
(gateway-service/) deployable into an existing
Kubernetes cluster with plain YAML — two stateless replicas behind a
Service for redact/reinstate, an optional single-replica session boundary
for the protect → release loop, health probes, resource limits, restrictive
security defaults, and no-PHI logging. Text is pseudonymised before it
reaches any external AI service, and the processing boundary stays under
your control. Walkthrough (local kind cluster included):
gateway-service/k8s/README.md ·
concepts: docs/KUBERNETES.md.
Deploying somewhere a DPO will ask questions? There's a one-page security &
data-protection summary at
pharmatools.ai/redacta-security.
Redacta is a strong first line of defence, not a guarantee. It won't catch every
possible identifier and isn't a substitute for formal data-protection processes.
Always review the redaction report before sharing text.