Zero-Trust PII & secrets sanitizer for AI agents with in-memory local redaction.
The Model Context Protocol (MCP) server io.github.moxno/privacyscrubber-mcp is a “Zero-Trust PII & secrets sanitizer.” It locally scrubs data in-memory before sending context to LLMs, aiming to support privacy-by-design and privacy-first workflows focused on protecting sensitive information.
🛠️ Key Features
Zero-Trust PII and secrets sanitization
In-memory local scrubbing before LLM context is provided
Supports PII detection and masking/redaction
🚀 Use Cases
Preventing PII exposure in LLM prompts and context
Applying privacy-focused data protection and anonymization
Masking/redacting sensitive fields such as phone data
⚡ Developer Benefits
Implements PII detection, masking-methods, and pii-redaction concepts
Fits MCP server integrations and cursor-plugin usage patterns
Includes topics spanning pii-anonymization and privacy-tools
⚠️ Limitations
Description emphasizes local in-memory scrubbing; it does not specify network, storage, or compliance scope (e.g., HIPAA) beyond included topics.
CISO-Approved Zero-Trust PII & Secrets Redaction MCP Server for Cursor, Windsurf, and Claude Desktop.
Locally scrubs PII, secrets, credentials, and custom regex rules from files and text contexts before they reach remote LLM providers to prevent API leaks and ensure HIPAA/SOC 2 compliance at the developer endpoint.
⭐ Support Zero-Trust Open Source: If PrivacyScrubber protects your API keys and code from leaks, please Star this repository or run gh repo star moxno/privacyscrubber-mcp in your terminal!
🔒 Zero-Trust Data Flow
All sensitive parameters, identifiers, and variables are intercepted locally inside your machine's RAM. They are replaced by tokens (e.g. [EMAIL_1]) before being sent to the AI. Once the AI responds, the tokens are safely swapped back to original values in your local context.
Need direct, in-memory zero-trust PII sanitization in your backend microservice, Next.js app, or RAG vector pipeline rather than an MCP server? Use our official zero-dependency SDK:
bash
npm install @privacyscrubber/sdk
typescript
importOpenAIfrom'openai';
import { wrapOpenAI } from'@privacyscrubber/sdk';
// Transparently masks PII before sending to LLM and rehydrates responses:const openai = wrapOpenAI(newOpenAI({ apiKey: process.env.OPENAI_API_KEY }));
// Outbound prompt is sanitized in local RAM before leaving your machine:// "Schedule a call with [NAME_1] at [EMAIL_1] regarding API key [AWS_KEY_1]."const completion = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Schedule a call with Alice Smith at alice@acme.com with AKIAIOSFODNN7EXAMPLE.' }]
});
// Incoming LLM answer is automatically rehydrated with "Alice Smith (alice@acme.com)":console.log(completion.choices[0].message.content);
⚡ Why @privacyscrubber/sdk vs Microsoft Presidio?
Microsoft Presidio is the Python standard, but deploying it in a Node.js / TypeScript stack requires running heavy Python microservices, Docker containers, and 500MB+ spaCy NLP models with 35–120ms latency. @privacyscrubber/sdk runs 100% in-process with zero dependencies:
🤖 AI Orchestrator Recipes: LangChain, LlamaIndex, CrewAI & AutoGPT
If you are building autonomous agents, RAG vector pipelines, or backend services rather than single-user IDE prompts, use @privacyscrubber/sdk to sanitize data in-memory:
1. LangChain.js (LCEL & RAG Document Ingestion)
typescript
import { ChatOpenAI } from'@langchain/openai';
import { createLangChainTransform, createDocumentTransformer } from'@privacyscrubber/sdk';
// A. RAG Pre-Ingestion: sanitize documents before embedding into vector storesconst docTransformer = createDocumentTransformer({
profile: 'General',
sensitiveMetadataKeys: ['account_owner', 'submitter_email'],
attachTelemetry: true
});
const sanitizedDocs = await docTransformer.transformDocuments(rawDocs);
// B. LCEL Runtime Chains: sanitize prompts and restore responses in local RAMconst transform = createLangChainTransform({ defaultProfile: 'Dev' });
const { scrubbedText, tokenMap } = transform.preprocess(
"Deploying database with user admin and pwd postgresql://user:SecretPass123@db.internal:5432/prod"
);
const model = newChatOpenAI({ model: 'gpt-4o' });
const response = await model.invoke(scrubbedText);
const finalOutput = transform.postprocess(response.content, tokenMap);
2. LlamaIndex.TS & Universal Vector DB Ingestion (Chroma, Pinecone, Qdrant)
typescript
import { createVectorIngestionGuard, createLlamaIndexTransform } from'@privacyscrubber/sdk';
// Initialize in-memory Vector Ingestion Guard (<1ms per batch, zero network egress)const guard = createVectorIngestionGuard({
profile: 'Finance',
sensitiveMetadataKeys: ['contractor_email', 'billing_contact']
});
// A. Sanitize Chroma columnar batches ({ ids, documents, metadatas })const cleanChroma = guard.sanitizeRecords(chromaBatch);
// B. Sanitize Pinecone / Qdrant record arrays ({ id, text/payload, metadata })const cleanRecords = guard.sanitizeRecords(pineconeRecords);
// C. Align search query tokens with sanitized vector space before embeddingconst { query: alignedQuery } = guard.sanitizeQuery("Search user john@example.com records");
// D. Transparent Vector Store Proxy: auto-sanitizes on write, re-hydrates on queryconst guardedStore = guard.wrapVectorStore(nativeVectorStore);
Complete runnable recipes are included inside the npm package under @privacyscrubber/sdk/examples/ (or run npx @privacyscrubber/sdk).
🛡️ Architecture & Security Deep-Dive (Zero-Trust vs Cloud DLP)
When AI IDEs (Cursor, Claude Desktop, Windsurf) connect to model providers, developer credentials, database connection strings, and internal customer PII are at continuous risk of prompt exfiltration. The @privacyscrubber/mcp-server enforces four immutable architectural guarantees:
Stdio Air-Gapped Transport:
The MCP server communicates exclusively over local standard input/output (stdio) child processes spawned by your IDE. It opens zero external listening ports and initiates zero remote network requests.
Volatile RAM-Only Token Map:
The mapping table between synthetic tokens ([AWS_KEY_1], [EMAIL_1]) and raw cleartext is maintained exclusively in ephemeral node memory and is wiped the moment your IDE session closes.
Deterministic AST Lookarounds vs Cloud Proxy Overhead:
Unlike cloud DLP gateways (Nightfall, Skyflow) that add 200–400ms latency and transmit unencrypted code to third parties, @privacyscrubber/mcp-server runs locally in <2ms with zero data egress.
claude mcp add privacyscrubber -- npx -y @privacyscrubber/mcp-server
🛠️ Provided Tools & JSON-RPC Specifications
1. sanitize_text
Redacts PII, secrets, API keys, and credentials from a text block and populates the volatile local replacement mapping.
Arguments:
text (string, required): The raw content or logs to sanitize.
profile (string, optional): Gated industry detection profile (e.g., 'General', 'Dev', 'Medical', 'Legal', 'Compliance'). Defaults to 'General'.
JSON-RPC Call Example:
json
{"method":"tools/call","params":{"name":"sanitize_text","arguments":{"text":"Contact me at dev-key-1234 or jane.doe@company.com","profile":"General"}}}
Response Example:
json
{"content":[{"type":"text","text":"Contact me at [SECRET_1] or [EMAIL_1]"}]}
2. reveal_text
Detokenizes the AI response back to the original values locally.
Arguments:
text (string, required): The response from the LLM containing tokenized placeholders.
JSON-RPC Call Example:
json
{"method":"tools/call","params":{"name":"reveal_text","arguments":{"text":"Please reach out to [EMAIL_1] regarding the update."}}}
Response Example:
json
{"content":[{"type":"text","text":"Please reach out to jane.doe@company.com regarding the update."}]}
3. sanitize_file
Reads a local file, extracts text, sanitizes it, and returns the redacted template for LLM analysis.
Supported Formats: Plain text (source code, logs, CSV, JSON, markdown) and Microsoft Word (.docx) documents.
Arguments:
filePath (string, required): Absolute file path to read and sanitize.
profile (string, optional): The industry detection profile.
4. guard_exec (Command Execution Firewall)
Safely executes terminal commands in an isolated child process, masking stdout/stderr PII, database credentials, and API keys in local RAM before passing them to the AI agent. Includes a CISO audit receipt in stderr.
Arguments:
command (string, required): The shell command to execute (e.g. cat .env, docker logs web, git diff).
cwd (string, optional): Working directory.
profile (string, optional): Detection profile (defaults to Dev).
timeout_ms (number, optional): Timeout in ms (defaults to 15000).
Reads files (.env, configs, source code, database dumps) and tokenizes all passwords, JWTs, and PII in volatile memory, returning safe redacted content for AI reasoning.
Arguments:
file_path (string, required): Path to file.
profile (string, optional): Detection profile (defaults to Dev).
max_lines (number, optional): Line cap for large files (defaults to 500).
6. guard_git_diff (Pre-Commit & Diff Sanitizer)
Inspects staged (--cached) or unstaged repository diffs, redacting any newly introduced secrets or PII in local RAM before AI code review or commit message generation.
Arguments:
staged (boolean, optional): If true, inspects staged changes (git diff --cached). Defaults to false.
cwd (string, optional): Working directory.
profile (string, optional): Detection profile (defaults to Dev).
7. guard_apply_patch (Safe Patch Applicator)
Reverses token placeholders ([API_KEY_1], [SECRET_1]) in AI-generated code or text by looking up the local RAM session map, creating a .bak backup, and writing authentic cleartext directly to disk. The remote LLM never sees real secrets.
Arguments:
file_path (string, required): Path to target file.
content (string, required): Content containing tokens to restore on disk.
create_backup (boolean, optional): Backup existing file before write (defaults to true).
Returns a visual dashboard showing your current tier, session request count, active profiles, and upgrade instructions. Use it at any time to check your license status or get setup help.
By default, the server runs under the Free Tier (restricted to 15,000 characters per request and the basic General PII profile). To unlock 30 specialized engineering, medical, legal, and financial PII profiles, as well as team-wide custom rules, you can purchase a commercial license.
After purchasing a PRO license at privacyscrubber.com/pricing, you will receive a license key. Add it to your MCP client config as an environment variable: PRIVACYSCRUBBER_KEY.
Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
The Zero-Trust Data Sanitization (ZTDS) architecture, in-memory deterministic tokenization, cryptographic session handoff, and stdio execution methods implemented in this package are proprietary technology of Ilya Sibiryakov (BrandMeWeb) and are protected under Patent Pending status:
Patent Office: State of Israel Ministry of Justice, Patent Office (ILPO)