Fail-closed privacy layer that pseudonymises enterprise API data for LLM agents and MCP clients.
io.github.AindriuB/data-prism MCP Server
The io.github.AindriuB/data-prism Model Context Protocol (MCP) server provides a privacy layer for enterprise APIs. It is described as refusing to start unless a reviewed “adapter jar” is provided per source. It focuses on pseudonymisation and is associated with Java and Spring Boot environments.
🛠️ Key Features
Privacy layer for enterprise APIs
Start-up depends on a reviewed adapter jar per source
Supports pseudonymisation
🚀 Use Cases
Deploying privacy controls for enterprise API integrations
Enforcing reviewed adapters in Java-based API environments
⚡ Developer Benefits
Clear gating via reviewed adapter jars before the server starts
Suitable for teams working with Java, MCP, and Spring Boot
⚠️ Limitations
Cannot start without a reviewed adapter jar per source
Fail-closed privacy layer that pseudonymises enterprise API data for LLM agents and MCP clients.
Data Prism is an open-source privacy layer for Java/Spring teams putting LLM agents or MCP clients in front of internal APIs holding customer data. It pseudonymises personal data per privacy scope, refuses anything unclassified, and can keep a hash-chained audit trail.
What it is. A privacy enforcement layer for enterprise AI. It gives LLM
agents useful, correlated business context — one consistent synthetic
identity per entity within a privacy scope — while controlling which
personal data crosses the AI boundary, and in what form.
Who it's for. Java/Spring platform and backend teams putting LLM agents
or MCP clients in front of internal APIs that hold customer data. If nothing
you run exposes personal data to a model, you don't need this.
Status: the walking skeleton and every slice through S9a are built, with 18
Maven submodules (19 Maven projects in the reactor counting the root
pom-packaged aggregator itself) and a passing test suite. The privacy
engine, correlation and consistency findings, parallel mTLS connectors,
an embedded Hazelcast identity cache and read budget (shared only across members that have joined one cluster; see multiple instances), an OAuth2 resource server
with session-derived PrivacyContext, audit and metrics are all real and
exercised end to end. The standalone server is the primary deployment
surface; the Spring Boot starter is the embedded option. A one-command local
Compose quickstart also exists: see "Try it" below. Two MCP tools ship
today, get_entity_context and compare_entity_sources — the other two
named in the design review, search_entity_data and
describe_entity_model, are not yet built (docs/tools.md "Not yet
built"). A durable, append-only, hash-chained audit sink and an offline
AuditChainVerifier ship as of 0.3.0, opt-in via
dataprism.audit.sink: hash-chained; the verifier catches an edit or
deletion inside a writer's chain, but cannot detect truncation of a writer's
most recent records or the deletion of a whole process boot's records, and
the trail does not resist an operator, or anyone else, who already has
write access to the file (docs/audit.md "What this does and does not
prove"). Not built: the re-identification operator surface (deferred past
V1 by decision, see docs/architecture.md#decisions-worth-knowing) and the
Elasticsearch connector and its search tools. See docs/plan/PLAN.md for
what is open.
The problem
An organisation wants an LLM to investigate live business data spread across
several systems. Giving the model direct API access is not acceptable: those APIs
carry personal and confidential data, each system represents the same entity
differently, and raw identifiers let anything downstream correlate across
sessions.
The obvious fix — redact everything sensitive — destroys the investigation. Once
three systems' names for one person are all [REDACTED], the model cannot tell
whether it is looking at one person or three.
What Data Prism does
It sits between the two and does two things that are easy to confuse:
It makes identity consistent. One subject gets one synthetic identity across
every source, derived deterministically from (scope, subject, namespace, algorithm version, key) — never random, never stored in plaintext, and
reproducible without the cache. The same person in three systems reads as one
person to the model.
It leaves the data inconsistent, and says so. If those three systems disagree
about a name, the answer carries a finding that says they disagree. The platform
never makes enterprise data look cleaner than it is. That distinction is the
point of the project:
Identity representation becomes consistent. Underlying data inconsistencies
become more visible, not less.
Pseudonyms are scoped. The same person in two different investigations gets two
different synthetic identities, so nothing correlates across cases by accident.
What it is not
Not an API gateway, not an ETL platform, not a master-data system, not an
identity provider, and not an entity-resolution engine — correlation requires a
key the sources already share, behind a documented SPI. It carries no business
domain: no Customer, Taxpayer or Employee type exists outside the example
application.
It is not anonymisation. Under GDPR Art. 4(5), pseudonymised data is still
personal data. Sending Data Prism output to a third-party model is still
processing, and still needs a lawful basis, a DPIA, and a transfer mechanism
where the provider is outside the EU. The platform reduces exposure; it does not
remove the obligation.
Try it
The fastest way to see a real MCP call answered by the real privacy engine —
no local JDK, no Maven install, one command:
bash
docker compose up
pulls the published ghcr.io/aindriub/data-prism-quickstart-<name> images
(pin one with QUICKSTART_IMAGE_TAG=0.5.0; run
docker compose -f compose.yaml -f compose.build.yaml up --build instead to
build every image from source) and brings up the standalone server, a
synthetic fixture API and a local HTTPS JWT issuer, proving an
agent-compatible get_entity_context call returns a pseudonymised response.
Walk through it in
docs/quickstart.md; connect your own agent client to
either that stack or a real deployment via
docs/agents/.
Once you have seen the demo, protect your own API: docs/quickstart.md ends
with a "What next" section pointing at
docs/protect-your-own-api.md, a YAML-only
walkthrough from a real JSON REST API to a working get_entity_context call.
If you found this on the MCP registry
The ghcr.io/aindriub/data-prism-server image listed there is published as a
multi-architecture manifest list covering linux/amd64 and linux/arm64,
each built and verified natively — docker run on Apple Silicon or any other
arm64 host pulls the arm64 image directly, no emulation required.
It is not a one-command install, on either architecture. docker run alone
yields a server that refuses to start: DataPrismContractValidator demands a
reviewed DataSourceAdapter bean for every configured source, and
DataPrismProperties.validate() demands a full deployment configuration (JWT
issuer/audience/JWKS, caller-claim mappings, security policy, HMAC key
reference, audit sink, metrics sink, Hazelcast topology). Neither ships in the
image. Two things an operator must supply themselves before it serves
anything:
A reviewed DataSourceAdapter (and IdentityResolver) jar for each
API you are protecting, mounted onto the image's loader path.
A deployment configuration satisfying the dataprism.* vocabulary.
docs/configuration.md is the authoritative,
complete contract for both. The "Try it" section above is a local Compose
fixture for evaluation, not this image or that configuration.
docs/plan/PLAN.md is the working queue. GitHub Issues is the front door for
anything coming from outside — file there, not in PLAN.md.
Stack
Java 21 bytecode, Spring Boot 4.1.x, Maven multi-module, Hazelcast, Model Context
Protocol via the official MCP Java SDK. The library modules require Java 21 or
newer and Spring Boot 4.1.x; the published images run Java 25. Boot 3
consumers stay on 0.4.x.
Artifacts publish under group io.github.aindriub as data-prism-<module>, with
package root io.github.aindriub.dataprism.
Building and running
Building needs JDK 21 or newer (the build compiles with --release 21, so a
newer local JDK is fine; CI builds on 21 and 25) and Maven >= 3.6.3 (pom.xml:282-284 enforces this).
bash
mvn -B --no-transfer-progress verify
This is the same command CI runs (.github/workflows/build.yml). It builds all
18 submodules plus the root aggregator, runs the full test suite, the
ArchUnit boundary rules, and the enforcer rule that keeps the classpath on
Jackson 3, with Jackson 2 banned apart from jackson-annotations.
data-prism-server is the primary executable distribution. Its /health
liveness probe is public and carries no deployment detail; its configured MCP
path (normally /mcp) requires a verified bearer JWT. It deliberately contains
no source schema, fixture adapter, or key. Supply configuration described in
docs/configuration.md, plus a reviewed adapter for
each configured source. Two ways to get one: a Java adapter extension with an
annotated response model (the general case — nested objects, any transport),
or, when the source's response is one flat JSON object, the published
data-prism-connectors-rest artefact — loaded via -Dloader.path, configured
entirely in YAML, no Java required. See
docs/extending.md for both paths and exactly where the
configuration-driven one's coverage ends (it never descends into a nested
object).
Adapter extensions are ordinary jars containing Spring Boot auto-configuration
that declares the required DataSourceAdapter beans and an explicit reviewed
IdentityResolver (use PassThroughIdentityResolver only when every source
genuinely shares the same identifier). Register that configuration in
META-INF/spring/org.springframework.boot.autoconfigure.AutoConfiguration.imports,
then load reviewed extension jars without rebuilding the server:
The process refuses startup if configuration, secrets, operational bindings, or
the exact configured adapter set is missing. data-prism-integration-tests is
the reactor's cross-module integration test suite, not a fixture-only demo,
and is never packaged into a deployable artefact. Container packaging and
Compose orchestration for a real, locally runnable instance of this exist too
— see "Try it" above and docs/quickstart.md.
Contributing
See CONTRIBUTING.md. The short version: read docs/conventions.md before
opening a pull request, and expect the privacy rules in it to be enforced
literally.
Licence
Apache License 2.0 — see LICENSE and NOTICE.
Install
Configuration
Environment variables
LOADER_PATHrequireddefault /app/adapters
Directory Spring Boot's PropertiesLauncher scans for extension jars; already set to /app/adapters by the image, but startup still fails with MISSING_SOURCE_ADAPTER until you bind-mount a reviewed DataSourceAdapter/IdentityResolver jar there (see the -v arguments above)
DATAPRISM_SECURITY_JWT_ISSUERrequired
OAuth2/OIDC issuer that mints the caller's JWT; required for every protected deployment
DATAPRISM_SECURITY_JWT_AUDIENCErequired
Expected JWT audience claim for this deployment; required for every protected deployment
DATAPRISM_SECURITY_JWT_JWK_SET_URIrequired
HTTPS JWKS location used to verify caller JWTs; exactly one of this or DATAPRISM_SECURITY_JWT_ISSUER_DISCOVERY_URI is required, never both
DATAPRISM_SECURITY_JWT_ISSUER_DISCOVERY_URI
Alternative HTTPS OIDC issuer-discovery location; set this instead of DATAPRISM_SECURITY_JWT_JWK_SET_URI, never both
Example only — declare DATAPRISM_SECURITYPOLICY_ROLES_<ROLE> per operator-defined role (no underscore between SECURITY and POLICY: Spring Boot's map-key enumeration under a hyphenated dataprism.security-policy.roles.<role> segment only binds the concatenated prefix, verified by binding this property directly against Spring Boot 4.1.1), a comma-separated list of known MCP tool capabilities; at least one role-to-capability mapping is required
DATAPRISM_PRIVACY_PROFILErequired
Name of the reviewed privacy profile implementation to apply; required
Name of the environment variable holding the HMAC key material; exactly one of this or DATAPRISM_PRIVACY_HMAC_KEY_PROVIDER_REFERENCE is required, never both, and never a literal key value
DATAPRISM_PRIVACY_HMAC_KEY_PROVIDER_REFERENCE
Reference to an approved secret provider holding the HMAC key material; set this instead of DATAPRISM_PRIVACY_HMAC_KEY_ENVIRONMENT_VARIABLE, never both
DATAPRISM_AUDIT_SINKrequired
Audit sink implementation: one of approved-sink, slf4j, hash-chained; required, never downgraded to no-op
DATAPRISM_AUDIT_WRITER_IDrequired
Writer/instance identity recorded on every audit entry; required
DATAPRISM_AUDIT_FILE_PATH
Path to the durable, hash-chained audit log; required when DATAPRISM_AUDIT_SINK=hash-chained, refused as MISSING_AUDIT_FILE_PATH if absent for that sink, ignored otherwise
DATAPRISM_METRICS_SINKrequired
Metrics sink binding, currently only micrometer; required in production, never the framework no-op
DATAPRISM_HAZELCAST_TOPOLOGYrequired
Cluster read-budget topology: embedded (shared across members that have joined one cluster; set cluster name and join mode) or single-node (enforced per process); required, never defaulted
DATAPRISM_HAZELCAST_CLUSTERNAME
Hazelcast cluster name; required when DATAPRISM_HAZELCAST_TOPOLOGY=embedded, never dev, refused if set with single-node
DATAPRISM_HAZELCAST_JOIN_MODE
How cluster members find each other: tcp-ip, kubernetes or none; required when DATAPRISM_HAZELCAST_TOPOLOGY=embedded (none is an explicit single member bound to 127.0.0.1), refused if set with single-node
DATAPRISM_HAZELCAST_JOIN_MEMBERS
Comma-separated member addresses (host or host:port); required when DATAPRISM_HAZELCAST_JOIN_MODE=tcp-ip, otherwise unused
DATAPRISM_HAZELCAST_MEMBER_PORT
Hazelcast member port, default 5701, never auto-incremented; must differ from the server, management and operator ports. Member traffic is unencrypted, so keep it on a private network. Optional
DATAPRISM_OPERATOR_ENABLED
Enables the operator surface; optional, off by default
DATAPRISM_OPERATOR_PORT
Port of the operator listener; required when DATAPRISM_OPERATOR_ENABLED=true, must differ from the server, management (OPERATOR_PORT_SHARED otherwise) and member ports, and has no fixed default (publish it explicitly)
DATAPRISM_OPERATOR_REQUIREDAUDIENCE
JWT audience an operator token must carry; required when DATAPRISM_OPERATOR_ENABLED=true
DATAPRISM_OPERATOR_REQUIREDSCOPE
JWT scope an operator token must carry; required when DATAPRISM_OPERATOR_ENABLED=true
DATAPRISM_SOURCES_CUSTOMER_BASE_URLrequired
Example only — declare DATAPRISM_SOURCES_<NAME>_BASE_URL (HTTPS) per configured source; at least one source, each with its own reviewed DataSourceAdapter bean, is required
DATAPRISM_SOURCES_CUSTOMER_TIMEOUTrequired
Example only — declare DATAPRISM_SOURCES_<NAME>_TIMEOUT (positive duration) per configured source; required alongside its base URL