AI governance MCP server: trust scoring, guardrails, bias/hallucination detection, compliance.
This MCP server provides AI governance capabilities focused on trust scoring and policy-oriented guardrails. It supports detection for bias and hallucinations and includes compliance-related functionality, as described by its description.
π οΈ Key Features
Trust scoring
Guardrails
Bias detection
Hallucination detection
Compliance
π Use Cases
Evaluating and scoring trust for AI outputs
Enforcing governance guardrails in agentic AI workflows
Adding bias and hallucination checks
Applying compliance requirements during runtime
β‘ Developer Benefits
Topics aligned with AI safety and AI governance (e.g., responsible-ai, policy-engine)
Designed for MCP-based integration (mcp, mcp-server)
β οΈ Limitations
No tool list or concrete interface details (e.g., toolCount) are provided in the available server data
WhitePact β an independent runtime authority, governance, and assurance layer for autonomous systems: a five-way governance decision engine (ALLOW / ALLOW_WITH_REDACTION / REQUIRE_APPROVAL / DENY / QUARANTINE), trust scoring, bias detection, guardrails, hallucination detection, compliance mapping (NIST AI RMF / EU AI Act / ISO 42001), cost intelligence, drift monitoring, a public Trust Index / leaderboard / AI Incident Database, and an MCP server (30 tools, 20 resources) with LangChain, LangGraph, and Google ADK trust-gate integrations.
Every team deploying AI in production faces the same gap: no unified way to
prove a model β or an autonomous agent's actions β is safe, fair, compliant,
and accountable. Audits are manual, bias is discovered in production,
compliance is a spreadsheet, an agent's tool calls go ungoverned, and nobody
knows what the LLM bill will be next month.
WhitePact gives you one platform β a REST API, a Python SDK, an MCP server,
and a live dashboard β that covers the full governance lifecycle:
Problem
Module
Output
Should this agent action be allowed, redacted, held for approval, denied, or quarantined?
WhitePactRuntimeGateway (governance core)
A five-way GovernanceDecision, deterministic, no LLM call in the decision path
Is this model trustworthy?
TrustScoreEngine
0β100 score, AβF grade, risk level
Does it comply with regulations?
ComplianceEngine
NIST AI RMF, EU AI Act tier, ISO 42001
Is it exposing PII?
GuardrailsEngine
Block / redact with audit log
Is it hallucinating?
HallucinationDetector
Risk score, unsupported claims
Can it be attacked?
RedTeamSimulator
10 vectors, CVE IDs, safe-refusal rate
How much is it costing?
CostTracker + ModelRouter
Per-model USD, routing to cheapest viable model
Is it getting worse over time?
TrustDriftMonitor
7/30-day trend, severity alerts
Is it biased?
BiasBuster
6 demographic probes, CI gate
Is this data labeled privately?
PrivacyLabel
Federated DP labels, never leaves device
Is this media real?
DeepfakeDetector
Ensemble confidence, method detected
Can I trust a third-party MCP server before connecting to it?
Free lookup, plus a real block/pause gate in-agent
Can any MCP client govern every AI call?
MCP Server
30 production governance tools over stdio, Streamable HTTP, or legacy HTTP+SSE
Install
bash
# Governance platform + REST API
pip install "rai-governance-platform[dashboard]"# With PostgreSQL support
pip install "rai-governance-platform[dashboard,postgres]"# With Redis + OpenTelemetry
pip install "rai-governance-platform[dashboard,redis,telemetry]"# With LLM providers
pip install "rai-governance-platform[dashboard,openai,anthropic]"# Everything
pip install "rai-governance-platform[all]"
The published PyPI package name (rai-governance-platform) and the import
name (responsibleai) predate the WhitePact rename and are kept as-is β
see MIGRATION_WHITEPACT_V2.md Section 3 and docs/PACKAGE_IDENTITY.md
for install vs import vs product naming (do not use pip install whitepact
unless PyPI documents that distribution).
30-second quickstart
bash
# Start the governance dashboard
pip install "rai-governance-platform[dashboard]"
uvicorn responsibleai.dashboard.app:app --port 8765
# Evaluate a model (no LLM key needed β supply your own scores)
curl -X POST http://localhost:8765/api/evaluate \
-H "Content-Type: application/json" \
-d '{
"model_name": "gpt-4o",
"provider": "openai",
"fairness": 0.80,
"privacy": 0.85,
"security": 0.82,
"robustness": 0.78,
"compliance": 0.90,
"authenticity": 0.88
}'
Open http://localhost:8765 for the live dashboard and
http://localhost:8765/api/docs for interactive API docs.
Governance core β five-way decisions, not a binary block/allow
src/responsibleai/governance/ (see SPEC.md Sections 4-8 for the full
architecture contract) is a deterministic runtime authority sitting in front
of agent tool calls:
Evidence (governance/evidence.py) β every decision is written to a
per-org, hash-chained EvidenceRecord; verify_chain() detects tampering.
Raw argument values are never stored, only field-name keys.
Approval workflow (governance/approval.py) β REQUIRE_APPROVAL
decisions queue a real, race-safe ApprovalRequest with a resolution API,
not just a log line.
Supply-chain scanner (src/responsibleai/supplychain/) β before an
agent trusts a third-party MCP server or tool, SupplyChainScanner returns
one of three explicit verdicts (VERIFIED_FACT / INFERRED_SIGNAL /
UNKNOWN) β never a single opaque trust score β from typosquat detection,
tool-description scanning, and known-incident cross-reference.
Identity Bridge (integrations/identity_bridge.py) β maps Entra ID,
Google Workspace, Okta, and AWS (Cognito / IAM Identity Center) ID token
claims into IdentityContext, plus map_groups_to_authority() to turn
IdP group membership into a granted-action-types AuthorityContext. See
MACHINE_AUTHORITY_V1.md's Identity Bridge section for exactly what's
verified (claim-shape correctness against each provider's public docs)
versus not (live-tenant testing, Graph/Admin-SDK group-name resolution,
AWS's non-JWT SigV4 path).
No governance decision is LLM-based; see
DETERMINISTIC_VS_PROBABILISTIC.md for why.
See it end-to-end: examples/08_whitepact_enterprise_scenario.py runs a
full scenario (an org onboarding an autonomous finance agent) through all
eight machine-authority invariants β ceiling, delegation, attenuation,
approval quorum, workflow composition, autonomy budget, memory firewall,
evidence bundle β against real code, no API keys required:
MCP Server β govern every AI call from Claude Code, Claude Desktop, or any MCP client
The MCP (Model Context Protocol) server exposes WhitePact as 30 tools and
20 resources (10 canonical resource URIs, dual-advertised under both
whitepact:// and rai:// schemes β see MIGRATION_WHITEPACT_V2.md) to any
MCP-compatible client β Claude Code, Claude Desktop, Cursor, Windsurf, or your
own agent runtime. Three transports are supported: stdio, Streamable HTTP
(/mcp, current MCP spec), and legacy HTTP+SSE (/sse + /messages/, kept
for older clients). When a team's client points at this server, every AI
interaction is automatically governed β five-way governance decisions, trust
scoring, guardrails, compliance checks (NIST AI RMF / EU AI Act / ISO 42001),
bias evaluation, drift detection, cost tracking, and hash-chained audit
evidence run on any call without code changes.
Setup
bash
# Install
pip install "rai-governance-platform[dashboard,mcp]"# Start the REST API (MCP tools call it internally)
RAI_DB_PATH=/var/lib/rai/governance.db \
RAI_API_KEYS=your-key-here \
uvicorn responsibleai.dashboard.app:app --host 127.0.0.1 --port 8765 &
# Add to Claude Code (~/.claude/claude_desktop_config.json or via /mcp)
whitepact-mcp and responsibleai-mcp are the same entry point β see
pyproject.toml's [project.scripts]; both will keep working, use whichever
name you prefer.
Available tools (27)
Tool
What it does
rai_scan
Detect and redact PII + harmful content before it reaches a log
rai_trust_score
Composite AI Trust Score (0-100) across 6 governance dimensions
rai_compliance
NIST AI RMF / EU AI Act / ISO 42001 compliance evaluation
rai_hallucination
Hallucination risk from hedging, consistency, unsupported claims
Free public Trust Index lookup for a third-party model/tool, before an agent invokes it β unlike every other tool above, which evaluates output the caller itself produced
Agent-framework integrations β LangChain, LangGraph, Google ADK
src/responsibleai/integrations/ wires rai_check_trust directly into three
agent frameworks so an agent can be gated on a tool's public trust score
before invoking it, not just log the call after the fact:
LangChain (langchain_middleware.py) β TrustGateMiddleware, a
wrap_tool_call middleware that blocks a call outright when its score is
below threshold. Requires pip install "rai-governance-platform[langchain]".
LangGraph (langgraph_gate.py) β make_trust_gate_node(), a node that
pauses the graph with interrupt() for a human approve/reject decision on
a below-threshold call, instead of a hard block. Requires
pip install "rai-governance-platform[langgraph]".
Google ADK (adk_toolset.py) β build_stdio_toolset() /
build_http_toolset(), thin factories over ADK's McpToolset, which
auto-discovers this project's MCP server's tools with no custom glue code.
Requires pip install "rai-governance-platform[adk]".
All three, or any subset, install via pip install "rai-governance-platform[agent-frameworks]".
See GAME_CHANGER_BUILD_PLAN.md Phase B for the reasoning behind each.
Available resources (20)
10 canonical resources, each advertised under both the whitepact:// and
rai:// URI schemes (dual scheme is additive β see
MIGRATION_WHITEPACT_V2.md; the table below shows the canonical URI):
Resource
URI
Contents
Health
whitepact://health
Current health status of the governance service
Model pricing catalog
whitepact://models/catalog
Supported models with per-token pricing
Compliance frameworks
whitepact://compliance/frameworks
NIST AI RMF, EU AI Act, ISO 42001
Red team categories
whitepact://redteam/categories
Adversarial attack categories
Trust dimensions
whitepact://trust/dimensions
The 6 dimensions behind the Trust Score
Bias probe catalog
whitepact://bias/probes
Available bias probes and scoring interpretation
Governance policy template
whitepact://governance/policy
Default policy template for rai_policy_check
Trust grade reference
whitepact://trust/grades
Grade thresholds, risk tiers, deployment guidance
NIST AI RMF checklist
whitepact://compliance/checklist/nist
Actionable NIST implementation checklist
EU AI Act checklist
whitepact://compliance/checklist/eu-ai-act
Compliance checklist for high-risk operators
MCP directory listings
WhitePact is listed and queryable today on real MCP directories β not
aspirational, all verified live:
Official MCP Registry β server.json at the repository root
(schema 2025-12-11, listing version 1.2.3) is published as
io.github.Guruprasath-Annadurai/whitepact, confirmed queryable at
registry.modelcontextprotocol.io.
Advertises both the PyPI/stdio package (whitepact-mcp, self-hosted,
free, unrestricted) and a remotes entry pointing at the hosted
Streamable HTTP and SSE transports (whitepact-mcp-http.onrender.com)
β a one-click remote connector, not just an installable package.
Antigravity CLI plugin β plugins/whitepact/ at the repository
root follows the official Antigravity plugin manifest
format, connecting to the
same hosted Streamable HTTP transport via serverUrl. No official
Antigravity plugin directory exists yet, so this is distributed
directly from the repo β see plugins/whitepact/README.md.
Smithery β listed as
guruprasathannadurai-official/whitepact,
30 tools and 20 resources discovered against the hosted Streamable
HTTP transport (whitepact-mcp-http.onrender.com/mcp, a separate
Render service from the main dashboard). This deployment has no
OAuth authorization server configured β only static Bearer API
keys β so a public, unauthenticated
/.well-known/mcp/server-card.json serves the same live
TOOL_DEFS/RESOURCE_DEFS the server itself advertises, for
directories whose scanners can't complete a live authenticated
crawl.
See compliance/MCP_DISTRIBUTION_GUIDE.md for the full distribution
plan, including directories not yet submitted to.
Platform integrations
WhitePact connects to the major AI platforms as one MCP server through
standards-compliant clients β no per-platform forks, no per-platform
governance logic. See docs/integrations/ for the
canonical compatibility matrix (PLATFORM_COMPATIBILITY.md), per-platform
setup docs (GitHub Copilot, Microsoft Copilot, Claude, Grok, Gemini,
Amazon Q, AWS Bedrock AgentCore, Mistral Le Chat, Cursor), and
FOUNDER_ACTIONS.md for what still needs a human. Run
python scripts/integration_smoke.py for a live protocol-level preflight
against the hosted endpoint.
from responsibleai import GuardrailsEngine
guardrails = GuardrailsEngine()
result = guardrails.scan("Customer SSN is 123-45-6789, email: alice@company.com")
print(result.is_blocked) # Trueprint(result.pii_count) # 2print(result.redacted_text) # "Customer SSN is [SSN], email: [EMAIL]"
Hallucination detection
python
from responsibleai import HallucinationDetector
detector = HallucinationDetector()
result = detector.analyze(
"AI will replace all human jobs by 2025.",
candidates=[
"AI will automate some repetitive tasks.",
"AI creates new job categories alongside displacing others.",
],
)
print(f"Risk: {result.hallucination_risk:.2f} Level: {result.risk_level}")
Free, public self-assessment against the open Trust Index standard
GET
/api/trust-index/verify/{passport_id}
Verify a cited Trust Index score (no auth)
GET
/api/trust-index/check
Free, public β trust score + incident count for a named model/tool, by exact name (no auth); what rai_check_trust and the LangChain/LangGraph/ADK integrations call
GET
/api/trust-index/registry
Every assessed model/tool, certified and self-reported, newest first (no auth) β data source for the public /registry page
GET
/api/trust-index/certified
Directory of certified passports (no auth)
POST
/api/trust-index/certify/{passport_id}
Certify a passport β super-admin only
GET
/api/trust-index/badge/{passport_id}.svg
Embeddable trust badge (Self-Assessed / Certified), no auth
POST
/api/incident-db/report
Report a publicly observed AI incident (no auth, rate-limited)
GET
/api/incident-db
Browse published incidents β filter by model, provider, severity, type (no auth)
GET
/api/incident-db/check
Pre-deployment exact-match incident check for a model/provider β PRO/ENTERPRISE
GET
/api/incident-db/verify
Recompute the hash chain over every published entry (no auth)
POST
/api/orgs/{org_id}/keys/{key_id}/mfa/enroll
Enroll an API key in TOTP MFA
POST
/api/orgs/{org_id}/keys/{key_id}/mfa/verify
Verify a TOTP code / backup code
GET/POST
/api/governance/evidence
Read/write hash-chained governance evidence records
GET/POST
/api/governance/approvals
Queue and resolve REQUIRE_APPROVAL decisions
Interactive docs at /api/docs. Public leaderboard page at /leaderboard β
see compliance/LEADERBOARD_METHODOLOGY.md for the published scoring
methodology and scripts/run_leaderboard_eval.py to run evaluations. Open
Trust Index standard and passport verification at /verify/{id} β see
compliance/TRUST_INDEX_SPEC.md. Free, zero-signup self-assessment at
/assess; browse every assessed model/tool at /registry. /llms.txt
points AI crawlers/answer engines at these as canonical sources β see
GAME_CHANGER_STRATEGY.md for why.
Schema changes are managed with Alembic. Run alembic history for the
current, authoritative migration count and table list β this number changes
frequently enough that a hardcoded count here goes stale fast; the command
itself is the source of truth.
bash
# Upgrade to latest schema
RAI_DB_PATH=/var/lib/rai/governance.db alembic upgrade head# PostgreSQL
RAI_DB_URL=postgresql://user:pass@host:5432/responsibleai alembic upgrade head# Show migration history
alembic history# Generate a new migration after changing engine.py
alembic revision --autogenerate -m "add_new_column"
All migrations use render_as_batch=True so they run on both SQLite and
PostgreSQL without changes.
Webhook notifications
Register an endpoint and receive signed events when governance thresholds fire.
Deliveries are persisted to the database. If the server restarts during a
retry cycle, the background worker picks up where it left off on next boot.
Retry schedule: 1 s β 5 s β 30 s β 2 min β 10 min.
Verify payloads with the X-RAI-Signature-256: sha256=<hex> header.
Docker
bash
git clone https://github.com/Guruprasath-Annadurai/Whitepact.git
cd Whitepact
python3 -c "import secrets; print(secrets.token_urlsafe(32))"cp .env.example .env# Edit .env β set RAI_API_KEYS
docker compose up -d
# Dashboard: http://localhost:8765# API docs: http://localhost:8765/api/docs
PostgreSQL + Redis (horizontal scaling)
bash
# .env
RAI_DATABASE_URL=postgresql://rai:secret@db-host:5432/responsibleai
RAI_REDIS_URL=redis://redis-host:6379/0
RAI_OTEL_ENDPOINT=http://otel-collector:4318
pip install "rai-governance-platform[dashboard,postgres,redis,telemetry]"# Run migrations before first start
RAI_DB_URL=postgresql://rai:secret@db-host:5432/responsibleai alembic upgrade head
The async database layer uses SQLAlchemy with connection pooling
(pool_size=10, max_overflow=20, pool_pre_ping=True). Rate limiting
switches to Redis-backed storage when RAI_REDIS_URL is set.
BiasBuster β bias evaluation in CI
bash
# Fail CI when demographic bias exceeds threshold
biasbuster run \
--provider openai --model gpt-4o \
--probes gender-bias,racial-bias,cultural-bias \
--threshold 0.20 \
--output report --format html
Full SQLAlchemy URL β takes priority over RAI_DB_PATH
RAI_DATABASE_URL
(unset)
Alias for RAI_DB_URL
RAI_API_KEYS
(empty = auth off)
Comma-separated bearer tokens
RAI_AUTH_ENABLED
true
Toggle auth enforcement
RAI_REDIS_URL
(unset = in-memory)
Redis URL for distributed rate limiting
RAI_RATE_LIMIT_DEFAULT
100/minute
Per-org rate limit (keyed by Bearer token)
RAI_OTEL_ENDPOINT
(unset = disabled)
OTLP HTTP endpoint
RAI_OTEL_SERVICE_NAME
responsibleai
Service name for traces
RAI_ALERT_THRESHOLD
5.0
Trust score drop that triggers drift alert
RAI_MONTHLY_BUDGET_USD
10000.0
Monthly AI spend limit
RAI_LOG_LEVEL
INFO
Log level
RAI_LOG_JSON
true
Structured JSON logs
RAI_HOST
127.0.0.1
Bind address
RAI_PORT
8765
Port
Dual-prefixed WHITEPACT_* equivalents for these are also read where
MIGRATION_WHITEPACT_V2.md documents them β the RAI_* names remain the
primary, always-supported form.
Development
bash
git clone https://github.com/Guruprasath-Annadurai/Whitepact.git
cd Whitepact
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"# Full test suite
PYTHONPATH=src pytest tests/ -ra
# Enterprise SaaS Layer 1 (identity, RBAC, verified principal gate)
PYTHONPATH=src pytest tests/test_enterprise_saas_layer1.py tests/test_enterprise_saas_layer1_pg.py -ra
# Dashboard tests only
RAI_DB_PATH=:memory: RAI_AUTH_ENABLED=false pytest tests/test_dashboard_api.py
# Webhook persistence tests
pytest tests/test_webhook_persistence.py
# MCP server tests
pytest tests/test_mcp_server.py
# Lint + type check
ruff check src/ tests/
mypy src/responsibleai src/biasbuster
Roadmap
See ROADMAP.md for the canonical NOW/NEXT/LATER plan. The list below is a historical, version-by-version changelog summary kept for reference.
v0.1 β BiasBuster: gender probe, 4 providers, CLI, CI integration
v0.2 β Racial / age / religious / occupational probes, HTML reporter, PrivacyLabel federated DP
v0.3 β Cultural bias, intersectional analysis, DeepfakeDetector ensemble
v1.1 β MCP server (10 tools, 5 resources), audit log API, red team API, billing API, Alembic migrations, per-org rate limiting, DB-persisted webhook retry queue
v1.2 β Public Leaderboard, Trust Index/Passports + embeddable badges, AI Incident Database, TOTP MFA, expanded field encryption, DB-persisted webhooks, full dashboard UI rebuild, white-label branding, a genuinely live hosted instance β see CHANGELOG.md for the full list
WhitePact migration (1.2.0 β 1.2.2) β governance decision core, MCP Streamable HTTP + OAuth/OIDC, risk tiering + policy engine, hash-chained evidence, approval workflow, multi-approver quorum + delegation chains, upstream MCP tool discovery, MCP trust/supply-chain scanner, HA Helm deployment, supply chain security (SBOM/provenance), release engineering, open source governance, live listings on the official MCP Registry and Smithery β see MIGRATION_WHITEPACT_V2.md for the full phase-by-phase log and what's still not done
v2.0 onward β see VERSION_ROADMAP.md for the phase-by-phase plan through v6.0
Strategic direction β GAME_CHANGER_STRATEGY.md lays out an infrastructure-first bet (free public trust registry, an agent-native trust-check primitive, AI-answer-engine citability) as an alternative to the enterprise-SaaS path, with GAME_CHANGER_BUILD_PLAN.md breaking it into concrete engineering phases against the current codebase
Security & Open Source Assurance
The official OpenSSF/OSPS BadgeApp project
currently records OpenSSF Best Practices Silver and OSPS Baseline Level 1.
They are voluntary project evidence, not an independent audit, penetration test, SOC 2,
or ISO certification. Current technical and claim boundaries are maintained in
WHITEPACT_TRUST_STATUS.md and
PUBLIC_TRUST_CLAIMS.md.
Release consumers can review the signed-tag evidence,
release process, security policy,
SLSA evidence boundary, and
consumer verification guide. The reusable trusted-builder
pipeline is present on main. Release v1.2.6 completed that path: its wheel and sdist
were reproduced, hashed, attested, independently verified in the publish job, published
to PyPI without rebuilding, hash-matched to PyPI, and attached to the GitHub Release with
the CycloneDX SBOM. Independent consumer verification was repeated on 2026-08-31. The
release-specific evidence is assessed as satisfying SLSA v1.2 Build L3; SLSA is a
conformance framework, not a certification or a guarantee that an artifact is secure.
MACHINE_AUTHORITY_V1.md β inventory of the eight core machine-authority invariants (Delegation Graph, Autonomy Budget, Memory Firewall, Evidence Bundle, and more)
ENFORCEMENT_BOUNDARY.md β precisely where each invariant's authority stops: inline enforcement vs. voluntary chokepoint
LEGACY_TO_MACHINE_AUTHORITY_MAP.md β mapping RBAC/OAuth/IAM concepts onto their WhitePact equivalents, for readers coming from traditional access control
compliance/SOC2_ALTERNATIVE_PATH.md β real, free, independently verifiable trust signals for now; the honest path to a real SOC 2 when there's budget for one
compliance/PROJECT_CONTINUITY_PLAN.md β the access/recovery checklist a second person would need if the founder became unavailable; stated honestly as a plan, not proof of bus-factor redundancy (no second person holds this access yet)