Metacognitive AI agent oversight: adaptive CPI interrupts for alignment, reflection and safety
io.github.PV-Bhat/vibe-check-mcp-server (MCP)
The io.github.PV-Bhat/vibe-check-mcp-server is an MCP server focused on “metacognitive AI agent oversight.” It provides adaptive CPI interrupts intended for alignment, reflection, and safety. Its documentation also notes the project’s ongoing maintenance status and community support under the MIT license.
🛠️ Key Features
Metacognitive AI agent oversight
Adaptive CPI interrupts for alignment
CPI interrupts for reflection and safety
MCP and MCP-server support
🚀 Use Cases
Oversight for agentic AI workflows
Agent alignment and reflection during execution
Safety-oriented interrupt handling for agent actions
⚡ Developer Benefits
Standardized integration via Model Context Protocol (MCP)
Topics cover workflow automation, error-handling, and “cpi” / “rli”
Research claims: success improved +27%, harmful actions reduced -41%
⚠️ Limitations
In maintenance mode: active feature development ended
Only security and bug fixes are published; v2.9.0 is latest maintenance release
This project is in maintenance mode. Active feature development has ended; only maintenance patches (security and bug fixes) are published. v2.9.0 is the latest maintenance release. The server remains fully functional. Community forks and contributions are welcome under the MIT license.
KISS overzealous agents goodbye. Plug & play agent oversight tool.
Based on research:
In our study agents calling Vibe Check improved success +27% and halved harmful actions -41%
Featured on PulseMCP “Most Popular (This Week)” • 5k+ monthly calls on Smithery.ai • research-backed oversight • STDIO + streamable HTTP transport
Gemini_Generated_Image_kvdvp4kvdvp4kvdv
Plug-and-play mentor layer that stops agents from over-engineering and keeps them on the minimal viable path — research-backed MCP server keeping LLMs aligned, reflective and safe.
curl http://127.0.0.1:2091/healthz to confirm the service is live.
Send JSON-RPC requests to http://127.0.0.1:2091/mcp.
npx downloads the package on demand for both options. For detailed client setup and other commands like install and doctor, see the documentation below.
Vibe Check MCP keeps agents on the minimal viable path and escalates complexity only when evidence demands it. Vibe Check MCP is a lightweight server implementing Anthropic's Model Context Protocol. It acts as an AI meta-mentor for your agents, interrupting pattern inertia with Chain-Pattern Interrupts (CPI) to prevent Reasoning Lock-In (RLI). Think of it as a rubber-duck debugger for LLMs – a quick sanity check before your agent goes down the wrong path.
Overview
Vibe Check MCP pairs a metacognitive signal layer with CPI so agents can pause when risk spikes. Vibe Check surfaces traits, uncertainty, and risk scores; CPI consumes those triggers and enforces an intervention policy before the agent resumes. See the CPI integration guide and the CPI repo at https://github.com/PV-Bhat/cpi for wiring details.
Vibe Check invokes a second LLM to give meta-cognitive feedback to your main agent. Integrating vibe_check calls into agent system prompts and instructing tool calls before irreversible actions significantly improves agent alignment and common-sense. The high-level component map: docs/architecture.md, while the CPI handoff diagram and example shim are captured in docs/integrations/cpi.md.
The Problem: Pattern Inertia & Reasoning Lock-In
Large language models can confidently follow flawed plans. Without an external nudge they may spiral into overengineering or misalignment. Vibe Check provides that nudge through short reflective pauses, improving reliability and safety.
Key Features
Feature
Description
Benefits
CPI Adaptive Interrupts
Phase-aware prompts that challenge assumptions
alignment, robustness
Multi-provider LLM
Gemini 3.6, Claude 5, GPT-5.6, and OpenRouter support
flexibility
History Continuity
Summarizes prior advice when sessionId is supplied
context retention
Optional vibe_learn
Log mistakes and fixes for future reflection
self-improvement
What's New in v2.9.0 (Security & Model Refresh)
Maintenance Notice: This project is in maintenance mode and is no longer under active feature development. It remains fully functional and available under the MIT license. Community forks are welcome. For details, see the Changelog.
Current models: Gemini 3.6 Flash, Claude Sonnet 5 / Opus 5 / Fable 5, and GPT-5.6 Sol / Terra / Luna are now the supported defaults, defined in one registry (src/utils/models.ts)
Native Google AI Studio: migrated from the retired @google/generative-ai package to the unified @google/genai SDK
HTTP hardening: CORS now defaults to loopback origins instead of *, Host headers are validated to block DNS rebinding, and the JSON body cap is explicit and validated
Security:npm audit is clean — 10 advisories resolved across axios, the MCP SDK's Hono stack, form-data, fast-uri, postcss and the test toolchain
Dependencies: MCP SDK 1.29, axios 1.18, OpenAI SDK 6.x, vitest 4.x; the unused body-parser direct dependency was dropped
Session Constitution (per-session rules)
Use a lightweight “constitution” to enforce rules per sessionId that CPI will honor. Eg. constitution rules: “no external network calls,” “prefer unit tests before refactors,” “never write secrets to disk.”
API (tools):
update_constitution({ sessionId, rules }) → merges/sets rule set for the session
Gemini runs natively against Google AI Studio (the Gemini Developer API) through the unified @google/genai SDK. Any model ID the provider accepts will work — the table lists the defaults and the suggestions surfaced to agents in the vibe_check tool schema.
Set the default globally with DEFAULT_LLM_PROVIDER / DEFAULT_MODEL, or per call with modelOverride. DEFAULT_MODEL names a model of DEFAULT_LLM_PROVIDER; a call that overrides the provider without naming a model falls through to that provider's default rather than reusing it.
If a Gemini call fails, the server retries once against gemini-3.5-flash-lite before falling back to static questions.
HTTP transport hardening
These apply only to --http mode; stdio is unaffected.
Variable
Default
Purpose
CORS_ORIGIN
loopback origins only
Comma-separated browser origin allowlist. * restores the pre-2.9 wildcard.
MCP_ALLOWED_HOSTS
localhost, 127.0.0.1, ::1
Host header allowlist (DNS-rebinding protection). * disables the check.
MCP_MAX_BODY_SIZE
100kb
JSON body cap. Unparseable values are ignored rather than silently disabling enforcement.
Upgrading to v2.9.0 over HTTP: if you serve Vibe Check on a non-loopback hostname (Docker, a reverse proxy, a hosted deployment), set MCP_ALLOWED_HOSTS to that hostname — or * — or requests will be rejected with HTTP 403.
See API Keys & Secret Management for supported providers, resolution order, storage locations, and security guidance.
Transport selection
The CLI supports stdio and HTTP transports. Transport resolution follows this order: explicit flags (--stdio/--http) → MCP_TRANSPORT → default stdio. When using HTTP, specify --port (or set MCP_HTTP_PORT); the default port is 2091. The generated entries add --stdio or --http --port <n> accordingly, and HTTP-capable clients also receive a http://127.0.0.1:<port> endpoint.
Client installers
Each installer is idempotent and tags entries with "managedBy": "vibe-check-mcp-cli". Backups are written once per run before changes are applied, and merges are atomic (*.bak files make rollback easy). See docs/clients.md for deeper client-specific references.
Claude Desktop
Config path: claude_desktop_config.json (auto-discovered per platform).
Default transport: stdio (npx … start --stdio).
Restart Claude Desktop after installation to load the new MCP server.
If an unmanaged entry already exists for vibe-check-mcp, the CLI leaves it untouched and prints a warning.
Cursor
Config path: ~/.cursor/mcp.json (provide --config if you store it elsewhere).
Schema mirrors Claude’s mcpServers layout.
If the file is missing, the CLI prints a ready-to-paste JSON block for Cursor’s settings panel instead of failing.
Windsurf (Cascade)
Config path: legacy ~/.codeium/windsurf/mcp_config.json, new builds use ~/.codeium/mcp_config.json.
Pass --http to emit an entry with serverUrl for Windsurf’s HTTP client.
Existing sentinel-managed serverUrl entries are preserved and updated in place.
Visual Studio Code
Workspace config lives at .vscode/mcp.json; profiles also store mcp.json in your VS Code user data directory.
Provide --config <path> to target a workspace file. Without --config, the CLI prints a JSON snippet and a vscode:mcp/install?... link you can open directly from the terminal.
VS Code supports optional dev fields; pass --dev-watch and/or --dev-debug <value> to populate dev.watch/dev.debug.
Uninstall & rollback
Restore the backup generated during installation (the newest *.bak next to your config) to revert immediately.
To remove the server manually, delete the vibe-check-mcp entry under mcpServers (Claude/Windsurf/Cursor) or servers (VS Code) as long as it is still tagged with "managedBy": "vibe-check-mcp-cli".
Research & Philosophy
CPI (Chain-Pattern Interrupt) is the research-backed oversight method behind Vibe Check. It injects brief, well-timed “pause points” at risk inflection moments to re-align the agent to the user’s true priority, preventing destructive cascades and reasoning lock-in (RLI). In pooled evaluation across 153 runs, CPI nearly doubles success (~27%→54%) and roughly halves harmful actions (~83%→42%). Optimal interrupt dosage is ~10–20% of steps. Vibe Check MCP implements CPI as an external mentor layer at test time.
flowchart TD
A[Agent Phase] --> B{Monitor Progress}
B -- high risk --> C[CPI Interrupt]
C --> D[Reflect & Adjust]
B -- smooth --> E[Continue]
Agent Prompting Essentials
In your agent's system prompt, make it clear that vibe_check is a mandatory tool for reflection. Always pass the full user request and other relevant context. After correcting a mistake, you can optionally log it with vibe_learn to build a history for future analysis.
Example snippet:
code
As an autonomous agent you will:
1. Call vibe_check after planning and before major actions.
2. Provide the full user request and your current plan.
3. Optionally, record resolved issues with vibe_learn.
When to Use Each Tool
Tool
Purpose
🛑 vibe_check
Challenge assumptions and prevent tunnel vision
🔄 vibe_learn
Capture mistakes, preferences, and successes
🧰 update_constitution
Set/merge session rules the CPI layer will enforce
This repository includes a CI-based security scan that runs on every pull request. It checks dependencies with npm audit and scans the source for risky patterns. See SECURITY.md for details and how to report issues.
Roadmap
Note: This project is in maintenance mode (latest maintenance release: v2.9.0). The roadmap below is preserved for community forks that may wish to continue development.
Structured output for vibe_check: Return a JSON envelope such as { advice, riskScore, traits } so downstream agents can reason deterministically.
LLM resilience: Wrap generateResponse with retries and exponential backoff.
Input sanitization: Validate and cleanse tool arguments to mitigate prompt-injection vectors.
Prompt externalization: Move hardcoded prompts to configuration files for transparency and auditability (see PR #71).