Spec-driven dev workflow MCP: start_* orchestration, memory, GitNexus, quality gates.
mcp-probe-kit is a spec-driven dev workflow MCP server focused on “Know the Context, Feed the Moment.” It provides protocol-level support for introspection and context hydration, and uses delegated orchestration to help an AI understand project intent and choose a precise workflow.
🛠️ Key Features
Introspection
Context Hydration
Delegated Orchestration
🚀 Use Cases
Helping AI understand project intent
Selecting an appropriate workflow based on that context
⚡ Developer Benefits
Spec-driven workflow support at the protocol level
Structured context gathering to “show the Context” to AI
⚠️ Limitations
Available metadata does not include tool names, implementation details, or a tool count (toolCount not provided).
mcp-probe-kit is a protocol-level toolkit designed for developers who want AI to understand project intent, choose a precise workflow, and retain validated experience without flooding the model with internal actions.
🚀 AI-Powered Complete Development Toolkit - Covering the Entire Development Lifecycle
A powerful MCP (Model Context Protocol) server with 24 model-visible tools by default, 30 when Memory is configured, and a 34-tool compatibility surface available through MCP_TOOLSET=full. It covers the complete workflow from product analysis to final release and supports structured output.
Runtime: Node.js 20 or newer. MCP_PROTOCOL_MODE=auto is the default; use legacy or modern only for compatibility diagnosis.
🎬 v4 in action
v4 turns delegated Agent work into an observable and verifiable delivery loop. The animations below are rendered from the same MCP App source shipped in the npm package—not separate marketing mockups.
Feature Workbench — parent-child specs, active step, outputs, evidence, and cross-session recovery.
Memory Center — semantic search, full-content inspection, lifecycle state, evidence, stale marking, and confirmed deletion.
Convergence Gate — blocks closure when steps or requirements/spec/implementation/test/review evidence are incomplete.
Five native MCP Apps: Memory Center, Feature Workbench, Bug Workbench, Product Workbench, and Convergence Gate.
Resumable delegated plans: plan_heartbeat persists real progress; resume_plan restores the next executable step.
Evidence-based convergence: converge gates delivery and long-term Memory writes.
Managed GitNexus Sidecar: version/platform/architecture/Node isolation, integrity verification, real FTS probe, and safe degradation.
Version-locked CLI fallback: project-local probe wrappers reach the same Tool Registry when a host drops the MCP tool lease.
Parent-child specifications: complex releases are decomposed and recursively validated instead of being flattened into one oversized spec.
v3 → v4 Migration Guide - Tool surfaces, protocol, Apps, plan state, Memory, and compatibility
MCP Apps Live Demos - Five real read-only workbenches generated from the shipped App source
✨ Core Features
📦 Tool Surfaces
The default compact surface keeps every independently useful workflow while removing competing internal and maintenance entries from the model context.
🧭 Routing (1) — workflow
🔁 Plan State & Convergence (3) — plan_heartbeat, resume_plan, converge
That is 24 model-visible tools by default. When the full Memory stack is configured, six Memory tools are added dynamically, bringing the model-visible surface to 30:
For compatibility and diagnostics, MCP_TOOLSET=full restores all 34 model tools. The compact surface deliberately omits add_feature, fix_bug, sync_ui_data, and ask_user: their implementations remain available through orchestration, maintenance scripts, or full compatibility mode.
workflow is a fallback tool-selection guide, not a natural-language intent classifier. The Agent normally chooses the appropriate MCP tool directly from the current conversation, Skill, and tool descriptions. scenario=auto returns guidance only (firstTool=null); an explicit scenario returns deterministic guidance for a scenario the Agent has already selected.
🔁 Delegated Plan State, Recovery, and Convergence
Every v4 delegated plan declares executionStatePolicy and instructs the Agent to create a local checkpoint on the first step.
plan_heartbeat persists completed/skipped steps, unresolved items, evidence, and the last verified revision under .mcp-probe-kit/plans/.
resume_plan recalculates ready and blocked steps from stored dependencies after interruption, restart, or Agent handoff.
converge refuses closure while steps, unresolved items, or requirements/spec/implementation/test/review evidence are incomplete. Formal long-term memory writes are allowed only after convergence passes.
These tools track and validate Agent execution; they do not move file, shell, Git, or implementation work into the MCP server.
🛡️ Quality Constraints (single source of truth)
All hard quality rules live in one module (src/lib/quality-constraints.ts) and are injected into code_review, the add_feature task templates, and the UI tools. Change once, apply everywhere — inspired by taste-skill and impeccable.
Code limits: single file ≤ 500 lines (split into modules/components when exceeded), function ≤ 50 lines, nesting ≤ 4, parameters ≤ 3.
Completeness blacklist: code_review flags placeholder/elision patterns (// ..., // TODO, // rest of code, bare ...) as CRITICAL — "a partial output is a broken output".
Anti-laziness task templates: add_feature tasks now carry a Scope-lock deliverable count, a mandatory evidence block (read code before writing), a per-file line budget, and a binary zero-tolerance rule for placeholders. check_spec validates these (missing Scope-lock = error, thin task without evidence = warning).
UI hard red lines: numeric, machine-checkable rules — 4pt spacing scale, WCAG contrast (4.5/3/3), type scale ≥ 1.25, hero font ≤ 6rem, OKLCH, eight interaction states, cognitive load ≤ 4, motion 150-300ms.
UI banned list + Pre-Flight checklist: match-and-refuse blacklist for AI slop (default Inter/Roboto, AI purple-blue gradients, gradient text, cookie-cutter card grids, em-dash, cream/beige body backgrounds, nested cards) plus a delivery-gate self-check matrix.
🧠 Code Graph Bridge (GitNexus)
code_insight bridges GitNexus by default for query/context/impact analysis
The bridge prefers an explicitly configured or system GitNexus CLI, then a version-locked managed Sidecar; GitNexus is not bundled into the main package and is never globally installed
init_project_context bootstraps baseline graph docs under docs/graph-insights/; if docs/project-context.md already exists, it preserves the old context docs and only backfills graph docs plus the index entry
start_feature refreshes the GitNexus index and runs task-level query/context/impact narrowing before spec generation to reduce over-scoping
start_bugfix refreshes the GitNexus index and runs task-level graph analysis before TBP RCA to constrain failure boundary and blast radius
Older projects that already have project-context.md but no graph docs are bootstrapped automatically through the init_project_context step
If GitNexus is unavailable, the server falls back automatically without breaking orchestration
Real graph queries read the .gitnexus index; docs/graph-insights/latest.md|json are readable snapshots for humans and AI agents
MCP resources in MCP client settings list 2 entries (probe://status, probe://project/bootstrap). Graph runtime snapshots (probe://graph/latest, etc.) and probe://project/skill|agents|context|graph remain readable via resources/read when tools expose URIs
Graph snapshots are persisted to .mcp-probe-kit/graph-snapshots (customizable via MCP_GRAPH_SNAPSHOT_DIR)
Tool responses include _meta.graph with snapshot URI and local JSON/Markdown file paths
start_bugfix runs graph narrowing, then delegated SRC-8 plan (metadata.plan.steps src8-1~8) before repair and tests
fix_bug returns delegated plan (src8-1~8), src8Checklist, rootCauseWorksheet (Step 4 core), and hard gates (no code change until root-cause worksheet is closed)
Highlights vs manufacturing TBP: repro contract, attribution layers (including agent_behavior), contributing factors, memorize_asset for cross-repo learning
Inherited from Toyota TBP: gap thinking, Plan-before-Do, no skipping to root-cause analysis, fact-based investigation, countermeasures over symptoms, evaluate then standardize.
Our elevation: genchi-genbutsu → read code/logs/repro; Step 4 worksheet; guidance-only MCP that forces discipline while the Agent executes.
🧠 Memory Retrieval
Memory tools use Qdrant as the vector database backend
Embedding service supports two modes:
ollama
openai-compatible
Memory tools:
search_memory - Semantic search across the shared memory pool (optionally prefer type / tags); text output includes id, score, summary, description, and a --- content --- body (default up to 1500 chars via MEMORY_SEARCH_CONTENT_MAX_CHARS)
memorize_asset - Persist an already validated MemoryCandidate into vector memory; for delegated workflows, call it only after converge passes
read_memory_asset - Read full asset content by asset_id (text output includes the full content body)
update_memory_asset - Update an existing asset by asset_id (preserves ID; content changes re-embed)
delete_memory_asset - Delete an asset by asset_id from the shared pool
scan_and_extract_patterns - Extract reusable patterns from code/file/directory before deciding whether to persist
Cross-repo memory pools: do not rely on source_project / source_path for shared retrieval; put file paths in content instead. Search injection hides foreign sourcePath unless MEMORY_REPO_ID matches or MEMORY_SEARCH_SHOW_SOURCE=true.
Core and orchestration tools support structured output, returning machine-readable JSON data, improving AI parsing accuracy, supporting tool chaining and state tracking.
⏱️ Native Tasks, Progress, and Cancellation
Uses an SDK-independent Internal Task Runtime, with the current SDK task protocol exposed through a Legacy Adapter
Advertises capabilities.tasks.requests.tools.call so clients can create tasks for tools/call
Falls back to synchronous execution when protocol task storage is unavailable
Emits notifications/progress when client provides _meta.progressToken
Ignores late progress after terminal completion; tool/task result is the final completion signal
Handles request cancellation via AbortSignal and preserves a clear cancelled state
Long-running orchestration tools (start_*) and sync_ui_data support cooperative cancellation/progress callbacks
Internal task persistence defaults to memory. Set MCP_TASK_STORE=json to use .mcp-probe-kit/tasks.json, or set MCP_TASK_STORE_PATH to choose another JSON path. Interrupted tasks that cannot reconstruct their executor are explicitly marked failed on restart instead of being reported as still running.
🔌 Official MCP Apps and Memory Center
v4.0.0 uses the official @modelcontextprotocol/ext-apps SDK and the stable io.modelcontextprotocol/ui extension.
MCP Apps are enabled by default and can be disabled with MCP_ENABLE_UI_APPS=0.
UI metadata and ui:// resources are exposed only after the client advertises support for text/html;profile=mcp-app.
Five stable Apps are included: Memory Center, Feature Workbench, Bug Workbench, Product Workbench, and Convergence Gate.
Memory Center uses a responsive master-detail layout for historical browsing, semantic search, full-content inspection, lifecycle state, evidence, stale marking, and confirmed deletion.
Feature and Bug Workbenches render a live plan stepper. The App polls resume_plan while visible, and progress advances only after the Agent records real step state through plan_heartbeat.
Product Workbench and Convergence Gate use the same developer-console design system for delivery paths, blockers, and evidence gaps.
list_memory_assets is an App-only action with _meta.ui.visibility=["app"]. It may appear in the raw tools/list response of an Apps-capable host, but compliant hosts must not offer it to the model. The model-visible count remains 24 by default or 30 with Memory.
Clients without MCP Apps support continue to receive the normal text and structuredContent responses; no GUI capability is required for existing workflows.
Trace metadata passthrough remains available through MCP_ENABLE_EXTENSIONS_CAPABILITY=1.
🧪 Tool and Real-Agent Contract Verification
bash
# Deterministic server-side audit across compact, Memory, full, App-only, and Legacy surfaces
npm run audit:tools
# Optional real-host audit: Claude Code calls and evaluates all 34 model tools
npm run audit:tools:agent
The direct audit verifies non-empty readable text, structuredContent, and that every referenced MCP tool exists on the active surface. The real-Agent audit additionally checks whether an Agent understands each tool, can follow the returned guidance, sees no text/structured contradiction, and can execute the stated next step. It is intentionally separate from release:verify because it requires a configured Claude Code account and incurs model usage.
🧭 Delegated Orchestration Protocol
All start_* orchestration tools return an execution plan in structuredContent.metadata.plan.
AI needs to call tools step by step and persist files, rather than the tool executing internally.
When requirements are unclear, use requirements_mode=loop in start_feature / start_bugfix / start_ui.
This mode performs 1-2 rounds of structured clarification before entering spec/fix/UI execution.
add_feature supports template profiles, default auto auto-selects: prefers guided when requirements are incomplete (includes detailed filling rules and checklists), selects strict when requirements are complete (more compact structure, suitable for high-capability models or archival scenarios).
Example:
json
{"description":"Add user authentication feature","template_profile":"auto"}
Applicable Tools:
start_feature passes template_profile to add_feature
start_bugfix / start_ui also support template_profile for controlling guidance strength (auto/guided/strict)
Template Profile Strategy:
guided: Less/incomplete requirements info, regular model priority
strict: Requirements structured, prefer more compact guidance
For version-level or epic work, start_feature defaults to spec_layout: "auto" and selects parent-child when the requirement spans multiple modules, stages, or capability domains. If child boundaries are not known yet, the delegated plan first returns a decompose-spec step. You can still explicitly pass flat or parent-child; add_feature remains an atomic tool and defaults to flat unless the layout and subspecs are already defined. The MCP server returns templates and pendingFiles; the calling Agent creates the parent spec, spec-manifest.json, and child specs after review. check_spec then validates the complete hierarchy recursively.
start_feature uses query-only GitNexus narrowing with an 8-second degradation budget, so graph cold starts do not block specification planning. Automatic index refresh is disabled by default; set MCP_GITNEXUS_AUTO_REFRESH=1 when the MCP process should refresh the index before graph queries.
json
{"feature_name":"commerce-v2","description":"Upgrade the commerce domain while preserving v1 compatibility","spec_layout":"parent-child","subspecs":[{"id":"01-foundation","title":"Data foundation","fr":["FR-1"]},{"id":"06-inventory-ledger","title":"Inventory ledger","fr":["FR-2"],"dependsOn":["01-foundation"]}]}
🔄 Workflow Orchestration
6 intelligent orchestration tools that automatically combine multiple basic tools for one-click complex development workflows:
start_feature - New feature development (Requirements → Design → Estimation)
If some skills are missing, workflow continues with MCP main plan and marks unavailable skills in metadata.
Why use sync_ui_data?
Our start_ui tool relies on a rich UI/UX database (colors, icons, charts, components, design patterns, etc.) to generate high-quality design systems and code. This data comes from npm package uipro-cli, including:
🎨 Color schemes (mainstream brand colors, color palettes)
Skill & AGENTS auto-bootstrap (v3.6.3+): Every MCP tool call writes .agents/skills/mcp-probe-kit/SKILL.md and merges the mcp-probe:context block into AGENTS.md. Workspace root is auto-detected (Cursor injects WORKSPACE_FOLDER_PATHS; OpenCode project opencode.json sets cwd). No per-client MCP_PROJECT_ROOT unless global MCP cannot resolve the workspace — then set MCP_PROJECT_ROOT or pass project_root in tool args.
Multi-harness adapters (v3.6.8+): AGENTS.md and the canonical Skill stay the single rule source. If the project already has .trae/, .lingma/, .comate/, .codebuddy/, or .claude/, matching thin adapters (skill mirror or rules pointer) are written automatically — no env vars.
Version-locked CLI fallback (v4.0.0+): Bootstrap also writes .mcp-probe-kit/bin/probe.cmd|probe.ps1|probe and .mcp-probe-kit/runtime.json. If a modified host or third-party Agent provider connects the MCP server but omits its tools from the Agent session, the generated Skill and Cursor rule instruct the Agent to invoke the same Tool Registry through the project wrapper. The wrapper pins the exact MCP package version, does not install globally, and does not modify the project's package.json.
Memory for CLI fallback: install-agent also creates .mcp-probe-kit/local.env (and local.env.example). The CLI fallback path (probe.* exec ...) does not inherit IDE mcp.json env; edit local.env with the same MEMORY_* keys.
Direct CLI examples:
bash
# JSON from stdin is the most portable optionprintf'%s''{"intent":"build a task board","scenario":"feature","project_root":"."}' \
| ./.mcp-probe-kit/bin/probe exec workflow --stdin
# Repair or install the project wrappers without a working MCP tool lease
npx --yes mcp-probe-kit@<exact-version> install-agent --project-root .
Note: OpenCode uses opencode.json with a different schema from Cursor/Claude Desktop. The key mcp replaces mcpServers, command is an array, type: "local" is required, and environment variables use environment instead of env. See OpenCode MCP docs for details.
No memory backend: scan_and_extract_patterns (local scan only; persist via memorize_asset when ready)
For full write/search you need both:
A Qdrant vector database
An embedding service in either ollama or openai-compatible mode
Note (CLI fallback): If you run the project wrapper (./.mcp-probe-kit/bin/probe* exec ...) instead of native MCP, Memory env is read from .mcp-probe-kit/local.env (created by install-agent).
Full guide (Docker Compose for Qdrant + Infinity, ports 50008 / 50012, MCP env, smoke tests):
MEMORY_EMBEDDING_PROVIDER: ollama or openai-compatible
MEMORY_EMBEDDING_URL: Embedding endpoint URL
MEMORY_EMBEDDING_API_KEY: Optional for Ollama, usually required for hosted OpenAI-compatible providers
MEMORY_EMBEDDING_MODEL: Default is nomic-embed-text
MEMORY_SEARCH_LIMIT: Default search result count is 3
MEMORY_SUMMARY_MAX_CHARS: Default summary truncation length is 280
Notes
Memory write capability is enabled only when MEMORY_QDRANT_URL, MEMORY_EMBEDDING_URL, and MEMORY_EMBEDDING_MODEL are configured
Memory read capability only requires MEMORY_QDRANT_URL
Qdrant collections are auto-created on first write with Cosine distance
Vector size is inferred from the first embedding response
GitNexus Managed Runtime
Applies to code_insight, start_feature, start_bugfix, and init_project_context.
GitNexus is not bundled into the mcp-probe-kit npm tarball because it includes native, platform-specific dependencies and uses the PolyForm Noncommercial license. The runtime policy is:
Use MCP_GITNEXUS_COMMAND when explicitly configured.
Otherwise reuse an already validated managed Sidecar from the mcp-probe-kit user cache.
Otherwise use a compatible gitnexus CLI already available on PATH.
If no runtime is installed, graph analysis degrades immediately instead of blocking the main workflow. The Agent can run doctor gitnexus --install and retry automatically.
Validated compatibility:
Node.js
Managed GitNexus
20-21
Managed Sidecar disabled; use a system GitNexus CLI or degraded mode
22+ / Windows、macOS、Linux
1.6.9
Each managed installation is isolated by GitNexus version, operating system, CPU architecture, and Node.js major version. npm integrity is checked against the pinned release metadata before the runtime is accepted. The installer then runs gitnexus doctor plus a real TypeScript indexing probe and rejects any runtime that silently disables FTS/BM25 search.
Install or repair the managed Sidecar through the project launcher:
powershell
# Windows
& ./.mcp-probe-kit/bin/probe.cmd doctor gitnexus --install
bash
# macOS / Linux
./.mcp-probe-kit/bin/probe doctor gitnexus --install
The first installation can take several minutes because GitNexus includes native parsers, LadybugDB, ONNX Runtime, and post-install grammar builds. It runs outside the project and does not modify the project package.json or node_modules.
Available modes:
MCP_GITNEXUS_MODE=auto — default; explicit/system/existing managed runtime, otherwise fast degradation.
MCP_GITNEXUS_MODE=managed — require the managed Sidecar and allow installation during the graph request.
MCP_GITNEXUS_MODE=system — use only explicit/system GitNexus; never install.
MCP_GITNEXUS_MODE=off — disable GitNexus.
MCP_GITNEXUS_AUTO_INSTALL=1 — allow auto mode to install synchronously; not recommended for latency-sensitive clients.
Some GitNexus dependencies use native modules. On Windows, LadybugDB FTS also requires the OpenSSL runtime shipped with Git for Windows; mcp-probe-kit discovers its mingw64/bin directory and exposes it only to the managed child process. Set MCP_GITNEXUS_WINDOWS_RUNTIME_BIN to an equivalent directory when Git is installed in a nonstandard location. A failed prebuilt-binary download may still require Visual Studio Build Tools with the C++ workload. Installation failure never prevents the mcp-probe-kit workflow from continuing in degraded mode.
npx -y mcp-probe-kit@4.0.0 2>&1 | tee ./mcp-probe-kit.log
Q2: Client not recognizing tools after configuration?
Restart client (completely quit then reopen)
Check config file path is correct
Confirm JSON format is correct, no syntax errors
Check client developer tools or logs for error messages
Q2b: Cursor shows connected but 0 tools / Agent says No MCP servers available?
This is a known Cursor-side issue: stderr may report a valid compact tool surface, while Mcp FileSystem Writer shows lease returned 0 tools and toolCount=0 — the Agent lease layer silently dropped the tool list.
Common causes:
Symptom in logs
Likely cause
tools/list ≈ 50+ KB then lease returned 0 tools
Cursor internal payload size limit (whole list dropped silently)
Windows mcpProcess utility failed; legacy fallback discovers tools but Agent lease stays empty
Settings green dot, Agent No MCP servers available
Renderer ↔ shared-process MCP routing not wired for this session
What we do:tools/list omits outputSchema by default, and v4.0.0 defaults to the 24-tool compact model surface. Structured output still works through structuredContent on tools/call. Restore output schemas with MCP_INCLUDE_OUTPUT_SCHEMA=1, or restore the 34-tool compatibility surface with MCP_TOOLSET=full.
What you can try:
Reload MCP or fully quit Cursor (not just close window) and reopen
Check Output → MCP for lease returned 0 tools / ipcReady / MessagePort
In Composer, open the tools panel — ensure the server toggle is on (some versions default off)
Upgrade Cursor (3.7.36+ had Windows ipcReady regressions; try latest or roll back to a known-good build)
If still broken after server update, report to Cursor with: connected=true, stderr tool count, lease toolCount=0, and shared-process MCP routing disabled
Fallback when the Host Agent path is replaced or does not bridge MCP tools:
If the MCP panel and tool cache are healthy but the actual Agent request is handled by a third-party provider with no MCP tool bridge, restarting the server cannot fix that path. Use the project wrapper generated by bootstrap:
powershell
# Windows
'{"intent":"continue the current feature","scenario":"feature","project_root":"."}' |
.\.mcp-probe-kit\bin\probe.cmd exec workflow --stdin
bash
# macOS / Linuxprintf'%s''{"intent":"continue the current feature","scenario":"feature","project_root":"."}' \
| ./.mcp-probe-kit/bin/probe exec workflow --stdin
The Skill automatically selects this route when native MCP tools are absent. plan_heartbeat, resume_plan, and converge use the same project files across separate CLI processes and native MCP sessions.
This folder is written by Cursor (Mcp FileSystem Writer), not by mcp-probe-kit. After a successful tool lease you should see:
text
mcps/user-mcp-probe-kit/
├── SERVER_METADATA.json
├── STATUS.md
├── tools/ ← one JSON per model-visible tool (~24 by default); Agent reads these for CallMcpTool
│ ├── init_project.json
│ └── ...
└── resources/ ← from resources/list (may exist even when tools/ is empty)
State
Meaning
resources/ exists, tools/ missing or empty
resources/list OK but tools lease failed (matches lease returned 0 tools)
tools/ has fewer entries than the selected model surface (24 default, 30 with Memory, 34 full)
Partial write or session interrupted; Reload MCP
STATUS.md says server errored
Cursor marked the server unhealthy for Agent even if Settings is green
Healthy session: tools/ should auto-populate within seconds of MCP connect — no manual setup, no repo config.
Q3: How to update to latest version?
npx method (Recommended):
Use @latest tag in config, automatically uses latest version.
Global installation method:
bash
npm update -g mcp-probe-kit
Q4: Why can the first GitNexus installation take a long time?
GitNexus includes native parsers, a graph database, ONNX Runtime, and post-install grammar builds. A cold managed installation may take several minutes, especially on Windows or a slow network.
The normal feature and bug-fix workflows do not wait for this installation in default auto mode. They return a structured managed_install_required degradation result, and the Agent can automatically run:
powershell
& ./.mcp-probe-kit/bin/probe.cmd doctor gitnexus --install
The installation is stored in the mcp-probe-kit user cache, uses an exact compatible version and npm integrity pin, and does not modify the business project. If native installation fails, graph analysis remains degraded while the rest of the workflow continues normally.