Content-addressed code graph with 22 MCP tools for AI agents.
io.github.blackwell-systems/knowing MCP Server
This MCP server provides a content-addressed code graph for AI agents, described as “knowing.” It exposes 22 MCP tools (and is associated with MCP tooling/resources) and centers on Merkle-tree based representations such as hierarchical Merkle, merkle proofs, and intelligence versioning.
🛠️ Key Features
Content-addressed code graph
22 MCP tools for AI agents
Merkle-tree / hierarchical Merkle structure
Merkle-proof support
Intelligence versioning, audit, and compliance-oriented provenance
Opentelemetry and retrieval pipeline alignment (as indicated by topics)
🚀 Use Cases
Code intelligence for AI agents
Retrieval pipeline workflows over a code graph
Static analysis using a code-memory/developer-memory model
Software supply-chain provenance and audit/compliance tasks
⚡ Developer Benefits
Code-memory and developer-memory oriented tooling
Provenance, audit, and compliance-related capabilities
Intelligence versioning for code-graph state tracking
⚠️ Limitations
Documented surface focuses on tool availability and code-graph concepts; the provided excerpt does not enumerate specific tool names or resource details.
Your architecture diagram says service A calls service B. Can you prove it?
knowing can. It builds a content-addressed graph of extracted code relationships, snapshots it as a Merkle tree tied to a git commit, and generates cryptographic proofs that verify offline. Agents use it for ranked context. Security teams use it for audit. Platform teams use it to compare code against production traces.
It gets better every time you use it. When code changes, stale knowledge expires automatically.
That's it. The MCP server auto-indexes your repo on first launch. No model downloads, no API keys. Your agent now has ranked context, blast radius, test scope, and implicit noise demotion that improves results during active sessions.
Verify it works: Ask your agent: "Use the context_for_task tool to find symbols related to [something you know exists in your code]." You should see ranked symbols with scores and file paths from your codebase. If results are empty, the repo is still indexing (10-30 seconds on first launch). If results seem unrelated, see Troubleshooting.
knowing is three products built on one foundation (content-addressed graph with hierarchical Merkle trees):
1. Context engine for AI agents
One call returns the most relevant symbols for a task, ranked by graph centrality, recency, and learned usefulness, packed to fit your token budget. 263 framework equivalence classes bridge vocabulary gaps when keywords fail. 47% fewer tool calls. 84% fewer tokens. Results improve with feedback.
2. Audit primitive for compliance
Every graph state is a Merkle root tied to a git commit. knowing prove generates a cryptographic proof that a relationship existed. knowing verify checks it offline. knowing fsck verifies the entire graph in 98ms. Supply chain detection extracts credential access, process spawning, and network exfiltration edges to flag structurally suspicious code.
3. Noise demotion that learns
Symbols returned but never used by the agent get demoted on future queries. When code changes, feedback expires automatically (verified via package Merkle roots). The system gets more precise during active sessions. That is the property knowing is built around.
These aren't separate features. They're structural consequences of content-addressing: the same hash that makes context cacheable also makes it provable, and the same Merkle root that detects staleness also expires stale feedback.
What It Answers
For your agent:
"I'm changing this function. What breaks?" (blast radius across callers, tests, routes, repos)
"Give me 50,000 tokens of context for this task." (graph-ranked, not grep-searched)
"Which tests should run?" (call-graph traversal, 98% precision)
For your platform team:
"Is this route used in production?" (static analysis + OTel runtime traces)
"What did the service graph look like at a specific snapshot?" (snapshot chain, each root tied to a git commit)
For your security team:
"Prove service A calls service B at this commit." (Merkle proof, verifiable offline)
"Prove this dependency does NOT exist." (absence proof via sorted leaves)
"Generate a compliance report." (knowing audit -proofs, one command)
"Does this package read credentials and spawn processes?" (knowing audit-supply-chain --scan-all)
All benchmarks are reproducible. The cross-system benchmark (P@10=0.330) uses 17 repos pinned to exact commits with a corpus manifest and setup script for full from-scratch reproduction. See METHODOLOGY.md for protocol details.
Quick Start
Path A: MCP server (recommended for AI agents)
bash
# 1. Install
brew install blackwell-systems/tap/knowing
# Or: npm install -g @blackwell-systems/knowing# Or: pip install knowing# Or: go install github.com/blackwell-systems/knowing/cmd/knowing@latest# 2. Add to your agent config (.mcp.json, Claude Code settings, etc.)# See "MCP Integration" below for the config block.# The server auto-indexes your repo on first launch. Done.
Path B: CLI usage (explore the graph yourself)
bash
# 1. Install (same as above)
brew install blackwell-systems/tap/knowing
# 2. Index your repo
knowing add .
# 3. Verify the index worked
knowing stats
# You should see node and edge counts. A healthy TypeScript repo with 50K LOC# typically produces 2K-10K nodes and 5K-30K edges. If you see very few edges,# the extractors may not have found your code (check language support below).# 4. Get context for a task
knowing context -task "refactor auth middleware" -format gcf
# 5. Check graph integrity
knowing fsck
Verify your setup
After indexing, run these commands to confirm everything is working:
bash
# Show node/edge counts, repos, snapshots
knowing stats
# Search for a symbol you know exists in your code
knowing query "MyKnownFunction"# Check graph integrity (should report 0 errors)
knowing fsck
# If results seem wrong, check if the graph is stale
knowing stale
If knowing stats shows zero nodes or very few edges, see
Troubleshooting below.
More CLI commands
bash
# Find affected tests
knowing test-scope -files internal/auth/middleware.go
# Explain why a symbol ranked where it did
knowing why -task "refactor auth" -symbol "SessionHandler"# Prove a relationship exists (cryptographic Merkle proof)
knowing prove -source"AuthService" -target "SessionStore"# Verify offline (no database needed)
knowing verify proof.json
# Check if the graph is stale (CI gate: exits 1 if stale)
knowing stale
# Supply chain audit (scan all files for suspicious patterns)
knowing audit-supply-chain --scan-all
# Remove a repo (evicts all data: nodes, edges, snapshots, feedback)
knowing remove ./path/to/repo
The --watch flag re-indexes on file changes. Your agent always queries fresh data. No manual knowing index or database path needed: the MCP server auto-indexes the git repository on first launch and registers it in the roster for future sessions.
Embeddings are off by default (confirmed neutral on cold-start benchmarks). Use --embeddings to enable if experimenting. The graph structure and equivalence classes carry retrieval quality.
What your agent gets: The key tool is context_for_task. When your agent calls it with a task description, knowing returns ranked, relevant code symbols packed into a token budget. This replaces grep-read loops. Other useful tools: blast_radius (what breaks if I change this?), test_scope (which tests to run?), explain_symbol (why did this rank here?). See MCP Tools Reference for all 28 tools.
Verify it works:
Start a session with your agent
Ask: "Use the context_for_task tool to find symbols related to [something specific in your code]"
You should see ranked symbols with scores and file paths from your codebase
If results are empty: the repo may still be indexing (10-30 seconds on first launch). If results seem unrelated: use specific symbol names in your task description (e.g., "find the AuthMiddleware handler" not "find auth code"). You can also verify from the CLI:
bash
knowing stats # should show nodes and edges
knowing query "MyFunc"# should find symbols you recognize
Git versions files. knowing versions the understanding of code.
The entire system is built on one idea: content-addressed identity. Every symbol, relationship, and snapshot is SHA-256 hashed. This single choice gives you:
Staleness detection for free. Changed file = new hash = stale edges are known without scanning.
Caching for free. Same package root = same results. 93x speedup on unchanged queries.
Integrity for free. Verify all stored hashes and snapshot chain continuity. 98ms.
History for free. Each snapshot is a Merkle root tied to a git commit. Walk the chain.
Feedback expiration for free. Feedback stores the package Merkle root. Code changes = root changes = old feedback is invisible.
Proofs for free. Merkle path from leaf to root is a self-contained cryptographic proof.
Intelligence: computes blast radius, context packs, test scope, feedback, communities from the stored graph.
The boundary matters: intelligence features read the graph and produce derived results. They cannot corrupt graph facts. A bad ranking produces a bad recommendation; it cannot invalidate a proof.
All extractors fire per file via multi-dispatch; results are merged. Tree-sitter produces edges at confidence 0.7 (ast_inferred); go/packages and SCIP at 0.95-1.0 (ast_resolved, scip_resolved).
Service transport and caching: varint, length-prefixed
74% fewer bytes
JSON
Human debugging, generic consumers
Baseline
GCF uses |-separated fields and local IDs ($1 -> $3) instead of repeated qualified names. Parseable by LLMs while fitting 5x more graph context into the same token budget. Session-stateful deduplication reduces repeated symbols by 47%.
Current Boundaries
Static blast radius follows calls edges; other edge types provide context, not traversal.
Runtime tools require OpenTelemetry trace ingestion; without traces they have no observations.
LSP enrichment: Go, TypeScript, Python, Rust, Java, C#. Auto-detected from project markers. Others fall back to tree-sitter.
Embeddings are off by default (confirmed neutral on cold-start benchmarks, session 23). Use --embeddings to opt in.