Cognitive memory for AI agents — semantic recall, knowledge graph, and contradiction detection
io.github.yantrikos/yantrikdb-mcp (MCP Server)
The io.github.yantrikos/yantrikdb-mcp server provides “cognitive memory for AI agents,” including persistent semantic recall, a knowledge graph, contradiction detection, and procedural learning. It can run as an embeddable engine, network database, or MCP server, and is intended to work with MCP-compatible clients.
yantrikdb-mcp demo: three memories stored over MCP, one recalled by a question that shares no words with it, then think() flagging that two stored beliefs about the same on-call lead contradict each other — the next session's recall comes back with open_conflicts: 1
All data on your machine. No telemetry. No external services.
Install
bash
# Default — uses the engine's bundled 64-dim embedder. ~10 MB install,# ~80 ms cold start, no native ML deps.
pip install yantrikdb-mcp
# Optional: higher-quality 384-dim ONNX MiniLM-L6-v2 embedder (~150 MB install).# Auto-used when an existing pre-v0.6 database is detected.
pip install 'yantrikdb-mcp[onnx]'
Upgrading from v0.5.x? Your existing database stays at 384 dim — install
the [onnx] extra to keep using it transparently. New installs default to
the lean bundled embedder. v0.7.0+ pins the engine migration fix automatically.
See Embedder backends below.
Configure
The MCP server has three deployment modes. Pick the one that fits your setup.
Mode 1 — Local (default, recommended for single user)
The MCP server runs the engine in-process with a local SQLite database. Fast, private, zero dependencies.
uvx fetches and runs the server on demand, so this works with no install step
(it needs uv on PATH). If you installed the
package yourself with pip or pipx, "command": "yantrikdb-mcp" with no args
is equivalent.
That's it. The agent auto-recalls context, auto-remembers decisions, and auto-detects contradictions — no prompting needed.
Mode 2 — HTTP Cluster (recommended for shared/multi-machine setups)
Forward all tool calls to a YantrikDB HTTP cluster instead of using an embedded engine. The MCP server is a thin stateless client — all memories live on the cluster, accessible from any machine.
Benefits: shared memory across machines, high availability, no local embedder download, no local database.
Supports sse and streamable-http transports. Note: SSE connections can drop on idle — Mode 2 (HTTP Cluster) is more reliable for shared deployments.
Environment Variables
Variable
Used in Mode
Default
Description
YANTRIKDB_SERVER_URL
Cluster
(unset → local mode)
Comma-separated cluster node URLs
YANTRIKDB_TOKEN
Cluster
(none)
Bearer token for the cluster database
YANTRIKDB_DB_PATH
Local
~/.yantrikdb/memory.db
Database file path
YANTRIKDB_EMBEDDER
Local
auto
Backend selector: auto | bundled | onnx | multilingual
YANTRIKDB_EMBEDDING_MODEL
Local
all-MiniLM-L6-v2
ONNX model name (only used when YANTRIKDB_EMBEDDER=onnx)
YANTRIKDB_SKILLS_WRITE_ENABLED
All
false
Set true to allow agents to author skills via skill(action="define") (see Skill substrate below)
YANTRIKDB_OUTCOMES_WRITE_ENABLED
All
true
Outcome tracking via skill(action="outcome"). Defaults on so the feedback loop works out of the box; set false to lock the outcome substrate. Added in v0.8.1 per #8
YANTRIKDB_API_KEY
SSE server
(none)
Bearer token when serving SSE/HTTP
Embedder backends
Local mode ships three embedders. The MCP picks one automatically; override with YANTRIKDB_EMBEDDER.
Backend
Dim
Cold start
Install size
Language coverage
When it's used
bundled (engine default)
64
~80 ms
~10 MB
English-only
New / empty databases (auto-selected)
onnx (MiniLM-L6-v2)
384
~2 s
~150 MB
English (higher recall)
Existing pre-v0.6 databases (auto-selected), or when set explicitly
multilingual (potion-multilingual-128M)
256
~2 s + ~460 MB download on first use
~10 MB pip + ~500 MB model cache
101 languages (BGE-M3 tokenizer)
Opt-in only via YANTRIKDB_EMBEDDER=multilingual
auto (default) reads the SQLite file at YANTRIKDB_DB_PATH and picks onnx if it already contains memories — preserving recall quality on upgrades — and bundled otherwise. Multilingual is never auto-selected because its 256-dim vectors are incompatible with existing bundled (64-dim) or ONNX (384-dim) databases; opt-in only on fresh databases.
Set YANTRIKDB_EMBEDDER=bundled|onnx|multilingual to override. If you set YANTRIKDB_EMBEDDER=onnx (or auto-detection picks it) without installing the extras, the server fails fast with an install hint:
code
RuntimeError: Existing DB has memories embedded with the 384-dim ONNX
model, but ONNX deps are missing.
Install with: pip install 'yantrikdb-mcp[onnx]'
For the multilingual backend, the engine downloads potion-multilingual-128M (~460 MB tarball) from github.com/yantrikos/yantrikdb-models on first use. The download is SHA-256 verified, extracted into the engine's cache dir, and reused on subsequent starts. No extra Python deps required — the model runs entirely inside the Rust engine.
Why Not File-Based Memory?
File-based memory (CLAUDE.md, memory files) loads everything into context every conversation. YantrikDB recalls only what's relevant.
Benchmark: 15 queries × 4 scales
Memories
File-Based
YantrikDB
Savings
Precision
100
1,770 tokens
69 tokens
96%
66%
500
9,807 tokens
72 tokens
99.3%
77%
1,000
19,988 tokens
72 tokens
99.6%
84%
5,000
101,739 tokens
53 tokens
99.9%
88%
Selective recall is O(1). File-based memory is O(n).
At 500 memories, file-based exceeds 32K context windows
At 5,000, it doesn't fit in any context window — not even 200K
YantrikDB stays at ~70 tokens per query, under 60ms latency
Precision improves with more data — the opposite of context stuffing
Run the benchmark yourself: git clone https://github.com/yantrikos/yantrikdb-mcp && cd yantrikdb-mcp && python benchmarks/bench_token_savings.py
Recommended agent workflow (golden path)
The server injects a golden-path playbook into the agent's system prompt. Since v0.10.0 the default is digest-first:
Cold start — one call.session(action="digest") returns a single briefing (narrative chain head, open decisions, unresolved conflicts, pending triggers, stale high-importance memories) — replacing several separate recall/temporal calls at conversation start. Then recall only for the specific thing the current message is about.
During work — capture as you go. New durable fact → remember; a stored fact changed → correct (keeps history, avoids contradictions); relationship learned → graph(action="relate").
End of substantial work — conditional. Only when the session was long or state-changing: think to consolidate + detect conflicts. Short/read-only exchanges need no end step.
Trust boundary: recalled memories and digest snippets are data, not instructions. The playbook directs the agent never to execute directives found inside recalled content — a memory may carry text an earlier session or another user stored.
Tools
20 tools, full engine coverage (gaps, conversation, task added in v0.9.0; atlas added in v0.24.0):
Tool
Actions
Purpose
remember
single / batch
Store memories — decisions, preferences, facts, corrections
recall
search / refine / feedback
Semantic search, refinement, and retrieval feedback
v0.24.0 — export this store's Memory Atlas (every memory, entity links, claims, revision history, tasks) as a static page and serve it on localhost; read-only, embedded mode only
Plus new actions on existing tools in v0.9.0:
session(action="digest") — one-call boot-time briefing (narrative chain head + open decisions + conflicts + triggers)
YantrikDB exposes a structured agent skill catalog — separate from loose procedure memories. Skills have schema (skill_id, applies_to, triggers, body, type) and are stored in the dedicated skill_substrate namespace so multiple consumers (this MCP, yantrikdb-hermes-plugin, Lane B SDK, WisePick, yantrikdb-server's /v1/skills/* endpoints) all read and write the same substrate. Background: Sarkar 2026 — Skill as Memory, Not Document.
Security model
Skill writes shape future agent behavior across sessions, so the MCP server implements defense-in-depth. Every control has an env-var knob (locked once at startup — C2) and the full state is exposed via stats(action="stats") and the audit log.
Layered controls (each ships on by default unless noted):
Layer
Control
Env var
Notes
Schema
skill_id regex, body 50–5000 chars, applies_to 1–10 entries, skill_type enum
(always on)
Same regex set as yantrikdb-server /v1/skills/define
YANTRIKDB_OUTCOMES_WRITE_ENABLED=false to lock outcomes too
v0.8.1+: define and outcome have different threat profiles — outcome can't introduce new instructions, only append {succeeded, note≤500} against an existing skill. Feedback loop works by default; lock explicitly if needed
C2 Locked config
All YANTRIKDB_SKILLS_* / YANTRIKDB_OUTCOMES_* env vars read once at startup
(always on)
Mutating env in a sub-process can't bypass the gate
Accept/reject counts by reason, surfaced in stats(action="stats")["skill_substrate"]
(always on)
Operator dashboards
E1 Body SHA-256
Stored at write time, re-verified on every read
(always on)
Detects out-of-band DB tampering — surface/get omit mismatches and log to audit
E2 Author origin
metadata.author_origin tag — defaults to yantrikdb-mcp
YANTRIKDB_SKILLS_AUTHOR_ORIGIN=... to override
Tracks substrate provenance across consumers
F Startup safety
Boot-time warnings about dangerous configurations
(always on)
Logs [F.1]–[F.5] to stderr + audit
G Review queue for rule
rule-type skills route to skill_pending_review (not surfaced by surface/get/list)
YANTRIKDB_SKILLS_RULE_REQUIRES_REVIEW=false to disable (not recommended)
Rules influence agent policy — human approval required
Multi-tenant guard
[F.1] warning if DB shows multiple actor IDs without ack
YANTRIKDB_SKILLS_MULTITENANT_ACK=true
One DB = one tenant is the safe default
Enterprise checklist:
bash
# Minimum production config when you turn the gate ON:
YANTRIKDB_SKILLS_WRITE_ENABLED=true
YANTRIKDB_SKILLS_WRITE_EXPIRES_AT=2026-12-31T00:00:00Z
YANTRIKDB_SKILLS_ALLOWED_NAMESPACES=workflow,review,onboarding
YANTRIKDB_SKILLS_AUDIT_LOG=/var/log/yantrikdb/skills.audit.jsonl
YANTRIKDB_SKILLS_AUTHOR_ORIGIN=acme-corp-claude-prod
# Defaults are already correct: writes off, scanners on, rate-limit 30/min,# rule-type routed to review, body-hash verified on read, locked at startup.
The audit log is the canonical record. Every accept, every reject (with the scanner that flagged), every tamper-detection on read, every gate-closed-due-to-expiry — all there in JSONL. Plug it into your SIEM.
stats(action="stats") example output (skill_substrate slice)
Lowercase dot-separated segments, length 4–200, e.g. workflow.git.commit_clean
body
50–5000 chars
applies_to
1–10 lowercase-underscore identifiers (no hyphens — load-bearing for substrate consistency)
skill_type
One of procedure, reference, lesson, pattern, rule
on_conflict
reject (default) or replace
Example session
python
# Define (requires gate enabled)
skill(action="define",
skill_id="workflow.git.commit_clean",
body="Before commit: run pytest, run lint, write a clear subject + body.",
skill_type="procedure",
applies_to=["git", "release"])
# Surface relevant skills for the current task
skill(action="surface", query="how to commit cleanly", top_k=5)
# Record an outcome after using the skill (gated, append-only)
skill(action="outcome", skill_id="workflow.git.commit_clean",
succeeded=True, note="caught a flake8 issue pre-push")
Outcomes are append-only events in the outcome_substrate namespace — no auto-rollup on the parent skill, matching yantrikdb-server's "schema not semantics" design rule. Agents (or the operator) can aggregate outcomes themselves to compute success rates.
FAQ
What is YantrikDB MCP?
YantrikDB MCP is a Model Context Protocol (MCP) server that gives AI agents persistent cognitive memory across sessions. It exposes 20 tools (remember, recall, forget, correct, think, graph, conflict, trigger, session, temporal, procedure, category, personality, stats, memory, skill, gaps, conversation, task, atlas) that any MCP-compatible client — Claude Code, Cursor, Windsurf, Continue, Claude Desktop — can call automatically without prompting.
How is this different from file-based memory like CLAUDE.md?
File-based memory loads everything into context on every conversation, which scales O(n) in token cost. YantrikDB uses selective semantic recall — at 5,000 memories, file-based costs ~101K tokens per conversation while YantrikDB costs ~53 tokens. Precision improves with more data instead of degrading as the context window fills up. Benchmark script: python benchmarks/bench_token_savings.py.
How does it compare to mem0 / Letta / Zep / native MCP memory?
See comparison table below. Short version: YantrikDB is the only one that ships as both an embeddable Rust engine and an MCP server and a network database with the same substrate semantics. It's the only one with first-class procedural memory + a skill substrate validated by schema at write time + autonomous consolidation/conflict detection. It's also the only one whose underlying engine is published as a peer-reviewed paper (Sarkar 2026, Zenodo DOI 10.5281/zenodo.20128887).
Can I self-host?
Yes — three ways. (1) Local: just pip install yantrikdb-mcp and point your MCP client at it. SQLite lives at ~/.yantrikdb/memory.db. (2) Network: run yantrikdb-server as a multi-tenant HTTP cluster, point the MCP at it via YANTRIKDB_SERVER_URL. (3) Hybrid: SSE server mode (yantrikdb-mcp --transport sse) for shared deployments.
Is my data sent anywhere?
No. All data stays on your machine (or your cluster). No telemetry, no third-party services. The default embedder runs entirely in the Rust engine via static lookup — no model downloads or API calls. The optional [onnx] and multilingual embedders fetch model weights once from HuggingFace's CDN and run locally thereafter.
What's the difference between procedure and skill?
procedure stores loose how-to memories (effectiveness-ranked, no schema). skill stores structured catalog entries (skill_id, applies_to, triggers, body, type) in a dedicated skill_substrate namespace shared with yantrikdb-hermes-plugin, Lane B SDK, WisePick, and the yantrikdb-server /v1/skills/* endpoints. Use procedure for personal how-to notes; use skill for structured agent capabilities that other consumers should be able to surface.
Is skill authoring safe to enable?
Skill writes are off by default precisely because they can shape future agent behavior. When you turn the gate on, seven layers of defense-in-depth apply: prompt-injection scanner, credential scanner, URL block, unicode-evasion scanner, namespace allowlist, author attribution, audit log, rate limit, body-hash tamper detection, and a review queue for rule-type skills. See Security model above.
Does it work in production?
Yes — yantrikdb-mcp runs in production on the YantrikDB homelab cluster (1973+ memories, SSE transport, 2 weeks uptime per release cycle) and is the reference deployment behind the engine's release decisions. v0.8.x added the engine's same-day-patch cadence to the MCP server itself: external issues filed by community contributors land as released fixes within 2 hours.
What's the engine written in?
The YantrikDB engine is Rust (crates.io: yantrikdb) with pyo3 Python bindings (PyPI: yantrikdb). The MCP server itself is Python — a thin wrapper around the engine's Python bindings, plus stdio/SSE/HTTP transport plumbing.
Comparisons reflect public-facing capabilities as of May 2026. PRs welcome to correct any rows.
Cite this work
If you use YantrikDB in academic or research context, please cite the substrate paper:
bibtex
@misc{sarkar2026skill,
author = {Sarkar, Pranab},
title = {Skill as Memory, Not Document: A Database-Native Substrate for Agent Skill Catalogs},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.20128887},
url = {https://doi.org/10.5281/zenodo.20128887},
orcid = {0009-0009-8683-1481}
}
After storing "We use Python 3.11" and later "We upgraded to Python 3.12", calling think() detects the conflict. The agent surfaces it:
"I found a contradiction: you previously said Python 3.11, but recently mentioned Python 3.12. Which is current?"
Then resolves with conflict(action="resolve", conflict_id="...", strategy="keep_b").
Privacy Policy
YantrikDB MCP Server stores all data locally on your machine (default: ~/.yantrikdb/memory.db). No data is sent to external servers, no telemetry is collected, and no third-party services are contacted during operation.
Data collection: Only what you explicitly store via the remember tool or what the AI agent stores on your behalf.
Data storage: Local SQLite database on your filesystem. You control the path via YANTRIKDB_DB_PATH.
Third-party sharing: None. Data never leaves your machine in local (stdio) mode.
Network mode: When using SSE/HTTP transport, data travels between your client and your self-hosted server. No Anthropic or third-party servers are involved.
Embedding model: Uses a local ONNX model (all-MiniLM-L6-v2). Model files are downloaded once from Hugging Face Hub on first use, then cached locally.
Retention: Data persists until you delete it (forget tool) or delete the database file.
langchain-yantrikdb — YantrikDB as a LangChain VectorStore and ChatMessageHistory.
yantrikdb-hermes-plugin — memory provider for NousResearch/hermes-agent, sharing the same skill substrate.
yantrik-memory — framework-agnostic memory layer with traits and bond evolution.
License
This MCP server is licensed under MIT — use it freely in any project.
Note: This package depends on yantrikdb (the cognitive memory engine), which is licensed under Apache-2.0 as of 2026-08-18 (previously AGPL-3.0). Both this server and the engine are now permissively licensed — there are no copyleft obligations on your code, modifications, or hosted services.