Two-layer memory MCP server for AI agents with 37 tools, RAG, graphs, wiki, auth
io.github.Cipher208/ariel-memory MCP Server
The io.github.Cipher208/ariel-memory Model Context Protocol (MCP) server provides a two-layer memory system for AI agents. It includes 37 tools for agent memory features such as RAG, a real knowledge graph, a wiki, and authentication. It is implemented with local-first storage using plain SQLite files.
π οΈ Key Features
Two-layer agent memory for AI agents
37 tools
RAG and hybrid search
Real knowledge graph
Wiki support
Authentication
Plain SQLite file storage (local-first)
π Use Cases
Persisting episodic memory for agents so they do not forget
Enabling RAG over locally stored knowledge
Querying and navigating a knowledge graph and wiki content
β‘ Developer Benefits
No cloud and no external APIs (local-first)
FastAPI/asyncio-oriented implementation implied by server topics
Uses SQLite for straightforward local deployment and data access
β οΈ Limitations
Documentation excerpt provided does not specify performance characteristics, supported data formats, or operational constraints beyond βplain SQLite filesβ and βzero cloud/zero external APIs.β
Your AI agents forget. a-memory makes them remember.
4-tier agent memory with hybrid search and a real knowledge graph β all in plain SQLite files. Zero cloud. Zero external APIs.
Also available on PyPI: pip install a-memory β
optional extras: a-memory[embeddings] for real multilingual embeddings.
Why SQLite?
Every other memory server sends your agent's data through a cloud API or requires a separate vector database.
a-memory stores everything in SQLite files on your machine.
Zero infrastructure. No Docker, no database server, no embedding API keys.
Zero data leaving your network. Works air-gapped.
Layer-isolated by design. User facts and agent identity never share a namespace.
One directory = entire memory. Back up with cp, sync with rsync.
Why this exists
Three problems a-memory solves:
β Agent self-evolution β your AI stops repeating mistakes between sessions. It remembers decisions, errors, and corrections in a dedicated agent layer, and an hourly consolidation sweep promotes what matters into long-term facts.
β‘ User persona persistence β your agent knows who it's talking to even after weeks of silence. Preferences, history, emotional context live in the user layer, isolated from agent identity.
β’ Project continuity β project tracks per-project context: decisions with rationale and outcomes, artifact maps, a graphify-powered code index β so a fresh session picks up where the last one left off.
Get started
bash
pip install a-memory
a-memory # MCP server on stdio β connect from any MCP client
git clone https://github.com/Cipher208/a-memory.git
cd a-memory
uv sync
uv run ariel-memory
The five primitives
Agents see exactly six tools β one verb per intent (5 verbs + memory_hook), no tool-choice paralysis:
Primitive
Intent
What it does
think
remember
Routes content to the right layer (L4 facts / L3 episodes / wiki / graph) based on importance, emotion, and relations
dream
recall
Hybrid search across ALL layers (FTS5 + binary embeddings + wiki + graph), returns a token-budgeted digest
forget
let go
Context-aware deletion with Shadow Bin archival (exact / fuzzy / recent)
evolve
grow
Records personality/rules evolution for the agent
project
continue
Per-project identity, decision log, artifact map, code index
Quick demo β Python MCP client:
python
# think β routed to the right store automaticallyawait session.call_tool("think", {"text": "User prefers dark mode", "layer": "user"})
# dream β finds it across every store, a week later
res = await session.call_tool("dream", {"query": "dark mode preference"})
print(res["summary"])
65 fine-grained operations exist in total, grouped into coherent opt-in tiers: the 6 primitives are exposed by default; add context (recall protocol, /new session recap, smart context budget, steering hints, tool-output compression), insight (Memory Query DSL, provenance fact-blame, quality loop, reflections, stats), write (typed memory schemas, declarative rules engine, scratchpad, counterfactuals, episodes), plus wiki, brief, and review (staged mutations) β e.g. ARIEL_EXPOSE=primitives,context,insight,write,wiki,brief,review (57 tools; the remaining 8 are admin-tier, exposed only via ARIEL_EXPOSE=all).
β οΈ Env sanitization gotcha (stdio): MCP clients pass a sanitized environment to stdio servers β setting ARIEL_EXPOSE in your shell profile does nothing. Define the tier set in your MCP client config (the env block of the server entry β see configuration guide). The server logs its resolved surface at startup (tool exposure: N/M tools) β if your agent reports seeing only the primitives, check that line first, then restart the client session (tool lists are cached per session).
Features
Category
What's inside
π§ Memory
L1 Reflex (atomic persistence) β L2 Sessions β L3 Episodic β L4 Core, importance scoring, typed memory kinds with TTL policies, layer isolation; bi-temporal fact history (is_current view hides superseded rows globally, changed_since delta-polling, drill-down to raw source surviving cold archival), hash-chained L0 journal with hot/warm/cold tiers; 65 tools (tiered exposure; 57 on the common combo, 6 primitives by default) including /recall protocol (multi-axis + disclosure triggers), session continuity recap (/new recovery pack), steering hints, tool-output compression + recall verification, provenance fact-blame, Memory Query DSL (faceted tags), typed memory schemas, a declarative rules engine, smart context budget (weighted token floors), reflections, counterfactuals, was_useful quality loop, operator diagnose/heal + integrity score
FTS5-indexed markdown files β edit in Obsidian/VS Code, search from MCP, 6 analytical perspectives (wiki_summarize), schema lint on save, external-dir sync
Architecture
graph TD
A[LLM Agent] -->|MCP Protocol| B[mcp_server]
B --> C{Importance Scoring}
C --> D[L1: ReflexBuffer]
D --> E[L2: SessionStore]
E --> F{EmotionTrigger?}
F -->|high emotion| G[L3: EpisodicMemory]
F -->|normal| H[L4: CoreMemory]
B --> I[RAG Engine]
I --> J[FTS5 Search]
I --> K[MIB Binary Search]
I --> L[Hybrid RRF Ranking]
B --> M[Wiki System]
M --> N[.md Files]
M --> O[SQLite Index]
B --> P[Knowledge Graphs]
P --> Q[Epistemic Graph]
P --> R[Temporal Graph]
B --> S[Project Store]
S --> T[Decisions / Artifacts / Code Index]
U[Hourly Sweep] -->|consolidate| G
U -->|promote| H
U -->|auto-VACUUM| V[(SQLite)]
Comparison
a-memory
mem0
letta (memgpt)
chroma
MCP native
β 6 primitives
β no MCP server
β
β
Layer isolation
β User vs Agent namespaces
β
β
β
Local-only (no cloud)
β SQLite β 0 infra
β οΈ API or self-host Docker
β needs LLM API
β local OSS + Cloud option
Own semantic search (no API)
β FTS5 + MIB binary hybrid
β οΈ BM25+entity (LLM-dependent)
β LLM-only
β οΈ hybrid on Cloud only
Knowledge graph
β Typed nodes + edges + temporal timeline
β οΈ entities only
β
β
Envelope encryption (secrets)
β NaCl SecretBox (auth/saga secrets; memory data is plaintext SQLite)
β
β
β
Lifecycle hooks
β 19 names, per-layer, config-gated
limited
limited
none
Self-maintenance
β Hourly consolidation + auto-VACUUM
β
β
β
Backup / restore
β Auto-cron + saga rollback
β
β
β
Notes (Sep 2026): mem0 now ships a self-hosted Docker image and a managed cloud with hybrid BM25+entity search; chroma is 29kβ and added hybrid+FTS5 to its Cloud tier (OSS server remains vector-only). What still differentiates a-memory: zero-infra SQLite (no Docker), NaCl-encrypted auth/saga secrets, layer isolation, hourly self-maintenance, and the temporal graph timeline.
Roadmap
4-layer memory hierarchy with layer isolation
Hybrid search (FTS5 + MIB binary embeddings)
Knowledge graphs (epistemic + temporal)
Hourly consolidation sweep + DB self-maintenance
mcp 2.x native SDK
Repo renamed to Cipher208/a-memory; PyPI package live (pip install a-memory)
Temporal timeline wired end to end (think/evolve/project events + dream recent digest)
S20 eval β β11 ablation on MINI + LongMemEval-S (50-question stride): full dual-route arm wins on both datasets, published-baseline comparison (GPT-4o long-context league, above ChatGPT-memory); dense e5-small run 3: no parity gain over hash on this split β hash stays prod
Screenshot / asciinema demo in README
LLM-assisted consolidation on top of the deterministic sweep
Stage 2 β tool-surface redesign (slot system, URI keys, inject/key consolidation β planned with the A-remainder)
LongMemEval full 500-question split + dense-aware threshold retraining (e5-small parity shown, not a win on the 50-question stride)