Long-term memory for AI agents: semantic facts, episodic events, and procedural workflows
io.github.alibaizhanov/mengram MCP Server
This MCP server provides long-term memory for AI agents, focusing on semantic facts, episodic events, and procedural workflows. The server is identified as io.github.alibaizhanov/mengram and is associated with topics including AI memory, RAG, and knowledge-graph style retrieval.
๐ ๏ธ Key Features
Long-term memory for AI agents
Storage and use of semantic facts
Tracking episodic events
Representation of procedural workflows
Model Context Protocol (MCP) server
๐ Use Cases
AI-agent memory beyond a single session
Semantic search over stored agent knowledge
Retrieval-augmented generation (RAG)
Episodic logging and later recall
Workflow/procedure reuse for agent operations
โก Developer Benefits
Supports agent-memory patterns aligned with cognitive-architecture terminology
Includes toolchain topics relevant to embeddings and vector search (e.g., multilingual embeddings, pgvector)
References alternatives and ecosystems (e.g., Letta alternative, mem0-alternative)
โ ๏ธ Limitations
Available documentation excerpt is incomplete/truncated, so specific tools, configuration, and APIs are not described in the provided data.
pip install mengram-ai # or: npm install mengram-ai
mengram try # see what memory would know about you โ local only,# no account, nothing leaves your machine
python
from mengram import Mengram
m = Mengram(api_key="om-...") # Free key โ mengram.io
m.add([{"role": "user", "content": "I use Python and deploy to Railway"}])
m.search("tech stack") # โ facts
m.ask("what's my tech stack?") # โ synthesized answer + citations
m.episodes(query="deployment") # โ events
m.procedures(query="deploy") # โ workflows that evolve from failures
Native multilingual: ask in Russian, Chinese, Spanish, Japanese โ Mengram retrieves and answers across 23 languages (Cohere multilingual embeddings + rerank).
Install in one prompt (any AI tool)
Paste this into Claude Desktop, Cursor, Codex, Claude Code, or Windsurf โ the agent reads our setup guide, installs the SDK, configures the MCP server, and verifies the round-trip end-to-end. No terminal context-switching.
code
Install Mengram for me. Fetch the canonical install guide at
https://mengram.io/agent-install.txt and follow it precisely.
My email is YOUR_EMAIL_HERE.
Works in any agent with shell + file-edit + web-fetch tools. Prefer doing it manually? See the plain-text guide โ it's structured for human eyes too.
Claude Code โ Memory That Survives /clear AND Auto-Compaction
Persistent memory that survives /clear, auto-compaction, machine switches, and team handoffs โ the SessionStart hook fires after every compact and re-injects your context. The summary can be lossy; the memory isn't.
bash
# 1. Get a free key at https://mengram.io and save it oncemkdir -p ~/.mengram && echo'{"api_key": "om-your-key-here"}' > ~/.mengram/config.json
# 2. Install the plugin (hooks + MCP server + skill)
claude plugin marketplace add alibaizhanov/mengram
claude plugin install mengram@mengram
# 3. Skip the cold start โ import your existing session history# (secrets are redacted on your machine before anything is uploaded)
mengram import claude-code
What happens:
code
Session Start โ Loads your cognitive profile (fires after /clear, compaction, and restarts)
Every Prompt โ Searches past sessions for relevant context (auto-recall)
After Response โ Saves new knowledge in background (auto-save)
Before a Bash โ If the command matches a learned workflow with a weak record, asks you first (policy gate, CLI hooks)
Before compact โ Writes the working state down โ last prompts, files edited, last commands, where Claude left off (CLI hooks)
After compact โ Puts that state back, verbatim, next to the host's summary โ and tells you so in one line
No manual saves. No tool calls. Claude just knows what you worked on yesterday โ even after compaction ate the transcript.
Prefer CLI-managed hooks instead of the plugin? pip install mengram-ai && mengram setup does the same via mengram hook install.
Same memory in Codex and Cursor
Codex and Cursor fire the same lifecycle events, so the same hooks run there โ one memory across all three. mengram setup finds them on your machine and installs their hooks; or add them one by one:
bash
mengram hook install --codex # ~/.codex/hooks.json: session context, recall on every prompt, compaction checkpoint, auto-save + task card
mengram hook install --cursor # ~/.cursor/hooks.json: session context, compaction checkpoint (back on the next tool call), auto-save
Cursor has no per-prompt context hook (beforeSubmitPrompt can only allow or block), so mid-conversation recall there is on request via the MCP tools; session start and compaction are automatic.
Every fact remembers which tool wrote it and when, and recall shows it: deploys to Fly.io from main (codex, 2026-09-15). Each hook also tells the server which OS and tool it runs under, so a workspace path recorded on your Mac never reaches a Claude Code session on a Linux box, and a Cursor settings fact never reaches Codex โ those are left out and counted (facts_left_out_for_host).
Pick a task up where it was left โ mengram resume
When an agent stops, the hooks write a task card for the repository and branch: files touched, last commands, the last test run and the commit it ran on, plus the agent's draft of the task, what is done and what remains. The next session on that branch โ tomorrow, another agent in another worktree of it, another machine โ gets it first, and is told when the code has moved since the last check. A session on a branch with no card of its own (a fresh Orca worktree, say) is only told that cards exist, since another branch's card is usually another agent's task; mengram resume shows it on request:
code
$ mengram resume
[Mengram resume โ where this task stands (2 h ago, from claude-code, branch main)]
Task (agent draft, not confirmed): prepare the Dify example for publication
Done:
- three workflows assembled
- save/recall round-trip checked through the API
Remaining:
- import the workflows into Dify and check the model's answers
Last check: `python3 -m pytest -q tests/test_dify.py` โ 4 passed in 0.31s
The last check ran on 57cdde9; HEAD is now a1b2c3d (4 commits later). Its result describes the earlier state.
Files touched: examples/dify-support/workflow.yml
Sources: session 5d274778 ยท commit 57cdde9
mengram resume --open serves a local page to correct the card, confirm it (confirmed text is never redrafted), pick another task or copy the context for a new session. The card lives in ~/.mengram/resume/ and needs no account; only the draft of task/done/remaining uses a model.
No account? Keep the memory in a folder
bash
pip install mengram-ai
mengram local init ./memory --provider anthropic --api-key sk-ant-... # or openai / ollama
mengram import claude-code --memory ./memory # seed it from your Claude Code sessions
mengram local map --memory ./memory --open # one page: who you are, what happened, what it learned
mengram hook install --memory ./memory # the same four hooks, all local
mengram server --memory ./memory # MCP for Claude Desktop, Cursor, any client
The folder is the memory: a memfmt tree of Markdown you own โ git diffs it, Obsidian draws it, memfmt validate checks it. Same procedures-with-outcomes as the cloud: versions, success/fail counts per step, the policy gate, and the regression gate that quarantines a fix that would break another workflow. Only extraction and a failure revision need a model, and that one you bring. Nothing expires and nothing asks for a key. Docs.
Local search matches words in Unicode text, including Russian; it does not use embeddings or translate queries. Multiple Mengram sessions coordinate writes with a folder lock and merge independent additions. Conflicting edits require a reload instead of overwriting another session. Each changed file is replaced atomically, but the whole folder is not a single crash-atomic transaction, and external editors do not participate in the lock.
Why Mengram?
Every AI memory tool stores facts. Mengram stores 3 types of memory โ and procedures evolve when they fail.
Mengram
claude-mem
Mem0
Zep
Letta
Semantic memory (facts, preferences)
Yes
Yes
Yes
Yes
Yes
Episodic memory (events, decisions)
Yes
Partial
No
No
Partial
Procedural memory (workflows)
Yes
No
No
No
No
Procedures evolve from failures
Yes
No
No
No
No
Cognitive Profile
Yes
No
No
No
No
Native multilingual retrieval (23 languages)
Yes
Partial
No
No
No
Ask & Citations (synthesized answer)
Yes
No
No
No
No
Multi-user isolation
Yes
No
Yes
Yes
No
Knowledge graph
Yes
No
Yes
Yes
Yes
Claude Code hooks (auto-save/recall)
Yes
Yes
No
No
No
MCP server
Yes
Yes
Yes
Yes
Yes
LangChain + CrewAI integrations
Yes
No
Partial
Partial
Partial
Import Claude Code history / ChatGPT / Obsidian
Yes
No
No
No
No
Pricing
Free tier
Free OSS (+cloud backup)
$19-249/mo
Enterprise
Self-host
Get Started in 30 Seconds
1. Install
bash
pip install mengram-ai
2. Setup โ one command does everything: account, Claude Code hooks, MCP configs for detected tools (Cursor, Claude Desktop, Windsurf), history import, and a round-trip check
bash
mengram setup
Or get a key manually at mengram.io and export MENGRAM_API_KEY=om-...
3. Use
python
from mengram import Mengram
m = Mengram(api_key="om-...")
# Add a conversation โ auto-extracts facts, events, and workflows
m.add([
{"role": "user", "content": "Deployed to Railway today. Build passed but forgot migrations โ DB crashed. Fixed by adding a pre-deploy check."},
])
# Search across all 3 memory types at once
results = m.search_all("deployment issues")
# โ {semantic: [...], episodic: [...], procedural: [...]}
File Upload (PDF, DOCX, TXT, MD)
python
# Upload a PDF โ auto-extracts memories using vision AI
result = m.add_file("meeting-notes.pdf")
# โ {"status": "accepted", "job_id": "job-...", "page_count": 12}# Poll for completion
m.job_status(result["job_id"])
javascript
// Node.js โ pass a file pathawait m.addFile('./report.pdf');
// Browser โ pass a File object from <input type="file">await m.addFile(fileInput.files[0]);
bash
# REST API
curl -X POST https://mengram.io/v1/add_file \
-H "Authorization: Bearer om-..." \
-F "file=@meeting-notes.pdf" \
-F "user_id=default"
This happens automatically when you report failures:
python
m.procedure_feedback(proc_id, success=False,
context="OOM error on step 3", failed_at_step=3)
# โ Procedure evolves to v3 with new step added
Every failure-driven revision records which assumption turned out false โ not just which step broke โ and derives a precondition that travels with the procedure at recall time:
json
{"version":3,"violated_assumption":"the build container had enough memory for a full build","preconditions":["check available memory before building"],"success_count":11,"fail_count":2}
An agent loading v3 doesn't repeat the two mistakes that produced it โ and knows what to verify before trusting the workflow.
Or fully automatic โ just add conversations and Mengram detects failures and evolves procedures:
python
m.add([{"role": "user", "content": "Deploy failed again โ OOM on the build step"}])
# โ Episode created โ linked to "Deploy" procedure โ failure detected โ v3 created
Ask Your Memory (RAG built-in)
m.ask() returns a synthesized answer with citations โ not a raw fact list.
Mengram embeds your query, retrieves the top relevant facts, and uses
Cohere Chat to write a grounded answer with native source attribution.
python
result = m.ask("what programming languages do I use?")
print(result["answer"])
# 'You use Python and Rust. Python is your daily language [1] and# Rust is your favorite [2]. You also know Java for enterprise# systems [3].'for cit in result["citations"]:
print(f' "{cit["text"]}" โ {cit["sources"][0]["fact"]}')
# "Python and Rust" โ uses Python daily for backend development# "favorite [2]" โ Rust is favorite language# "Java" โ specializes in Java/Spring Boot
Multilingual: ask in any of 23 languages, get an answer in the same language with citations linking back to facts in the original language they were stored. Premium feature (Pro / Growth / Business).
Cognitive Profile
One API call generates a system prompt from all memories:
python
profile = m.get_profile()
# โ "You are talking to Ali, a developer in Almaty. Uses Python, PostgreSQL,# and Railway. Recently debugged pgvector deployment. Prefers direct# communication and practical next steps."
Insert into any LLM's system prompt for instant personalization.
5 hooks: profile on start, recall on every prompt, save after responses, a policy gate before workflow-shaped Bash commands, and the outcome of each step written back after it runs.
The receipt. Memory that works is invisible, so each hook leaves a line behind and the next session opens with the sum: "Mengram, last session: recalled memories on 4 prompts ยท asked before 1 workflow with a weak record ยท recorded 3 step outcomes (2 ok, 1 failed)". Silent when nothing happened. mengram receipt shows the last session and the past 7 days.
Policy gate. Outcome history changes what the agent may do, not only how results rank. When a git push, deploy, migrate, kubectl, rm -rf โฆ matches a learned workflow that is untested, inherits its record from an earlier version (61% expected), or sits below the bar (58% reliable, default 70), the hook answers ask: you see why, Claude gets the steps on record, nothing runs on the agent's say-so. A proven workflow stays silent. Works against the cloud, or fully offline against a memfmt folder with MENGRAM_MEMORY_DIR=./memory. Only workflow-shaped commands trigger a lookup (one search each); ls and cat never do. Tune with MENGRAM_POLICY_MIN_RELIABLE=80, MENGRAM_POLICY_PATTERN='\bmake\b'; skip with mengram hook install --no-policy. The memory can ask; it never denies.
Mengram works as a persistent memory backend for autonomous agents. Your agent stores what it learns, and recalls it on the next run โ getting smarter over time.
python
from mengram import Mengram
m = Mengram(api_key="om-...")
# Agent completes a task โ store what happened
m.add([
{"role": "user", "content": "Apply to Acme Corp on Greenhouse"},
{"role": "assistant", "content": "Applied successfully. Had to use React Select workaround for dropdowns."},
])
# โ Extracts: fact ("applied to Acme Corp"), episode ("Greenhouse application"),# procedure ("React Select dropdown workaround")# Next run โ agent recalls what worked before
context = m.search_all("Greenhouse application tips")
# โ Returns past procedures, failures, and successful strategies# Report outcome โ procedures evolve
m.procedure_feedback(proc_id, success=False,
context="Dropdown fix stopped working")
# โ Procedure auto-evolves to a new version
Works with any agent framework โ CrewAI, LangChain, AutoGPT, custom loops. The agent just calls add() after actions and search() before decisions.
Self-Hosted (Ollama)
When running locally with Ollama, use models with 8B+ parameters and 8K+ context window. The extraction prompt is ~4,000 tokens โ smaller models will hallucinate or mix examples with real data.
Model
Parameters
Works?
llama3.1:8b
8B
Yes
mistral:7b
7B
Yes
gemma2:9b
9B
Yes
llama3.1:70b
70B
Best
phi4-mini:3.8b
3.8B
No โ context too small
API Reference
Endpoint
Description
POST /v1/add
Add memories (auto-extracts all 3 types)
POST /v1/add_text
Add memories from plain text
POST /v1/add_file
Upload file (PDF, DOCX, TXT, MD) โ vision AI extraction