Pseudolife-MCP

简体中文 ·
日本語 ·
한국어 ·
Português (BR) ·
Español
Persistent long-term memory for Claude Code, Codex, and other MCP clients.
An MCP server that gives coding agents a long-term memory that persists across
sessions — surviving context compactions and fresh tasks. Your coding agent is
the intelligence; this server is its memory on disk.

What you get:
- Associative memory with honest forgetting — a flat similarity store
ranked by hybrid dense-plus-lexical retrieval, with conflict detection
that admits potential updates while preserving earlier source notes;
whole-note replacement is explicit. (The measured verdict: a preregistered
ablation campaign found the previous 8-band continuum tied a flat store
on every gate, so the simpler structure ships; the continuum remains
one config line away.)
- Canonical facts, not vibes — one current value per
entity.attribute
slot (or a member set, for slots that hold many concurrent values);
corrections supersede rather than silently overwrite, and the full
version history survives.
- Dreams — a bundled local extractor, or any OpenAI-compatible endpoint
(a Claude model on your Max plan, a GPT-5.6 model on a ChatGPT plan, LM
Studio, Ollama, vLLM), consolidates the memory stream into facts and a
knowledge graph while you're not looking.
- Lessons from its own work — successes, dead-ends, and your corrections
become do/avoid guidance surfaced at the start of every session.
- A web console to watch it think — the Cortex Console above, plus cited
world facts, session episodes, and document RAG.
Measured, with receipts — the full 500-question LongMemEval sweep, all
six question types, and every number ships with its committed run artifact:
| LongMemEval oracle, 500 questions | naive RAG | commit-gated cascade |
|---|
| accuracy, all six question types | 0.688 | 0.690 |
| context tokens per question | ~1210 | ~883 |
| knowledge-update slice (78 of the 500) | 0.859 | 0.936 (retired — see below) |
Equal accuracy to naive RAG across the whole benchmark on ~73% of the
context, and better calibrated about what it does not know: on BEAM-100K's
abstention questions the fact spine scores 0.950 against naive RAG's
0.775, unchanged under two independent judges. Read that as calibration,
not recall — in the budget-matched five-arm run of 2026-09-02 (rag 0.725 there;
one replicate, local judge) an arm served no memory at all scores 1.000 on
the same questions, because refusing is the right answer there and an
empty context always refuses. The fact spine loses where an answer has to
be aggregated across sessions. The second claim to survive a judge swap is
a win rather than a wash: re-run on 2026-09-04 with the hybrid arm
budget-matched to the control at 6 turns, the same 500 questions give
hybrid 0.730 against naive RAG's 0.690 under the local judge and
0.736 against 0.694 under claude-opus-5 — paired +0.040 / +0.042,
p 0.015 / 0.013 — bought with more context, ~1229 tokens against the
control's ~1124, not less, and carried mostly by temporal-reasoning
questions. Graded by a local, byte-reproducible judge (the cross-judge check
names its second judge) — compare within rows, never against GPT-judged
leaderboards.
Retired 2026-08-25 (#188): the 0.936 knowledge-update headline. It was
measured on the 2026-07-30 bench stack (Qwen3.6-27B answerer and judge).
Re-running the same 78 questions after the 2026-08-17 migration to
Qwen3.8-27B puts the cascade at 0.846, below the naive-RAG control —
which lands on 0.859 on both stacks. The cascade serves the fact-spine
answer unless that channel says "I don't know", so it measures the
answerer's abstention behaviour as much as the memory: 32/78 abstentions
at 46/46 commit precision on the old stack, 22/78 at 0.839 on the new one.
The 500-question table above is on the older judge and has not been
re-judged, so read its cascade row as an upper bound.
Full tables, the per-type breakdown, both stacks side by side, and every
artifact: Benchmarks.
Quickstart
Install and register the lite tier. No Docker, no database to set up, no container runtime:
pip install "pseudolife-mcp[lite]"
claude mcp add --scope user pseudolife-memory -- pseudolife-mcp
Codex instead of Claude Code — same shape:
pip install "pseudolife-mcp[lite]"
codex mcp add pseudolife-memory --env PSEUDOLIFE_WRITER_ID=codex -- pseudolife-mcp
For Codex, finish setup before starting a fresh task. In the existing
[mcp_servers.pseudolife-memory] table in ~/.codex/config.toml, add
startup_timeout_sec = 240, tool_timeout_sec = 240, and required = true.
The shim can wait up to 180 seconds for a cold daemon; Codex's default
startup budget is 10 seconds. required makes missing memory visible at
startup and waits for its initial catalog. These are starting budgets,
not a promise that a first model download fits. The tool budget leaves time for
the shim's 180-second deadline to report a failure before the host cancels it; prewarm with
pseudolife-mcp serve in a terminal if needed.
Approve memory_message in that table's tool configuration
([mcp_servers.pseudolife-memory.tools.memory_message] approval_mode = "approve", or allow it once and keep the approval). Board mail can wake an
idle Codex task by default, and the woken task reads its mail with that tool:
without the approval it stalls on a prompt until someone answers it. Set
PSEUDOLIFE_CODEX_DOORBELL = "0" in the same env table to keep the task
from being woken (see Codex doorbell).
The agent board (peer awareness and addressed mail between sessions) is on by
default, behind bearer authentication. Without coordination.allowed_principals
in the daemon's config.yaml, only the singular PSEUDOLIFE_MCP_TOKEN
principal is admitted. A Codex bearer from a PSEUDOLIFE_MCP_TOKENS map stays
off the board until an operator lists it (allowed_principals: [default, codex]).
Until then a shim on the default setting leaves coordination off without an
error; one pinned with --enable shows an attach-unavailable hint instead.
python ops/setup-codex-coordination.py --check reports ready (default-on)
or names the cause. See
Codex CLI and desktop.
The MCP handshake delivers compact recall/capture/reflection instructions.
For the complete standing guidance, copy the
bundled memory block into your project
AGENTS.md or ~/.codex/AGENTS.md. For session briefings and per-turn
reminders, follow Codex hooks and verification.
Use one MCP registration and one hook source; an installed plugin may
already provide either. After the daemon is running, execute
pseudolife-mcp doctor from the same environment as the registered command.
It checks the handshake and annotations without calling bank tools.
Then in either coding agent: "remember that my staging box is haze-02" →
the agent calls memory_store; next session, "which box is staging?" →
memory_search finds it. Browse everything at the Cortex Console:
http://127.0.0.1:8765/ui/.
The rebuilt console (phase 1: Observatory and the agent Board) is at
http://127.0.0.1:8765/ui/next/; the rest of its views still open the classic one.
The first session auto-starts the daemon, which provisions an embedded
PostgreSQL 18 (pgvector included, via pg0-embedded) under a stable
per-user data dir and downloads the embedding model (~1.2 GB, one-time).
It is a real Postgres bank, not a cut-down one: pseudolife-mcp backup
writes a standard owner-free pg_dump archive (plus a state archive, 7-day
rotation) that restores into any PostgreSQL 18 target regardless of role —
the Docker tier included — so outgrowing lite is a dump/restore, not a
migration project (backups). For a
tier- and Postgres-version-independent copy, pseudolife-mcp export /
import move the whole bank as portable JSONL
(logical export / import).
Windows needs an ASCII-only data path
(PSEUDOLIFE_MCP_DATA_DIR).
What lite gives you, and the one thing it doesn't
| lite (pip) | durable (Docker) |
|---|
| Associative store, hybrid search, supersession, version history | yes | yes |
| Cortex facts, knowledge graph, lessons, world facts, episodes | yes | yes |
Cortex Console, document RAG, pseudolife-mcp backup | yes | yes |
| Dream consolidation filling the cortex on its own | no extractor ships | yes — bundled local CPU sidecar |
| External volumes, health-checked services, deploy/rollback tooling | no | yes |
The gap, stated plainly. Lite ships no extractor, so the dream
pass still runs, prunes, and acknowledges its input batch, but writes no
canonical facts: on this path memory_fact_set is the only cortex
writer. Everything else above works. Nothing about this is silent —
curl http://127.0.0.1:8765/health reports "extractor": "none", and the
stdio shim says the same on stderr at session start.
Any OpenAI-compatible endpoint closes it. The daemon inherits the
environment it starts from, so two variables are the whole fix — with a
local Ollama:
export PSEUDOLIFE_DREAM_BASE_URL=http://localhost:11434/v1
export PSEUDOLIFE_DREAM_MODEL=qwen2.5:7b
pseudolife-mcp serve
$env:PSEUDOLIFE_DREAM_BASE_URL = "http://localhost:11434/v1"
$env:PSEUDOLIFE_DREAM_MODEL = "qwen2.5:7b"
pseudolife-mcp serve
/health then reports "extractor": "configured". One gotcha: a daemon
that is already running keeps the environment it started with, and the shim
reattaches to it rather than spawning a new one — stop the old daemon
first. A hosted endpoint works too, and costs you the zero-egress
property: memory text leaves the machine. Extractor tiers, quality, and the
trade-offs: Dreaming.
Durable tier — Docker (recommended for a long-lived bank)
Everything above plus the bundled extractor, external volumes,
health-checked services, and backup/rollback tooling. Requires Docker and
at least one MCP-capable coding agent — Claude Code, Codex, and Gemini CLI
are wired end-to-end; anything else gets paste-ready config
(provider matrix). One command from clone to
first memory:
git clone https://github.com/Pseudogiant-xr/Pseudolife-MCP.git
cd Pseudolife-MCP
ops/install.sh
ops\install.ps1
The installer first asks where the memory bank lives: on this machine
(the default, and everything below), on another machine that already runs
it (a client-only install that asks for the daemon's URL and token), or on
this machine with other machines connecting to it (see
Sharing one bank across machines). It then
asks which agents to wire (multi-select, with a capability
matrix showing exactly what each one gets — session briefing, per-turn
discipline, standing file), runs the preflight (one exact fix line per
missing prerequisite), then asks which dream extractor should
consolidate memories —
- sidecar — the bundled local CPU model; no Claude plan needed, works
for everyone, and keeps every memory on the box (~11.8 GB image);
- claude-only — the lightest install: a Claude model via a CLI shim
(
claude-opus-5-5 by default since 2026-09-29, when it cleared the
extraction-ladder gate against claude-opus-5; needs a logged-in
Max-plan claude CLI);
the sidecar image is never built or pulled (~11.8 GB lighter; dreams
pause while the shim is down);
- claude-fallback — the Claude shim primary, the bundled sidecar as
automatic fallback (Max-plan CLI plus the ~11.8 GB image);
- openai-only / openai-fallback — the same two shapes on an OpenAI
subscription: a GPT-5.6 or GPT-6 model (Terra by default) via the Codex CLI
shim on a signed-in ChatGPT plan (extraction quality unmeasured — see
the dreaming guide);
- endpoint / endpoint-fallback — any OpenAI-compatible server you name
(
--extractor-url, --model: LM Studio, Ollama, vLLM, a hosted API),
alone or with the sidecar as fallback. The earlier names sonnet-* and
codex-* still work as deprecated spellings —
then brings the stack up, installs the selected clients' session hooks
(where the client has a hook system), registers the MCP transport (the
stdio shim by default, with a per-provider writer id; direct HTTP via
--transport http), and health-checks the daemon — finishing with a
per-agent ladder of what got wired and what that agent's platform cannot
support. Codex setup offers one choice to enable automatic memory briefings,
reminders, and session cleanup (plus, where the agent board is on, a board
check-in at session start and a new-mail hint per prompt), use standing
instructions only, or skip.
Automatic setup reuses an enabled PseudoLife plugin or installs the three
lifecycle hooks, backs up configuration, approves only their exact current
definitions, and verifies execution. If verification fails, the same approval
allows the standing memory block as a fallback; setup reports the remaining
repair step. Hook-less providers (Gemini CLI and generic agents) are offered
the standing block. --instructions append always writes the block from examples/CLAUDE.memory.md into
~/.claude/CLAUDE.md / ~/.codex/AGENTS.md / ~/.gemini/GEMINI.md
(useful for subagent visibility even with hooks).
Idempotent — re-run any time; --extractor <mode> switches extractor
setups. Non-interactive example:
ops/install.sh --extractor sidecar --client codex --codex-hook-trust yes.
Without explicit hook approval, unattended setup does not grant trust; use
--codex-hooks skip --instructions append for instructions only. Explicit
--instructions skip prevents fallback edits.
The installer also turns the agent board on: it mints a bearer token in
ops/.env (owner-only, never printed) and gives each shim client it wires an
owner-only token file. Re-running it on an existing install adds that file to
the Claude Code registration in place. The ladder's last line, and
pseudolife-mcp doctor, say whether the board is on or why it is off.
Use doctor --host codex or --host claude-code for a compact coordination
snapshot: reachability, bearer admission, registration evidence, tool inventory
and configured wake are separate. Default doctor is read-only; health and a
configured wake path do not prove delivery. The coordination diagnostic
quickstart
covers saved-instance checks and an explicitly disposable mail proof.
--no-token / -NoToken keeps an open-loopback install with the board
dormant (Turning the board on).
Linux (Docker Engine): your user must be in the docker group —
sudo usermod -aG docker $USER, then log out/in (the preflight checks this).
Image sizes, the Windows WSL2 memory cap, and what the installer automates:
the containerized install below.
Manual install (the steps the installer automates)
ops/preflight.sh --client codex
docker volume create pseudolife-mcp-bank
docker volume create pseudolife-mcp-state
docker compose -f ops/docker-compose.yml up -d --build
docker compose -f ops/docker-compose.yml -f ops/docker-compose.ghcr.yml pull pseudolife-pg pseudolife-daemon
docker compose -f ops/docker-compose.yml -f ops/docker-compose.ghcr.yml up -d
curl http://127.0.0.1:8765/health
pip install pseudolife-mcp
claude mcp add --scope user pseudolife-memory --env PSEUDOLIFE_MCP_NO_SPAWN=1 -- pseudolife-mcp
codex mcp add pseudolife-memory --env PSEUDOLIFE_MCP_NO_SPAWN=1 -- pseudolife-mcp
claude mcp add --transport http --scope user pseudolife-memory http://127.0.0.1:8765/mcp
codex mcp add pseudolife-memory --url http://127.0.0.1:8765/mcp
cat examples/CLAUDE.memory.md >> ~/.claude/CLAUDE.md
cat examples/CLAUDE.memory.md >> ~/.codex/AGENTS.md
Optional knobs live in ops/.env (cp ops/.env.example ops/.env — the
install/update scripts scaffold it too; every value is commented, a missing
file runs entirely on defaults).
What this is
A memory engine exposed over MCP. There's no chat UI and no LLM doing the
thinking — your coding agent is the intelligence; these are tools it calls to store and
recall what matters. (Models are bundled as plumbing: baked embedding
weights for retrieval, and the optional CPU extractor sidecar that
consolidates memories into facts while you sleep.)
Where it sits among the common approaches to agent memory — each column
is a fair tool for what it's for; this table is about what question each
one answers, not who's wrong:
| notes file (CLAUDE.md) | auto-journaling plugin | plain vector store | Pseudolife-MCP |
|---|
| Survives sessions and compactions | yes | yes | yes | yes |
| "What is X now?" has one current answer | if you curate it | no — replays what happened | no — every stored version competes at recall | yes — slot-keyed cortex |
| A canonical-fact correction replaces the old value | you edit the file | appended beside it | old and new both retrievable, unranked by recency of truth | cortex supersedes, with full version history kept |
| Facts know their age and go stale | no | no | no | dated, freshness-decayed, quarantined when stale |
| Distils do/avoid lessons from its own outcomes | no | no | no | yes |
| Benchmark numbers ship with their raw run artifacts | — | typically no | typically no | every published number, test-enforced |
Auto-journaling records what the agent did; Pseudolife curates what it
learned. Both are useful — they answer different questions. Named
alternatives — Mem0, Zep/Graphiti, Letta, Cognee, memU, Memori — and the
cases where one of them is the better pick:
Comparison.
It layers several complementary stores: the associative store (a flat
embedding store ranked by cosine similarity fused with a BM25 lexical pool
(on by default), with conflict-aware admission and explicit source-note
replacement; an 8-tier
banded layout is available as an opt-in preset); the cortex (slot-keyed canonical facts — one current
value per entity.attribute, or a member set for set-valued slots — with
provenance tiers and contender parking instead of silent overwrites); a typed knowledge graph over those facts
with a closed relation vocabulary and on-read inference; the world
cortex (durable cited facts about external reality, age-decayed trust);
procedural lessons learned from the agent's own work; and a ChromaDB
reference bank for document RAG. The canonical layers in depth:
the memory model; the graph and multi-hop
recall: retrieval.
State lives in Postgres (the durable source of truth) behind a single
long-lived daemon; every session attaches through a thin stdio shim
(installer default — per-session identity) or directly over HTTP
(single-session setups). The result: Claude can pick up where it left
off, correct itself when facts change, and reason over relationships —
without you re-explaining context each session.
Documentation
This README is the front door — install, wiring, and the basic loop. The
deep material lives in the user guide:
| Page | What's in it |
|---|
| Configuration | Env vars, tuned defaults, toolset tiers, stdio shim, LAN sharing, data layout, backups, schema history |
| Providers | Capability matrix per coding agent, memory instruction layers, AGENTS.md standard, Codex hook setup and verification, writer ids |
| Retrieval | Reranker, BM25 hybrid, abstention floors, ranking-trace debugging, memory_recall, the knowledge graph |
| Dreaming | Extractor tiers, the bundled sidecar, upgrading the extractor, Sonnet-fallback, cadence, deep dream, consolidation |
| Episodes & sessions | Daemon-owned session episodes, the briefing hook, nested sub-episodes, tags |
| The memory model | Cortex slots, provenance contenders, world cortex, lessons, temporal/HLC stamps |
| Benchmarks | LongMemEval results; why extraction quality dominates |
| Comparison | Mem0, Zep/Graphiti, Letta, Cognee, memU, Memori — the axes, and when to use something else |
| Security posture | Memory poisoning (ASI06): every shipped mitigation, and what is not defended |
| Sharing one bank across machines | Exposing the daemon over Tailscale, a LAN or a proxy; per-machine principals; client-only installs |
Plus evals/README.md (full benchmark methodology) and
CONTRIBUTING.
The surface was consolidated 2026-07-02 (55 → 32 tools; now 38 with
memory_toolset, the set-slot pair and coordination): lifecycle families became verb-dispatched tools
(memory_dream, memory_forget, memory_graph_review), and
dump/introspection views moved to the Cortex Console (REST) — the manifest
is agent context every session, so it stays lean.
| Tool | Purpose |
|---|
memory_store(text, source?, tags?, origin?, episode?, authority?, distortion_tolerance?) | Remember one durable fact / decision / observation (canonical facts reach the cortex via the dream pass or memory_fact_set); authority/distortion_tolerance label the speech act and how exactly it must survive — auto (default) is a deterministic form heuristic, no model call, and both labels are inherited through supersession unless restated |
memory_search(query, top_k?, filters..., rerank?, bm25?, explain?, verbose?) | Associative retrieval; canonical cortex facts surface ahead of recall hits, each dated (asserted_at / last_confirmed / human age, plus stale when it has rotted); explain=True attaches a ranking trace |
memory_recent(n?, sources?, episodes?, tags?, verbose?) | Newest stores, timestamp-ordered (debug + session catch-up) |
memory_supersede(old_text?, new_text, entry_id?) | Correct the selected entry by ID, or one unique exact-text match; ambiguous/missing targets fail closed. Keep the old entry as history; derived_flagged names canonical facts built on it (flagged, never rewritten) |
memory_reinstate(entry_id, operation_id, expected_..., evidence_packet_sha256, reviewer_ids, reason) | Reinstate one independently reviewed retired entry under its durable ID; Postgres-only, named-principal, exact-preimage, append-only and idempotent. Refuses any trace invalidation and never confirms derived cortex facts |
memory_forget(scope, ...) | Forget from one store: memory (by text/substring/source/episode/tag — several filters narrow the match, AND across kinds like memory_search; a match over memory.delete_confirm_threshold entries, default 20, is refused with would_delete until the call repeats with confirm_bulk=true) and fact hard-delete; world and lesson (by entity/attribute) retire the slot with an audit row — reversible via memory_graph_review(action="restore_slot") |
memory_stats() | Store occupancy, hit rates, totals |
memory_agents(action, project?, task?, status?, lease?, expect?, children?, park_reason?, park_needs?, park_clear_by?, park_resume?, park_expires?) | Experimental peer awareness or update of the caller's registered context, on by default for authenticated installs (coordination); lists peers active within the hour (three hours while holding a lease), each status with its age and a stale flag past two hours, and counts the rest as idle_omitted; update sets project, task, status, expect (seconds until the status is overdue), children (the labels of subagents working under the caller's address, which only read the board; the Claude Code plugin's subagent hooks keep their own entries there, and a Codex subagent's own row carries its parent_agent_id) and the park record (why the session stopped, what clears it, who can, what to do then; a plain status clears it), which every listed row carries; claim/release take or free an advisory lease; unknown episode scope stays unknown, and activity is not a resource reservation |
memory_message(action, to?, text?, request_id?, reply_to?, after?, message_id?, clears?, urgent?) | Experimental addressed mail: send to one agent (its id, or a unique prefix of 8+ hex characters), to project:<name> or to all (every attached, non-idle peer, at most 50, one request id for the burst, per-recipient receipts), non-destructive receive, or explicit recipient ack (one id or several comma-separated, ids or prefixes); requires authenticated adapter binding, remains outside memory retrieval, and never grants user approval; a subagent's own address (a Codex subagent) does not send (child_send_refused): its parent does. Each receipt carries the daemon's wake decision (hinted, not_needed, rung, withheld with the parked need, nudged, no_path, capped): a parked peer rings only for mail that clears what it declared it needs (park records and wake) |
memory_get(entry_id) / memory_reinforce(entry_id) | Dereference a memory id to its full episode (+ consolidated_into); reinforce it after finding it useful |
memory_fact_get(entity, attribute) | The one CURRENT canonical value at a slot (+ parked contenders); on an empty slot returns ranked candidates (same-entity, then similar slots); aged/contested facts carry a ready-made correct_with call (as do memory_search / memory_world_search hits) |
memory_fact_set(entity, attribute, value, origin?, confidence?, episode?, freshness_class?, authority?, distortion_tolerance?) | Assert a canonical fact deliberately (insert / confirm / supersede / contest); freshness_class (auto default) says how fast the slot rots — auto infers it from the entity's kind; authority/distortion_tolerance (auto = deterministic form heuristic, no model call) inherit the slot's labels unless restated |
memory_fact_resolve(entity, attribute, accept) | Settle a contested slot — adopt (true) or discard (false) the contender |
memory_set_add(entity, attribute, member) / memory_set_remove(entity, attribute, member) | Add/confirm or retract one member of a set-valued slot (many concurrent values, e.g. tags — not one NOW value); a scalar there converts to a set one-way on first memory_set_add, except a number-led aggregate scalar ("32", "$1,500"), which is protected — the add parks as a contender instead. Read with memory_fact_get, which returns {kind: "set", members, removed} for these slots |
memory_history(entity, attribute?) | With attribute: version timeline at a slot, with writer/temporal stamps. Without: the entity's causal chain — dated fact/entry/edge/lesson events ("what led to X") |
memory_world_set(entity, attribute, value, source_url?, ...) | Assert a cited WORLD fact (external knowledge; age-decayed trust by freshness class) |
memory_world_search(query, top_k?, verbose?) | Search world facts — each carries effective_confidence, a stale flag, and its citation |
memory_outcome(task, outcome, about?, detail?, polarity?, episode?, used_ids?) | Record a procedural outcome signal (success/failure/correction); the dream distils signals into lessons. used_ids names the search hits the work actually turned on — each credits every retrieval_events row in the session window that served it with a retrieval_uses label (used_via=outcome), the relevance signal a learned reranker trains on; same session, within use_window_seconds, or nothing is credited |
memory_lesson_search(query, top_k?, verbose?) | Recall learned lessons for the task at hand — heed polarity - dead-ends; re_verify flags lessons whose subject facts changed since |
memory_dream(action, limit?, commit_token?, apply?, snippets?, run_id?) | Drive the dream: status / pull / commit / run (server-side extractor) / runs (audit trail of recent passes) / rollback (revert the latest committed pass from its pre-image journal) / deep (full-corpus graph consolidation; dry-run unless apply, which snapshots the graph tables first; snippets=false omits candidate evidence; responses carry evidence-enriched merge_proposals for near-duplicate triage, with long lists capped and their full counts under truncated) |
memory_graph_review(action, proposal_id?, proposal_ids?, proposals?, scope?, src?, dst?, relation?, store?) | Work the review queue: list / propose / relate (link a pair and dismiss its duplicate proposal in one call) / dismiss_pair / dismiss_slot_pair / restore_slot / accept_link / reject_link / accept_merge / accept_junk / reject_entity (merge/entity decisions are audit-stamped decided_by=agent over MCP, human via Console); list takes scope=<memory source> (the Atlas project scope) to keep only analyzer findings whose entities carry that source — queued proposals (proposed_link / merge_candidate / junk_candidate) always list; omit or "all" for everything; it is not a finding-kind filter; proposal_ids settles many id-actions in one call; restore_slot undoes a memory_forget(scope="lesson"/"world") retirement — store + the retired `entity |
memory_session_title(title, episode?) | Name THIS session's auto-opened episode (default titles are generic); episode is your session handle from the briefing — concurrent sessions share one HTTP connection, so pass it to land the rename on your own episode |
memory_episode_start(title, hint?, episode?) / memory_episode_end(episode?) | Open/close a nested sub-episode for a substantial task; entries stored while open carry its id; episode is your session handle so the nest/pop lands in your own tree when several sessions run concurrently |
memory_episode_summary(id) | Stats + tag/source distribution + recent entries within an episode |
memory_consolidation_candidates(query?, episode?, ...) | Cluster near-duplicate memories ripe for consolidation |
memory_consolidate(replaces?, new_text, source?, tags?, entry_ids?) | Replace selected entries with one canonical note; validate every ID (or unique exact text) before any changes |
memory_graph_relate(src, relation, dst, ...) | Assert a typed edge (closed relation vocabulary; re-assertion bumps confidence) |
memory_graph_unrelate(src, relation, dst) | Retract an edge (superseded, kept for audit) |
memory_alias(entity, alias) | Bind an alternative name — lookups resolve aliases first |
memory_graph(entity, depth?, include_facts?, to?, relation_filter?) | Entity neighborhood (≤3 hops) with derived transitive/inverse edges and per-edge EXTRACTED/INFERRED/AMBIGUOUS provenance tags; to returns the shortest path between two entities |
memory_recall(query, hops?, top_k?, verbose?) | Multi-hop retrieval for relational questions; low_confidence: true → fall back to memory_search |
memory_relation_define(name, description, ...) | Grow the closed relation vocabulary (deliberate, rare act) |
document_ingest(path, source?) | Index a file (txt/md/pdf/html) verbatim in the reference bank — the lossless complement to agent-side distillation (division of labor) |
document_search(query, top_k?) | RAG search over the reference bank only |
memory_toolset(action) | Check or change this principal's visibility tier: status / expand / collapse |
Each tool returns plain JSON. See pseudolife_memory/mcp_server.py for
docstrings — those are what Claude reads to decide when to call which tool.
The five recall-path tools return compact entries by default (result
payloads are agent context on every retrieval; search and recent entries
keep their write date); pass verbose=true for full metadata. Full-table dumps and topology views live in the Cortex Console
(/api/*) and the pseudolife-mcp briefing CLI.
Toolset tiers. Three visibility tiers — minimal (9 tools), core
(24), full (38) — filtered per principal at
tools/list; a principal (the named bearer-token identity, or the writer
id for single-token installs) steps its own tier up or down with
memory_toolset before calling a hidden tool. Defaults, per-client mapping, and weak-model
deployments:
Configuration — toolset tiers.
Architecture
One memory daemon owns the bank and serves MCP over streamable HTTP
at /mcp; every Claude Code session (and any LAN agent) attaches to it.
Postgres 18 + pgvector (in Docker on the durable tier; the lite tier
runs the same Postgres embedded, no container) is the durable source of
truth —
the in-memory store is a write-through cache hydrated at startup
(a small weights.pt persists only counters — there are no MLP weights).
The daemon runs either containerized (recommended — portable, no host
Python) or as a host process. Claude Code attaches through a thin
torch-free stdio shim (the installer default — per-session identity,
needed for concurrent sessions) or directly over HTTP (simpler for
a single session):
Claude session A ─┐ stdio shim (installer default) or HTTP
Claude session B ─┼───────────────────► pseudolife-mcp daemon ─► Postgres (Docker)
LAN agent ────────┘ or stdio shim (single writer) pgvector
(per session) host proc OR Docker
This kills two v0.1 hazards by construction: a single writer means
concurrent sessions can't clobber each other, and entries are transactional
so a crash can't wipe the bank. The single writer is enforced: the daemon
holds a Postgres advisory-lock writer lease on its bank. A second daemon,
or a script or eval that opens the live bank through the service, refuses
to start and names the process that holds it. Stop the daemon before
offline maintenance such as ops/dedup_cortex.py. If the bank fails to
load at startup, the daemon serves nothing rather than a partly loaded
bank, and /health reports degraded with the reason until a retry
succeeds. On top of the associative store sit the
canonical layers — cortex, world facts, lessons, temporal/HLC stamps
(the memory model) — joined to a typed
knowledge graph walkable via memory_graph and multi-hop memory_recall
(retrieval & the graph).
Install — containerized (any OS)
What the durable tier
installer above does, by hand. The whole stack — Postgres and the memory daemon — runs in Docker.
No host Python, no torch install, no version skew; the daemon image bakes
in CPU-only torch and the embedding weights — Qwen/Qwen3-Embedding-0.6B
(the default retrieval backbone since schema v25) plus all-MiniLM-L6-v2
(kept baked for the ONNX-parity test path) — so it runs identically on
Windows / macOS / Linux. Requires only Docker; built once: ~5.0 GB daemon
image (measured 2026-07-29 on the deployed build) + ~0.6 GB Postgres +
~11.8 GB extractor sidecar (measured 2026-08-20 with the v3 multi-task
bake; skip the sidecar entirely with the installer's claude-only mode).
The ~12.6 GB and ~10.4 GB figures published before
2026-07-29 are retired: both were inflated by a CUDA torch build that a
dependency-resolution bug pulled into the image (see the CHANGELOG); the
daemon has always been CPU-only.
git clone https://github.com/Pseudogiant-xr/Pseudolife-MCP.git
cd Pseudolife-MCP
docker volume create pseudolife-mcp-bank
docker volume create pseudolife-mcp-state
docker compose -f ops/docker-compose.yml up -d --build
Or skip the ~5 GB daemon build entirely and pull the prebuilt images
(releases ≥ 0.14.0):
docker compose -f ops/docker-compose.yml -f ops/docker-compose.ghcr.yml pull pseudolife-pg pseudolife-daemon
docker compose -f ops/docker-compose.yml -f ops/docker-compose.ghcr.yml up -d
The extractor sidecar is not published and still builds locally; update the
pull path with pseudolife-mcp update (backup, pinned image, rollback tag;
see Updating), not a bare pull + up -d.
Upgrading from a pre-rename install (volumes ops_pseudolife_pgdata /
ops_pseudolife_data)? Don't rename those volumes — keep pointing at them by
creating ops/.env with PSEUDOLIFE_BANK_VOLUME=ops_pseudolife_pgdata and
PSEUDOLIFE_STATE_VOLUME=ops_pseudolife_data before up. See the compose header.
Windows: cap Docker Desktop's WSL2 VM, which otherwise claims up to
~50% of host RAM — how much the stack actually needs, the
ops/wslconfig.example template, and the daemon container's own memory
cap: Configuration — Windows / WSL2 memory.
The daemon serves MCP at http://127.0.0.1:8765/mcp and restarts with
Docker — no logon task needed. First build downloads the model into the
image (once); every container start after that is offline and fast. Wire
Claude Code in via the stdio shim (installer default) or directly over
HTTP (both below). Where the data actually lives, and
how to back it up:
Configuration — data layout.
Other machines, one bank: sessions on a laptop, a server or a container
can use this daemon too, through a client-only install that runs just the
shim. Exposing the daemon (Tailscale Serve, a LAN publish or a reverse
proxy), per-machine tokens and board admission:
Sharing one bank across machines.
Host-process install (Windows, for GPU / dev): run Postgres in Docker
but the daemon on host Python — for hacking on the daemon or running the
embedder on a local GPU. Steps, the pseudolife-mcp CLI modes, and the
logon autostart task:
Configuration — host-process install.
Updating
Upgrading from before the agent board (no bearer token yet): this
migration step applies to installer-managed Docker installations. Rerun
ops/install.ps1 (Windows) or ops/install.sh (Linux / macOS) with the
same client selection (sessions can stay open: the shim installs as a new
runtime beside the running one).
The installer creates a bearer token for the default shim install and
migrates the environment of the installer-managed pseudolife-memory
stdio registration for Claude Code in place. Custom registrations are
preserved; follow the installer's printed credential warnings for Gemini
or custom registrations. -All / --all and ops/update_clients.py do not
create the token or migrate the registration environment. Then refresh clients
with the Everything at once recipe below and restart them.
One command, no checkout: the installed shim updates the whole install
from a release:
pseudolife-mcp update
pseudolife-mcp update --tag 0.15.1
pseudolife-mcp update --check
The first update from 0.15.0 or earlier has no such command (that shim
answers "unknown mode", and that daemon never announces a release): from a
checkout, git pull, then ops/update.sh --all or ops/update.ps1 -All;
for a pipx or pip install, pipx install --force "pseudolife-mcp[lite]==<version>"
or pip install --upgrade "pseudolife-mcp[lite]==<version>" (the [lite]
extra only for a lite install). After that, pseudolife-mcp update exists.
On a client-only machine (the daemon on another host), update and
--clients-only move this machine's clients to the daemon's release.
On a Docker-tier install it pulls the pinned GHCR daemon image, backs the
bank up (the checkout's backup script when the compose project still has
one, else its own pg_dump + state-volume tar into ~/.pseudolife-mcp/ backups), tags the running image for rollback, recreates only the
daemon container from the compose files it was created with (or the
package's bundled copies when the checkout is gone), waits for /health
at the new version, then installs the release as a new shim runtime beside
the running one, refreshes the plugin cache and prints the Codex hook
step. The backup stays visible and deliberate: nothing here automates it
away. --clients-only / --daemon-only take one half; the shim itself
runs from wherever it was registered, so use its full path when
pseudolife-mcp is not on PATH. On a pip / lite install it upgrades
the package in the interpreter that holds it (pip, or pipx) and restarts
nothing: it says what to restart.
You are told when it is time: the daemon checks PyPI for the newest
release in the background (updates.check_releases, on by default) and
every session's briefing then opens with release X is available — run pseudolife-mcp update. With updates.unattended_clients: true in
config.yaml (off by default) a Docker-tier shim on the daemon's host
installs the daemon's release as its own new runtime and refreshes the
plugin cache by itself whenever the daemon is newer than it; the daemon
recreate stays a deliberate command.
On a headless host, pseudolife-mcp update --schedule 03:30 installs a
daily task or timer that applies a new release only while
updates.unattended_daemon: true is set and no session is active on the
agent board, with the same backup and rollback tag, and posts a board
notice either way. Codex keeps its hook approvals across updates (it
approves the hook definitions, not the scripts), and an update from a
checkout refreshes manual Codex hook copies itself; only a changed
hooks.json makes every update path print the approval steps.
See Updating.
Lite tier by hand: one command, bank untouched:
pip install -U "pseudolife-mcp[lite]"
On Windows, first close every Claude Code, Codex and Claude Desktop
session using the shim (quit Desktop from the tray): upgrading a package
that is running can leave it half-removed.
Docker tier from a checkout: after a git pull (or local code
change), redeploy the daemon only — safely, without touching Postgres
or the extractor:
.\ops\update.ps1 # Windows
Both are thin wrappers over the same Python deploy pseudolife-mcp update
runs (ops/update.py → pseudolife_memory/update_cli.py), so there is
one implementation. It backs up the bank (pg_dump + a state-volume tar), tags a rollback
image (when a previous one exists — it says so loudly when there isn't),
rebuilds + recreates only the daemon, and waits for /health.
It never runs down -v. (Host-process install: just restart the daemon —
pip install -e . is editable.) Build cache is pruned automatically after
every healthy deploy; see
Docker disk retention for the
weekly Scheduled Task and the manual .vhdx compact. Never run
docker system prune --volumes, which deletes volumes.
The image records the commit it was built from, and /health reports it
as build (git_sha, dirty, built_at, and source: checkout or
release). So the script refuses a tree
with uncommitted or untracked files, and lists them. It also refuses a
tree git cannot describe (no git, not a clone, or git's safe.directory
refusal, which it quotes). Commit or clean up first, or pass
-AllowDirty / --allow-dirty to deploy the tree as it is: stamped
dirty: true, or unknown when git cannot describe it. Each deploy
builds a new image, so the daemon container is recreated even when the
commit has not changed.
Everything at once: the daemon is one of three installs. The shim
your clients launch and the Claude Code plugin are separate and do not
move with it. -All / --all moves them in the same run, after the
daemon is healthy:
.\ops\update.ps1 -All # Windows
No session has to be closed for the shim step. Each shim version installs
into its own runtime (%LOCALAPPDATA%\pseudolife-mcp\runtimes\NNNNNN
on Windows, ~/.local/share/pseudolife-mcp/runtimes/NNNNNN elsewhere) and
every client registers one launcher path
(%LOCALAPPDATA%\pseudolife-mcp\bin\pseudolife-mcp.exe /
~/.local/share/pseudolife-mcp/bin/pseudolife-mcp) that starts the newest
complete runtime. The installer and the update step also make
pseudolife-mcp typed in a terminal reach the launcher. On Linux and macOS
~/.local/bin/pseudolife-mcp becomes a link to it; an older pipx or
pip --user entry there is moved aside as pseudolife-mcp.<kind>-<stamp>
(never deleted), and anything else there is left alone and named. When
~/.local/bin is not on PATH, the step prints the line to add to your
shell profile: export PATH="$HOME/.local/bin:$PATH". On Windows the
launcher directory goes first on your user PATH, ahead of pipx's and
pip's copies; open a new terminal for the name to reach the launcher,
because a terminal that was already open keeps its old PATH.
pseudolife-mcp doctor reports what the name resolves to
(path_resolution) and warns when that is not the launcher; the
launcher's full path always works.
The step installs the checkout as a new runtime beside the old one, moves
any registration that still names a runtime, pipx or virtualenv path to
the launcher in place (Claude Code, Codex, Claude Desktop and Gemini CLI;
each file is backed up first as <file>.bak-<stamp>), and removes older
runtimes once no process runs from them and no registration names them.
Sessions already running keep the runtime they started with; the next
session start uses the new one. A shim running straight from this
checkout's .venv is already live and is named instead. It then
refreshes the plugin cache when its bytes differ from the marketplace
clone (claude plugin update installs the new copy beside the one
running sessions use; nothing is uninstalled), and reports whether
Codex's hook copy matches the checkout (that refresh is a consent step:
python ops/setup-codex-hooks.py). It ends with a ladder of what moved
and which clients need a restart; a client-side step that fails is
reported, never a failed deploy. The same helper runs on its own:
python ops/update_clients.py, and python ops/shim_runtime.py manages
the runtimes by hand (install, list, prune, migrate). Custom
registrations are preserved and named; upgrade those in their own
interpreter. A daemon-side credential change with an old shim leaves the
two out of step: the Claude Desktop registrar refuses a shim that cannot
read the selected token file (exit 4), and names the upgrade.
- Upgrading past 2026-09-28: board mail starts waking idle sessions.
Once the plugin cache and the shim move (
-All, then a client restart),
a Claude Code session and a Codex task with a codex CLI are woken when
mail arrives that clears the need they parked on; before this both wake
paths were opt-in. Nothing rings for chatter, and rings are capped. To
keep the old behaviour, set PSEUDOLIFE_AGENT_WAKE_HOOK=0 in the env
block of ~/.claude/settings.json and PSEUDOLIFE_CODEX_DOORBELL = "0"
in the Codex server's env table (PSEUDOLIFE_AGENT_COORDINATION=0 in
either place turns off the board for that client). pseudolife-mcp doctor shows each client's wake path under wake.
The three tell on each other: /health reports the daemon's version
and a digest of the hook scripts it shipped with; the session briefing
opens with a one-line notice when the plugin's version differs from the
daemon's (naming /plugin marketplace update pseudolife-mcp then
/plugin update pseudolife-memory@pseudolife-mcp, or the daemon redeploy,
whichever side is behind), or when the version matches but the hooks do
not (naming the --all command above); the shim says so on stderr and
ahead of its instructions; and pseudolife-mcp doctor reports
version_mismatch.
Two upgrades are not automatic, because neither can be done safely
in place. Both have a step-by-step runbook — backup, dry run, apply,
verify, roll back — and a fresh install needs neither:
- A bank older than 0.11.0 (schema v25): every embedding column moved
from
vector(384) to vector(1024), so the daemon refuses to start
rather than half-migrate. Re-embed offline with
ops/migrate_embeddings.py —
the v25 migration runbook.
- A Docker-tier bank created before 2026-08-14 (PostgreSQL 16 → 18):
a Postgres major bump cannot reuse the old data volume. Run
pwsh ops/migrate-pg18.ps1 —
the PostgreSQL 18 migration runbook.
Wire into your coding agent
Plugin (hooks + commands). The installer adds it whenever Claude Code
is a selected client (--claude-plugin skip / -ClaudePlugin skip opts
out; an installed plugin is left alone). It wires the session hooks
(briefing + episode identity), the memory-loop instructions, and the
/dream + /memory-status commands. By hand, the same two commands inside
Claude Code do it:
/plugin marketplace add Pseudogiant-xr/Pseudolife-MCP
/plugin install pseudolife-memory@pseudolife-mcp
The plugin replaces the settings.json hook, and the daemon serves a compact
memory core and a live briefing as session context. The full CLAUDE.md block
below is not served; append it if you want the complete guidance.
It deliberately does not bundle the MCP server: Claude Code loads a
plugin server alongside any user-registered one with no deduplication, which
doubled every session's tool namespace next to the installer's registration
— so the transport is registered exactly once, by ops/install.* (stdio
shim by default — per-session episode identity) or the one-liner below.
Hooks an earlier install wrote to ~/.claude/settings.json would duplicate
the plugin's; the installer offers to remove them once the plugin runs
(--claude-legacy-hooks remove / -ClaudeLegacyHooks remove unattended).
Details, non-default ports/tokens, and migration:
plugin/README.md.
Manual transport registration. The installer's default (shim mode)
registers a thin stdio shim — one shim process per session, so every
session carries its own tier-1 identity. The same wiring by hand:
pip install pseudolife-mcp
claude mcp add --scope user pseudolife-memory --env PSEUDOLIFE_MCP_NO_SPAWN=1 -- pseudolife-mcp
PSEUDOLIFE_MCP_NO_SPAWN=1 belongs on Docker-tier registrations: the shim
then waits for the container instead of spawning a host-side fallback whose
port bind can race a still-booting Docker and shadow the real bank. On the
[lite] pip tier drop the --env — there the spawn fallback is the
zero-config path.
Direct HTTP works too — the daemon serves MCP over HTTP natively (no shim,
no host command, nothing OS-specific; concurrent sessions then share one
episode identity, so it fits single-session setups best):
claude mcp add --transport http --scope user pseudolife-memory http://127.0.0.1:8765/mcp
(--scope user registers it for every project; drop it to register for the
current project only.) Or write the equivalent JSON yourself — into
~/.claude.json under the top-level mcpServers key for user scope, or into
a .mcp.json at a project root for project scope:
{
"mcpServers": {
"pseudolife-memory": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
For a token-protected daemon, add a headers key to that Claude JSON entry:
"headers": { "Authorization": "Bearer <your-token>" }.
Claude Desktop (the app, including its Cowork and Code sessions) — the
installer writes the entry for you:
ops/install.sh --client claude-desktop
Desktop has no mcp add; its servers live in claude_desktop_config.json
(macOS ~/Library/Application Support/Claude/, Linux ~/.config/Claude/,
Windows %APPDATA%\Claude\ — except the Store/MSIX build, whose real file
is under %LOCALAPPDATA%\Packages\Claude_*\LocalCache\Roaming\Claude\; the
installer prefers that path when it exists). The equivalent entry by hand:
{
"mcpServers": {
"pseudolife-desktop": {
"command": "/absolute/path/to/pseudolife-mcp",
"env": {
"PSEUDOLIFE_WRITER_ID": "claude-desktop",
"PSEUDOLIFE_MCP_NO_SPAWN": "1",
"PSEUDOLIFE_MCP_DAEMON_URL": "http://127.0.0.1:8765"
}
}
}
}
The entry is named pseudolife-desktop on purpose. Desktop's Code tab runs
Claude Code, which starts its own per-session pseudolife-memory server; when
an app-level entry has the same name, Desktop sends the session's
mcp__pseudolife-memory__* calls to the app-level entry and the session's own
server gets none. Re-running the installer renames an entry it wrote under the
old name (its env sets PSEUDOLIFE_WRITER_ID to claude-desktop), keeping
any env keys you added and backing the config up first. It leaves any other
pseudolife-memory entry alone and warns about it. Chat and Cowork then list
the tools as mcp__pseudolife-desktop__*.
Two things differ from the CLI clients. Desktop launches MCP servers with a
sanitized environment — PATH plus a few system variables, none of your
shell's exports — so command must be the shim's absolute path (which pseudolife-mcp / Get-Command pseudolife-mcp), and a token-protected
daemon needs "PSEUDOLIFE_MCP_TOKEN_FILE": "/absolute/path/to/a/private/file"
in that env block: a bearer exported in your OS environment never reaches
the shim. The symptom of forgetting it is every session failing with
"Couldn't start for Cowork and Code sessions. Error: unhandled errors in a
TaskGroup (1 sub-exception)" — a 401 under the SDK's wrapper, which the
shim now names plainly on stderr at startup. When the daemon is
token-gated the installer writes that file (owner-only) from
PSEUDOLIFE_MCP_TOKEN in its environment or ops/.env, or migrates a
literal token already in the entry into it; with no token to write it
says so and exits 3. The shim must also be able to read that file:
releases through 0.15.0 only read the literal PSEUDOLIFE_MCP_TOKEN, so
the registrar probes <command> --help for the file form first and
refuses an older shim (exit 4, nothing written) rather than register an
entry that would fail with the same TaskGroup error — upgrade the shim
(re-run the installer, or python ops/update_clients.py --only shim from
the checkout: it installs a new runtime beside the running one, so no
session has to close) and re-run. After any edit, fully quit Desktop from the tray or
menu-bar icon and relaunch — closing the window does not reload the
config.
Codex — the installer's default (shim mode) wires the same stdio shim, so a
Codex session gets its own tier-1 identity instead of inheriting a
concurrent Claude session's episode:
pip install pseudolife-mcp
codex mcp add pseudolife-memory --env PSEUDOLIFE_MCP_NO_SPAWN=1 -- pseudolife-mcp
(Same Docker-tier note as the Claude wiring above: keep
PSEUDOLIFE_MCP_NO_SPAWN=1 when the daemon runs in Docker; drop it on the
[lite] pip tier.)
The HTTP one-liner works too (no pip package needed):
codex mcp add pseudolife-memory --url http://127.0.0.1:8765/mcp
Or add the equivalent user-level entry to ~/.codex/config.toml:
[mcp_servers.pseudolife-memory]
url = "http://127.0.0.1:8765/mcp"
bearer_token_env_var = "PSEUDOLIFE_MCP_TOKEN"
For that Codex HTTP configuration, export PSEUDOLIFE_MCP_TOKEN in the
environment that launches Codex. The token stays out of config.toml, and
Codex reads it when connecting. This is unnecessary for the default stdio shim.
Gemini CLI — same shape (-s user matters: Gemini defaults to project
scope; the -e env gives Gemini sessions their own write attribution, and
the same Docker-tier PSEUDOLIFE_MCP_NO_SPAWN=1 note as above applies):
pip install pseudolife-mcp
gemini mcp add -s user -e PSEUDOLIFE_WRITER_ID=gemini -e PSEUDOLIFE_MCP_NO_SPAWN=1 pseudolife-memory pseudolife-mcp
Or HTTP, no pip package needed:
gemini mcp add -s user -t http pseudolife-memory http://127.0.0.1:8765/mcp
Note: since 2026-06-18 Google no longer serves individual-tier accounts
(free, AI Pro, AI Ultra) through Gemini CLI — OAuth sign-in fails and
points at Antigravity. The wiring above stays correct, but individual
accounts need API-key auth (GEMINI_API_KEY) to actually run sessions —
or use Google Antigravity itself, which connects to the same bank via
~/.gemini/config/mcp_config.json; both are covered in
the providers guide.
Any other MCP-capable agent (Cursor, Windsurf, Zed, Copilot CLI, …) —
add the generic mcpServers entry to that tool's MCP config (ops/install.sh --client generic prints both shapes ready to paste):
{
"mcpServers": {
"pseudolife-memory": {
"command": "pseudolife-mcp",
"env": {
"PSEUDOLIFE_WRITER_ID": "mcp-client",
"PSEUDOLIFE_MCP_NO_SPAWN": "1"
}
}
}
}
What each agent gets — and what its platform can't support (hooks,
per-turn discipline): the provider matrix.
Verify: run claude mcp list, codex mcp list, or gemini mcp list
(the server should report connected), then ask the agent to "store a memory
that this install works" and check it
appears in the Stream tab of the Console at http://127.0.0.1:8765/ui/.
Preferring stdio (this is what the installer wires by default, for
per-session identity)? A thin torch-free shim proxies stdio to the
daemon:
stdio shim
· sharing one bank across machines
· backups & restore rehearsal
· agent mailbox recovery.
Recommended agent setup (CLAUDE.md / AGENTS.md)
The server's value depends on the agent using it. The MCP server advertises
the core loop through protocol-level instructions; the shim adds the
messageboard check-in when its coordination adapter is up.
The plugin's memory SessionStart hook (also used by verified Codex hooks)
delivers a short operating guide and a bounded briefing; without the plugin,
the installer's Claude Code settings.json hook
(pseudolife-mcp briefing --hook-json) delivers the same two.
A separate coordination hook asks the agent to set its project,
task and status, discover peers, and read pending messages, but only where
the board is on for that credential (it is on by default behind bearer
authentication, so an open install or a disabled board adds no check-in).
Detailed memory
guidance remains in the bundled standing block. Hooks add per-prompt reminders
and session bookkeeping; neither delivery method
guarantees that the model performs every requested memory operation.
Hooks serve at most that short guide; the detailed block reaches an agent
only as a standing copy. For the complete guidance, for subagent
visibility (subagents read CLAUDE.md but not hook output), or in place of
hooks, append it to Claude's global ~/.claude/CLAUDE.md, Codex's
global ~/.codex/AGENTS.md, Gemini's global ~/.gemini/GEMINI.md, or a
per-project CLAUDE.md / AGENTS.md:
cat examples/CLAUDE.memory.md >> ~/.claude/CLAUDE.md
cat examples/CLAUDE.memory.md >> ~/.codex/AGENTS.md
cat examples/CLAUDE.memory.md >> ~/.gemini/GEMINI.md
Add-Content "$env:USERPROFILE\.claude\CLAUDE.md" (Get-Content examples\CLAUDE.memory.md -Raw)
Add-Content "$env:USERPROFILE\.codex\AGENTS.md" (Get-Content examples\CLAUDE.memory.md -Raw)
For hook-less providers this standing block supplies the full memory policy,
but it cannot provide a live briefing or run session cleanup. AGENTS.md is the cross-vendor standard for standing
agent instructions (Linux Foundation-governed; read by Codex, Copilot,
Cursor, Gemini CLI, Zed, and 30+ others), so a per-project AGENTS.md
carrying the block reaches almost every agent at once. Claude Code is the
holdout — it reads CLAUDE.md — but a CLAUDE.md whose first line is
@AGENTS.md imports the shared file, so one copy serves every tool.
The block (examples/CLAUDE.memory.md) teaches
the loop: RECALL at the start (memory_search / memory_lesson_search /
memory_fact_get / memory_world_search), CAPTURE as you go
(memory_store with an honest origin, memory_fact_set for canonical
facts, memory_world_set for cited external facts, source="status" for
verbose logs so they stay out of the dream), REFLECT at the end
(memory_outcome, with used_ids naming the hits you actually used — the
dream distils these signals into the lessons surfaced at your next session
start).
For an existing Codex installation, run python ops/setup-codex-hooks.py.
The helper asks once, detects the hook source, backs up changed configuration,
persists scoped trust through Codex, and verifies startup briefing, prompt
reminder, and session cleanup, plus the agent-board check-in where the board
is on. The Docker installer runs this step for you.
See Codex setup options and fallback.
For Claude Code, use the plugin, or the legacy
ops/install-hook.ps1 -Client claude / ops/install-hook.sh --client claude
for briefing and reminder hooks. The legacy --client codex path remains
available but only writes hook definitions; it does not complete trust and
verification. Session episodes also work without hooks through the daemon:
Episodes & sessions.
Current Codex runtimes enable hooks by default, including Windows.
Availability depends on the application/runtime and policy, not the model.
If [features] hooks = false is intentional, keep it and use the standing
AGENTS.md block. Codex runs the plugin's Windows hooks as native
PowerShell 7 commands; Claude Code runs their Bash commands through Git
Bash, so it needs Git for Windows installed (see
plugin/README.md).
See the official hook protocol.
Codex hook trust: setup approval is limited to PseudoLife's current hook
definitions: the memory and coordination SessionStart and UserPromptSubmit
handlers and SessionEnd, plus the plugin's Stop entry (Claude Code's wake
hook, on by default; in Codex it runs only the park gate). It does not approve other plugins or bypass
future trust checks. Codex approves the definitions, not the scripts they
run, so the approval also covers the scripts later PseudoLife updates
install (for manual copies the update installs them itself, verified by
content). Changed definitions need approval again, and so does a
handler a plugin update adds; Codex skips an unapproved one silently in the
desktop app, which pseudolife-mcp update (from a checkout, ops/update.ps1 -All or ops/update_clients.py) reports as needs-approval. If automatic setup cannot
use the installed runtime's trust interface, it reports the problem and
asks you to open /hooks to review and trust the definitions. Approved standing
instructions remain available as fallback. Installed files alone do not
establish that hooks are working.
Usage patterns
At session start — loads what you've worked on before, persistent
across compactions:
memory_search("project context for X")
During work — store real decisions; skip fleeting chatter (the shipped
store gate is permissive, so deliberate, durable claims only):
memory_store("Decided to use stdio transport for the MCP because no port conflicts", source="pseudolife")
When corrected — marks the old fact superseded and stores the
correction; both surface in future retrieval, the new one ranked higher.
Select the entry by the id carried on the search or recent hit:
memory_supersede(
entry_id=417,
new_text="Provider interface uses async calls — sync version was the v0.7 prototype only"
)
old_text= still selects by the full stored text when that text is exactly
unique among live entries; it is the legacy selector and the only one file
mode has. Ambiguous or missing targets change nothing.
Hygiene — memory and fact scopes hard-delete (at least one filter
is required for scope memory, preventing accidental wholesale deletion;
several filters narrow the match, so text="probe", source="status" is
the probe entry in that source, not the whole source; a match over
memory.delete_confirm_threshold entries — default 20, 0 disables — is
refused with would_delete until the call repeats with
confirm_bulk=true); lesson and world scopes retire the slot with an
audit row and are reversible with memory_graph_review(action="restore_slot", store=..., src="entity|attribute"); for "keep the history but mark it
wrong" use memory_supersede instead:
memory_forget(scope="memory", source="test-noise") # refused past 20 matches
memory_forget(scope="memory", source="test-noise", confirm_bulk=True)
memory_forget(scope="fact", entity="test-entity")
Discovering what's in the bank: open the Cortex Console — sources, tags,
episodes, and full-table views all live there. Going deeper:
reranking, BM25, abstention, and trace debugging
· episodes + tags
· canonical facts, contenders, world facts, lessons
· the consolidation workflow.
Dreaming — consolidating memories into facts
A dream distils the recent associative stream into canonical cortex
facts while you're not looking: pull unconsolidated memories → extract
(entity, attribute, value) → acknowledge those exact entries durably.
New entries remain pending regardless of their timestamps; failed acknowledgement
can replay claim application. Manual commits use the token returned by pull.
Extraction is pluggable:
| Tier | How it runs | Needs | Quality |
|---|
| 0 — none | no extractor configured — the dream still runs, prunes, and acknowledges input batches, but writes no canonical facts | nothing | none (memory_fact_set is your only cortex writer) |
| 1 — agent-driven | the agent itself is the gateway: the /dream judgment session (its manual-extraction branch fires only when no endpoint is configured) | the agent you already run | highest |
| 2 — shipped default | daemon auto-sweep → the bundled CPU sidecar, or any OpenAI-compatible endpoint | nothing (sidecar) | high; free if local |
The stack ships tier 2 preconfigured (the bespoke Gemma 4 E4B extractor
fine-tune in a llama.cpp sidecar, internal-only). The sweep cadence,
pointing dreams at a bigger local model or at Claude Sonnet with automatic
sidecar fallback, the full-corpus deep dream graph pass, and the
privacy/cost trade-offs: Dreaming.
Benchmarks
The headline is the whole benchmark, not a slice: all six
LongMemEval question types, 500
questions, oracle variant, run end to end through the memory (qwen-27b
extraction under the v25 embedding backbone, BM25-on turn retrieval).
Single pass, graded by the local Qwen3.6-27B bench judge (2026-08-03):
| arm | accuracy | context tokens/question |
|---|
| naive RAG (top-6 turns) | 0.688 | ~1210 |
| cortex facts only | 0.416 | ~158 |
| hybrid (facts + top-3 turns) | 0.664 | ~842 |
| commit-gated cascade | 0.690 | ~883 |
The cascade is a serving policy, not a fourth pipeline: answer from
the consolidated facts when that channel commits, fall back to raw-turn
RAG when it abstains. Overall this is a wash on accuracy at ~73% of the
context — 0.690 vs 0.688 is one question in 500 on a single pass, and
nobody should read it as a win. The fact spine alone answers at ~13% of
RAG's token budget, at a large accuracy cost outside the types it is built
for. The structure is per type:
| question type | n | naive RAG | commit-gated cascade |
|---|
| knowledge-update (facts change) | 78 | 0.859 | 0.936 (retired — why) |
| single-session-user | 70 | 0.929 | 0.943 |
| single-session-assistant | 56 | 0.911 | 0.929 |
| single-session-preference | 30 | 0.800 | 0.700 |
| temporal-reasoning | 133 | 0.526 | 0.526 |
| multi-session | 133 | 0.504 | 0.474 |
The consolidated spine helps where a fact changes and where the answer
sits inside one session; it loses where the answer must be aggregated
across sessions or ordered in time, because per-fact consolidation is
exactly what discards that structure. BEAM-100K reproduces the same shape
independently. Its abstention questions are where the spine looks best —
the fact-spine arm scores 0.950 against naive RAG's 0.775, identical under
the local judge and under an independent Opus-class judge — but the
2026-09-02 five-arm run bounds that reading: a no-memory arm scores 1.000
there, so the edge is calibration (a small fact context refuses where raw
turns confabulate), not evidence that the memory recalled anything. Setup, caveats,
both bench stacks side by side, and the evidence that extraction quality is
the dominant factor: Benchmarks; full
methodology: evals/README.md.
Retrieval itself was re-measured on the same corpus before the v25 backbone
swap (150 questions, 74,183 haystack turns, 299 gold turns; pure recall — no
reader, no judge): Qwen/Qwen3-Embedding-0.6B reaches R@10 0.809 against
bge-base-en-v1.5's 0.742 and the previously-shipped all-MiniLM-L6-v2's
0.572, and beats bge-base head-to-head +32/−12 at k=10 (p=0.004).
Artifacts: embedder-recall-shootout-20260727.json,
embedder-recall-qwen-vs-bge-20260728.json.
Cortex Console (web UI)
An operator dashboard served by the daemon itself — point a browser at
http://127.0.0.1:8765/ui/ (the /health and /mcp endpoints are
unchanged; the console is additive). It's a read-mostly instrument panel for
seeing and steering the memory a human otherwise can't observe:
Observatory (health, per-layer counts, the memory store's capacity meter, dream
gauges), Cortex (canonical facts with provenance, version-history
timelines, inline Accept/Discard for contested slots), World / Lessons /
Episodes, Stream (live search with rerank/BM25 toggles and a
ranking-trace debugger), Graph (interactive force-directed visualiser, with a review drawer that
can Accept/Reject merges or — for a source file and its own bare concept,
band.py ↔ band — record an implements edge instead of forcing
merge-or-dismiss; proposals a background dream has already judged carry a
verdict chip — accept/reject/leave with confidence, the model's reason in
the tooltip — as a lead, never a decision), and Console (every safe config.yaml scalar with live-vs-restart
badges, diff-preview, and atomic save).
Coordination is a read-only board view: active roster/status age, reported
children, park needs/resume, resource holders and FIFO queues. Pending mail
counts are visible only for the caller's principal; the latest 100 retained
send/read/ack/wake/expiry events omit all message bodies. This view requires
a configured bearer on coordination.allowed_principals, even on an otherwise
open loopback installation. Refresh neither receives/acknowledges mail nor
settles expired leases. Snapshots are marked stale after a minute; use Refresh
for current state. The API is GET /api/agents?view=coordination&limit=50
(roster limit 1–50); the default /api/agents awareness response is unchanged.
Incoming audit visibility uses current recipient ownership, so incoming events
may no longer appear after the recipient address is pruned; audit retention
also bounds the timeline. A cold durable tier reports unavailable rather than
initializing the bank through this read.
Auth mirrors /mcp: /ui (static shell) and /health are open; /api/*
requires the same PSEUDOLIFE_MCP_TOKEN bearer when one is set (the console
prompts for it and stores it locally). No build step, no CDN, fully offline —
vanilla ES modules + vendored OFL fonts served straight from the daemon.
Developing the UI? A fixture-backed dev server (no Postgres, no torch)
renders the real frontend against canned data:
python -m pseudolife_memory.web.devserver → http://127.0.0.1:8770/ui/.
Its payloads self-announce ("fixtures": true on /health), and the
topbar shows a "DEMO DATA — fixture server, not a real bank" chip in
place of the live chip, so a fixture run is never mistaken for a real
bank.
Capabilities at a glance
| Capability | Status |
|---|
| Transport | Streamable-HTTP MCP daemon (/mcp); stdio shim is the installer default (per-session identity) — HTTP remains for single-session setups |
| Storage | Postgres 18 + pgvector (source of truth); ChromaDB for the reference bank |
| Associative store | Flat similarity store (default since the 2026-08-15 measured verdict; the 8-tier banded preset remains opt-in); hybrid dense + BM25 ranking (BM25 on by default); contradiction detection admits potential updates, including a deterministic slot-identity path regardless of embedding similarity, while retaining earlier source notes; whole-note supersession requires an explicit replacement operation |
| Canonical-fact cortex | Single-writer: LLM dream pass + memory_fact_* (regex auto-promote opt-in, default off) |
| Set-valued slots | memory_set_add / memory_set_remove for many-current-value slots; one-way scalar→set conversion, aggregate scalars guarded (park as contender); an assistant-origin add cannot convert or join another tier's set, and cannot retract another tier's member |
| Provenance contenders | Tier-rank guard user > action > agent > assistant; memory_fact_resolve |
| Fact currency | Every cortex fact is dated (asserted_at / age); freshness_class (evergreen / slow / volatile) decays effective_confidence and flags stale. Left auto, the class is inferred from the entity's kind (schema v24 entity_kinds) — only system entities can rot; artifacts and concepts stay evergreen |
| Write-time labels | authority (directive / observation / quoted — the speech act, orthogonal to the origin tier) and distortion_tolerance (constraint / procedural / belief / preference / episodic) on entries and facts, set at write time (explicit, or a deterministic heuristic under auto) and inherited through supersession unless restated. A constraint source is carried verbatim through the dream (with a post-dream guard) and pinned ahead of cosine in memory_search's cortex block and memory_recall when the query names its entity; a quoted source is low-trust for the two-man rule (schema v35) |
| Knowledge graph | Typed entities/edges, closed relation vocab, on-read closure (Postgres + NetworkX, no AGE/Cypher) |
| World cortex | memory_world_* — cited external facts + age-decayed freshness (manual ingest) |
| Procedural memory | memory_outcome (signals) → dream-synthesised lessons via memory_lesson_search; prefers/avoids graph edges; single-writer |
| Sense of time + multi-writer | Per-write stamp (tx/valid time, HLC ordering, writer/session); memory_history; relative age on reads; write_mode seam (snapshot live, occ Phase-2) |
| Episodes + tags | Session episodes daemon-owned, keyed by a resolved five-tier session identity; hook eager-open or lazy-open on first store + idle reaper + prune-empty + resume-after-reap; nested sub-episodes with subtree-expanded recall; multi-valued tags=[...] |
| Session briefing | SessionStart hook injects lessons + verified world facts + last-session recap once per conversation (resume/compact get the episode handle only); the plugin's per-turn note speaks only when lessons or other sessions' status notes changed (pseudolife-mcp briefing adds the unsure-graph section) |
| Consolidation | memory_consolidation_candidates + memory_consolidate |
| Optional components | Cross-encoder reranker (rerank=True, ~80 MB); ONNX embedding backend (pip install .[onnx] — load-only, and auto-selected when installed and the configured model's artifact is already on disk, ~3x faster CPU encode on MiniLM. The configured artifact must already exist locally: the daemon image provisions MiniLM's while building, while a pip install stays on torch until you provision it yourself. Models whose Transformer module loads from a subfolder use torch on native Windows, and the default Qwen3-Embedding-0.6B has no ONNX export at all); NLI contradiction scorer (pip install .[nli], ~278 MB) |
| Web console | Cortex Console at /ui/ — health/stats, fact review + history, graph visualiser, search/trace, config editor (read-mostly, token-gated like /mcp) |
| Schema version | v52 (Postgres meta version) — additive ADD COLUMN IF NOT EXISTS migrations on daemon start, except v25: the vector(384)→vector(1024) move is not additive, so the daemon refuses to start against an older-dimensioned bank until you run ops/migrate_embeddings.py; legacy file-mode .pt banks auto-migrate into Postgres; full version history |
Troubleshooting
Start with curl http://127.0.0.1:8765/health — it reports the daemon's
package version, the schema version, storage backend, auth state, and
persist_errors (non-zero means
writes are failing to reach Postgres; check docker logs pseudolife-mcp-daemon).
- The cortex stays empty (canonical facts never appear on their own).
/health reporting "extractor": "none" means no extractor is
configured, so the dream writes no facts and memory_fact_set is the
only cortex writer — expected on the lite tier. Point the daemon at an
OpenAI-compatible endpoint
(Quickstart) or use
the Docker tier's bundled sidecar. "extractor": "disabled" instead
means dreaming itself is switched off in config.
- Lite daemon refuses to start on Windows with a message about the data
path: the embedded Postgres runtime needs an ASCII-only data
directory. Set
PSEUDOLIFE_MCP_DATA_DIR to one (e.g.
C:\pseudolife-data) —
Configuration.
- First build is slow / big. The daemon image (~5.0 GB, several
minutes to build) bakes in CPU torch and the embedding weights (Qwen3-Embedding-0.6B
plus MiniLM); the extractor sidecar
adds a ~5.3 GB model download on its first build. Every start after that is
offline and fast — if a rebuild is re-downloading models, the Docker
layer cache was pruned.
- Daemon unreachable after
wsl --shutdown (Windows): the host port
forward is gone — docker restart pseudolife-mcp-daemon re-establishes it.
- Docker eating RAM (Windows): the WSL2 VM (
Vmmem) claims up to ~50% of
host memory by default. Copy ops/wslconfig.example to
%USERPROFILE%\.wslconfig, tune memory=, then wsl --shutdown.
- Port already in use: the stack binds
127.0.0.1:8765 (daemon) and
127.0.0.1:5433 (Postgres). Change the host side in
ops/docker-compose.yml if either collides.
- Console shows "offline" / Unauthorized: "offline" means the daemon
isn't reachable (see above); a 401 prompt means it runs with
PSEUDOLIFE_MCP_TOKEN — paste that token into the Console's Token dialog.
- The coding agent doesn't see the tools:
claude mcp list or
codex mcp list should show
pseudolife-memory ✓ connected. If not, re-check the URL
(http://127.0.0.1:8765/mcp — the /mcp path matters) and the bearer
header when a token is set. The daemon preloads the embedder on a warmup
thread at start (~5–10 s); a very early first call can race it and take a
few seconds.
- Tools vanish after an upgrade / the client log says "Connection
closed": the shim's registered command can live outside the repo venv,
and the MCP SDK v2 migration set an
mcp>=2.1 floor — an older SDK in
that environment crashes the shim on start. The shim detects this and
prints the interpreter path and the exact fix on stderr: pip install -U "mcp>=2.1,<3" in that interpreter, or re-run the installer (which
registers the project venv's shim).
- Claude Desktop says "Couldn't start for Cowork and Code sessions. Error:
unhandled errors in a TaskGroup (1 sub-exception)": the real exception
is at the bottom of
mcp-server-pseudolife-desktop.log
(mcp-server-pseudolife-memory.log for an entry not yet renamed) in the app's log
folder (%LOCALAPPDATA%\Claude\Logs on Windows, ~/Library/Logs/Claude
on macOS). If it is a 401, the daemon is token-gated and the
Desktop-launched shim holds no credential: Desktop sanitizes the
environment, so put PSEUDOLIFE_MCP_TOKEN_FILE in the entry's env (see
Claude Desktop under Wire into your coding
agent) or re-run ops/install.* --client claude-desktop, then fully quit and relaunch. On the Windows Store build
the file lives under
%LOCALAPPDATA%\Packages\Claude_*\LocalCache\Roaming\Claude\. If the
entry already carries PSEUDOLIFE_MCP_TOKEN_FILE and still 401s, the
shim predates token-file support (PyPI releases through 0.15.0 read only
the literal token): pseudolife-mcp --help from a capable shim lists
PSEUDOLIFE_MCP_TOKEN_FILE; upgrade the shim (python ops/update_clients.py --only shim from the checkout installs a new
runtime beside the running one, so no session has to close) and re-run
the installer, which now refuses to register an older one against a
token file.
- A harness "removed tools" notice is not an outage. A resumed session
can carry a larger tool roster in its transcript than the current
toolset tier serves —
that's a visibility filter, not a disconnect. Make one
memory_search
call before concluding the MCP is down.
Uninstall
Lite tier: remove the MCP registration (claude mcp remove pseudolife-memory / codex mcp remove pseudolife-memory), then
pip uninstall pseudolife-mcp. If you also want the bank gone, delete
the per-user data directory (%LOCALAPPDATA%\pseudolife-mcp on Windows,
~/.local/share/pseudolife-mcp on Linux, ~/Library/Application Support/pseudolife-mcp on macOS — or wherever PSEUDOLIFE_MCP_DATA_DIR
points). Back it up first: pseudolife-mcp backup works on lite too.
Docker tier — deletion is deliberate at every step:
docker compose -f ops/docker-compose.yml down
claude mcp remove pseudolife-memory
codex mcp remove pseudolife-memory
gemini mcp remove pseudolife-memory -s user
docker volume rm pseudolife-mcp-bank pseudolife-mcp-state
Host-process installs: also unregister the logon task
(Unregister-ScheduledTask -TaskName "Pseudolife-MCP Daemon") and remove
the SessionStart briefing hook — plus the UserPromptSubmit memory-change
hook (pseudolife-mcp prompt-hook; older installs: the discipline echo) — from ~/.claude/settings.json and/or
~/.codex/hooks.json (a timestamped .bak-* sits next to each edited file).
Testing
pip install -e .[dev], then pytest tests/. The suite covers every
layer, from the MemoryService surface to the Cortex Console REST API;
model-heavy pieces are stubbed so it stays fast and offline. The PG-backed
suites each target a throwaway per-run pseudolife_memory_test_<pid>
database on the bundled dev container (never your real bank; concurrent
runs can't collide), dropped on exit, and skip cleanly without Postgres.
The container's password is read from ops/.env (override with
PSEUDOLIFE_TEST_DATABASE_URL or PSEUDOLIFE_TEST_PG_PASSWORD); a
server that is reachable but rejects the credentials errors the PG-backed
tests rather than skipping them, so a rotated password can never produce a
green run by accident. Full dev setup: CONTRIBUTING.
What's not built yet
- Reflection via MCP sampling — would let the dream borrow Claude
itself as the extractor;
Claude Code doesn't yet support it.
- Cross-machine sync — memory lives on one PC's disk; syncing via
rclone / syncthing is left as an exercise.
- Automated world-knowledge ingestion — populating the world cortex
from the live web needs a web-fetch tool the standalone server doesn't
ship; an agent with web access can automate the fetch+cite step today
via
memory_world_set.
Support
Solo-maintained, best-effort. One person builds, tests, and runs this;
there is no support contract and no response-time commitment. That said,
issues are read and most get an answer.
- Something is broken → open a
bug report.
The form asks for your
/health output, schema version, install tier,
and client, because those four answer most questions before any
back-and-forth.
- Something is missing → open a
feature request.
Say what you were trying to do, not only what to add.
- A security problem → do not open a public issue. Use GitHub's
private vulnerability reporting — SECURITY.md. Memory
integrity specifically: security posture.
- Sending a patch → CONTRIBUTING and
CODE_OF_CONDUCT. The bar is "surgical, tested, and
explained", not "big".
If you need someone to call, Comparison — use something else
if names vendors who sell
support.
License
Apache-2.0 — see LICENSE and NOTICE.