io.github.lacausecrypto/sophon is an MCP server focused on “honest token economics” for MCP agents. It provides deterministic context compression so agents use less context. The project is described as one Rust binary with zero machine learning at query time and reproducible benchmarks and real-data measurements.
🛠️ Key Features
Deterministic context compression for MCP agents
One Rust binary
Zero ML at query time
Reproducible benchmarks and real-data measurements
“68% tokens saved” claim
🚀 Use Cases
Reducing token usage for MCP agent context handling
Running comparable compression benchmarks with real data
⚡ Developer Benefits
Deterministic compression behavior
Reproducible benchmark results
Simple deployment via a single Rust binary
⚠️ Limitations
The provided description does not specify supported MCP tools, configuration options, or runtime interfaces beyond context compression.
Sophon is a deterministic context layer for agents speaking the Model Context Protocol. It compresses prompts, conversation memory, code digests, file deltas, and shell output — without an embedding model at query time, without a GPU, and without API keys.
Single 5.2 MB Rust binary. MCP-native. cl100k_base-accurate. Default build pulls no Python, no ML weights, no network.
What it does, in 30 seconds
Tool
What it solves
compress_prompt
Long structured prompt → keep only sections relevant to the query
Real numbers — measured on this repo's own dev cycle
We built four independent benches that each capture a different chunk of an agent's tool traffic. All four run against this repo's actual git history + working tree on the operator's machine. Reproducible byte-for-byte by anyone with cargo build --release.
real_session_holistic.py runs all four sub-benches with --json, parses them, and produces the weighted blend. Default weights reflect this repo's observed shape; pass --weights "history=0.4,..." to model your own workload.
USD economy on Claude Opus 4.7
Saved per session
Naive input pricing ($15/MT)
$2.03
With prompt caching (25-turn reads at $1.50/MT)
$3.24
Pass --model sonnet or --model haiku to real_session_deep_dive.py if you're re-pricing for a cheaper tier.
Where each dimension falls short (we say it ourselves)
history measures only what git captures (commits + diffs) — typically ~5-10 % of a real session's tool traffic. The 94.6 % is the upper bound, not the typical case.
shell mixes commands that compress well (git diff 95 %) with commands that don't (gh repo view --jsonadds tokens, −9 %). 84.4 % is a real-world average, not a curated highlight.
filereads uncovered that compress_prompt on raw source files compresses by budget cap, not by query routing — same file with 3 different queries → identical output. Section detection only fires on structured input (Markdown headers, XML tags). Documented inline in the bench.
search depends entirely on YOUR repo's state. A repo with no TODOs gets 0 % on grep TODO.
The blended 84.7 % is napkin-math from a linear weighted average across four real measurements. Not a cherry-picked synthetic. Run the benches yourself to verify.
Other reproducible benchmarks (synthetic, on-thesis)
Sophon is not a memory platform, a recall system, an OCR stack, or a replacement for provider-side caching. It's a deterministic compressor that slots in front of whatever memory / cache / code-nav layer you already use, and attacks the tokens those layers can't.
In front of Anthropic / OpenAI prompt caching
Provider caching handles the static half of a request — system prompt, tool definitions, reused documents. It doesn't touch the dynamic half (growing conversation history, tool outputs). Sophon compresses exactly that half. The two stack cleanly.
+24 % tokens / +49 % $ saved on top of prompt caching on a 25-turn Claude session — because the uncached dynamic block is billed at 10× the cached rate. See sophon_plus_prompt_caching.py.
In front of mem0 / Letta / Zep / Graphiti
Memory systems retrieve the right memories. Sophon shrinks what gets sent to the LLM after retrieval. If mem0 returns 2 kB of raw memories, compress_prompt keeps only the sections the query actually references.
Honest caveat: on very short retrieved blocks (< ~200 tokens) Sophon's wrapper adds overhead and you should pass through. The bench reports this directly.
In front of Claude Code / Cursor / Cline
Primary use case. Every repeat file read becomes a read_file_delta; every shell command output goes through compress_output; every repeated boilerplate block gets a fragment_cache token. Install transparently with sophon hook install --agent claude --global.
In front of a RAG pipeline
navigate_codebase produces a PageRanked repo digest that a RAG retriever would otherwise spend expensive embedding calls to build. Tree-sitter / regex symbol extraction over 11 languages, sub-second.
When NOT to use Sophon
Long-form conversational recall above 80 % — Sophon caps at ~40 % on LOCOMO and we don't chase it. Run mem0 / Letta / Zep for recall, then optionally pipe their output through Sophon.
Multi-hop reasoning on massive documents — that's HippoRAG or GraphRAG.
OCR / PDF layout — out of scope. Use Docling / Marker / Unstructured upstream.
Very small inputs (< ~200 tokens) — Sophon's section scaffolding can cost more than it saves.
Quick start
Install via npm (recommended)
bash
npm install -g mcp-sophon
sophon doctor # verify install + show config
The postinstall script downloads the right prebuilt binary for your platform from the GitHub Releases page. Supported: macOS arm64/x64, Linux arm64/x64, Windows x64.
MCP protocol:2025-06-18. notifications/cancelled actually drops the response (since v0.5.4). Structured JSON-RPC error codes (-32000..-32099 reserved for Sophon). Infallible dispatcher — a malformed request can't kill the stdio loop.
Configuration
Run sophon doctor to see every SOPHON_* env var currently set with validation warnings. Full catalogue (24 flags) lives in runtime_flags.rs. The flags worth knowing:
Flag
Effect
Cost
SOPHON_RETRIEVER_PATH=/dir
Activate the semantic retriever (chunk store on disk)
~0
SOPHON_MEMORY_PATH=/file.jsonl
Persistent conversation memory across sophon serve runs
~0
SOPHON_HYBRID=1
BM25 sparse-lexical + HashEmbedder fused via RRF
~1 ms
SOPHON_ROLLING_SUMMARY=1
Build rolling summary at update_memory time, not at query time
LLM call moved to ingest
SOPHON_CHUNK_TARGET=500
Bigger chunks preserve cross-sentence context
~0
SOPHON_EMBEDDER=bge
Swap HashEmbedder for BGE-small (needs --features bge)
model load at startup
SOPHON_LLM_CMD="claude -p --model haiku"
LLM shell-out command (used by summarizer when configured)
per-call subprocess
Deprecated v0.4.0 recall-chasing flags — SOPHON_HYDE, SOPHON_FACT_CARDS, SOPHON_ENTITY_GRAPH, SOPHON_ADAPTIVE, SOPHON_LLM_RERANK, SOPHON_TAIL_SUMMARY, SOPHON_REACT, SOPHON_GRAPH_MEMORY, SOPHON_MULTIHOP_LLM — chase LOCOMO recall, an axis we no longer optimise. Still functional but sophon doctor flags them. Removed in a future major.
LOCOMO conversational recall caps at ~40 %. mem0 / HippoRAG hit 80-90 % with neural retrieval at query time — we chose determinism + sub-100 ms p99 instead. Pipe mem0 in front of Sophon if you need that recall.
HashEmbedder is keyword-bound. "favorite food" ↔ "weakness for ginger snaps" doesn't match. Activate BGE (SOPHON_EMBEDDER=bge) for semantic recall — costs +25 MB binary + model load.
No multimodal ingestion. Images / PDFs / audio out of scope. Run Docling / Marker / Unstructured upstream.
Rolling summary doesn't help on small sessions. When the un-summarised tail fits the budget, the rolling cache is a no-op. Useful for long-running sessions with SOPHON_LLM_CMD set.
Some commands don't compress.gh repo view --jsonadds tokens, git log --oneline saves 0.4 %. Sophon's job isn't to compress already-compact output — it's to compress redundant verbose output. The benches name the gaps explicitly.
Project layout
code
.
├── README.md ← you are here
├── BENCHMARK.md ← full per-section benchmark detail
├── CHANGELOG.md ← version history + deprecated numbers
├── benchmarks/ ← reproducible scripts for every number above
├── npm/ ← npm wrapper package
└── sophon/crates/ ← 11-crate Rust workspace
├── prompt-compressor/ compress_prompt
├── memory-manager/ compress_history, update_memory, rolling summary
├── delta-streamer/ read/write_file_delta
├── fragment-cache/ encode/decode_fragments
├── semantic-retriever/ chunker + HashEmbedder + BM25 + entity graph
├── output-compressor/ 21 command-aware filters + JsonStructural
├── codebase-navigator/ tree-sitter / regex + PageRank
├── cli-hooks/ transparent agent installer
└── mcp-integration/ stdio server, async dispatch, cancellation
Contributing
PRs welcome. Run the test suite:
bash
cd sophon && cargo test --workspace --lib --tests --exclude prompt-compressor # 405 testscd sophon && cargo test --features codebase-navigator/tree-sitter # +AST testscd sophon-py && .venv/bin/pytest tests/ # 4 Python tests
Every benchmark claim is reproducible — pointers to the scripts live in BENCHMARK.md. If a number doesn't reproduce on your machine, open an issue.
Particularly welcome:
TypeScript bindings (Python bindings ship in sophon-py/)
gh family filter (gh run list, gh pr list, gh repo view --json) — the bench shows this is currently a gap
SOPHON_EMBEDDER_CMD shell-out plugin pattern (mirror of SOPHON_LLM_CMD) for Voyage / OpenAI / Cohere
Multi-repo real_session_holistic.py runs against popular open-source repos