Index your codebase. AI searches instead of re-reading files. 94% token savings.
io.github.ai-elara/code-context-engine MCP Server
This MCP server indexes a codebase so AI can search for relevant content instead of re-reading files. The provided documentation excerpt claims “94% token savings” with reproducible benchmarking, and positions the server as a code-context engine for AI coding workflows.
🛠️ Key Features
Indexes a codebase for AI search
Intended to reduce token usage (“94% token savings” claim)
Topics indicated include code-indexing and LLM tools
🚀 Use Cases
AI-assisted coding where searching indexed code is preferred to re-reading files
Tool-based coding setups using MCP integrations
⚡ Developer Benefits
Lower token consumption when retrieving context (per the excerpt’s benchmark claim)
Faster context retrieval via search over indexed sources
⚠️ Limitations
No tool count, supported tools, or specific MCP endpoints are included in the provided data
Restart your editor. Done. Every question now hits the index instead of re-reading files.
Agent Plugin support: Run cce init --plugin to generate a portable
Agent Plugin directory that works with
VS Code, Cursor, Copilot, Codex, ChatGPT, and Kiro. The plugin uses
uvx to launch CCE on demand, so users don't need to pre-install the
Python package. See Agent Plugin below.
Already have Ollama? Skip [local] and use uv tool install code-context-engine instead. CCE auto-detects Ollama at localhost:11434 and uses nomic-embed-text.
System requirements
Python 3.11+ and a C compiler (for tree-sitter grammars).
Tested on macOS, Linux, Windows with Python 3.11/3.12/3.13.
cce init auto-detects your editor and writes the right config. To target a
specific agent, use --agent claude, --agent codex, --agent copilot, --agent pi, or
--agent all.
Multiple editors in the same project? All get configured in one command.
Codex note: Codex CLI reads MCP servers from ~/.codex/config.toml only —
it has no per-project config. cce init adds one [mcp_servers.cce-<project>-<hash>]
section per project so multiple projects coexist; cce uninstall removes only
the section for the current project.
Pi note: Pi does not support MCP natively. To use CCE with Pi, you need a
pi MCP adapter extension (e.g. pi-mcp-adapter)
that consumes the .mcp.json config and exposes CCE's tools to the Pi agent.
cce init sets up both .mcp.json and AGENTS.md (Pi loads the latter
automatically for startup instructions).
Supports Anthropic, OpenAI, and Google model pricing. Configure via pricing.model in ~/.cce/config.yaml.
Why this matters
Input tokens are 85-95% of your Claude Code bill. CCE cuts them by 94% (benchmarked on FastAPI).
code
Without CCE: Claude reads payments.py + shipping.py = 45,000 tokens
With CCE: context_search "payment flow" = 800 tokens
Without CCE
With CCE
Session startup
Re-reads files every time
Queries the index
Finding a function
Read entire 800-line file
Get the 40-line function
Cross-session memory
None
Decisions + code areas persisted
Token cost (Sonnet, medium project)
~$0.14/session
~$0.04/session
Benchmark: FastAPI (reproducible)
We benchmarked CCE against FastAPI (53 source files, 180K tokens) with 20 real coding questions. No cherry-picking, no synthetic queries.
Methodology: For each query, "without CCE" means reading the full content of every file the query touches. "With CCE" means the relevant chunks after compression.
Important baseline note: The 94% number is measured against full-file reads, not against what Claude Code actually does. In practice, Claude Code already uses grep, partial file reads, and targeted tools, so the real-world savings compared to normal Claude Code behavior will be lower than 94%. We use full-file as the baseline because it's reproducible and deterministic (no agent behavior variability). The benchmark measures CCE's retrieval efficiency, not a head-to-head comparison with Claude Code's built-in exploration.
Metric
Result
Retrieval savings
94% (83,681 → 4,927 tokens/query)
Compression (additional, on retrieved chunks)
89% (4,927 → 523 tokens/query)
Recall@10 (found the right files)
0.90
Latency p50
0.4ms
Queries tested
20
Per-Layer Savings (each measured independently)
Layer
What it does
Savings
Method
Retrieval
Full files → relevant code chunks
94%
measured
Chunk Compression
Raw chunks → signatures + docstrings
89%
measured
Grammar
Drops articles/fillers from memory text
13%
measured
Output compression (reducing Claude's reply length) provides additional savings (~65% estimated) but is not included in the headline number above.
Walk turn summaries for a session (drill into recall hits)
session_event
Inspect raw tool input/output for a specific event
record_decision
Save a decision for future sessions
record_code_area
Record which files were worked in
index_status
Check index freshness
reindex
Re-index a file or the full project
set_output_compression
Adjust response verbosity (off / lite / standard / max)
Live dashboard with donut charts, file health, and session history:
bash
cce dashboard
Dollar estimates with multi-provider pricing (Anthropic, OpenAI, Google):
bash
cce savings --all # see savings across all projects
How it works
Index: Tree-sitter parses your code into semantic chunks (functions, classes, modules). Stored as vector embeddings locally.
Search: Claude calls context_search. Hybrid vector + BM25 retrieval finds the right chunks. Code graph adds related files automatically.
Compress: Chunks are truncated to signatures + docstrings (or LLM-summarized if Ollama is running).
Remember: Decisions and code areas persist across sessions via session_recall.
Track: Every query is logged. cce savings shows exactly how much you saved.
Re-indexing after edits takes under 1 second (96% embedding cache hit rate). Git hooks keep the index current automatically.
What makes CCE different
It saves where the money is
Output compression tools (like Caveman) save 20-75% on output tokens. Output is 5-15% of your bill. Net savings: ~11%.
CCE saves on input tokens (94% retrieval savings on FastAPI, reproducibly benchmarked). Input is 85-95% of your bill.
It actually understands your code
Not a text search. Tree-sitter AST parsing creates semantic chunks. Hybrid retrieval merges vector similarity with BM25 keyword matching via Reciprocal Rank Fusion. A confidence scorer blends similarity (50%), keyword match (30%), and recency (20%). Graph expansion walks CALLS/IMPORTS edges to pull in related code.
It remembers
record_decision("use JWT for auth", reason="session tokens flagged by legal") is stored in SQLite and surfaces via session_recall in the next session. No re-explaining your architecture.
It tracks real savings
Not estimates. Actual tokens served vs full-file baseline, broken down by buckets (retrieval, compression, output, memory, grammar). Dollar costs fetched from Anthropic's pricing page. Savings summary shown at every session start.
It is secure by default
Secret files (.env, *.pem, credentials.json) are never indexed. Content is scanned for AWS keys, GitHub tokens, Slack tokens, Stripe keys, JWTs, and generic credentials. PII (emails, IPs, SSNs, credit cards) is scrubbed from memory writes. All MCP file paths are validated against path traversal.
Under the hood
Content-Hash Embedding Cache
SHA-256 fingerprint per chunk, salted with model name. Re-index skips unchanged code. Binary float32 storage (10x smaller than JSON). Typical re-index: 96% cache hit, under 1 second.
sqlite-vec: 2 MB instead of 217 MB
Replaced LanceDB with sqlite-vec. Same cosine-distance quality, 99% smaller install. WAL mode + PRAGMA NORMAL for 80% write speedup. Vectors, FTS5, code graph, and compression cache all in three SQLite files.
Deterministic Grammar Compression
Memory entries compressed without LLM calls. Drops articles, fillers, pronouns. Three levels (lite/full/ultra, 20-60% savings). Code, paths, URLs preserved byte-for-byte. Same input always yields same output.
Fail-Closed Hook Design
5 Claude Code lifecycle hooks capture session context. Every hook runs curl ... || true, so a crashed server never blocks the user. SessionStart injects bootstrap context; others capture silently.
Multi-Provider Pricing
Dollar estimates in cce savings support 15+ models across Anthropic, OpenAI, and Google. Static pricing ships with CCE, live Anthropic pricing is fetched and cached 7 days. Configure pricing.model (e.g. gpt-4o, gemini-2.5-pro, sonnet) or override with pricing.input / pricing.output for custom rates.
Resource Governor (Multi-Instance Safety)
Running dozens of cce serve processes (one per project per AI session) can exhaust system memory. The resource governor caps ONNX Runtime threads per process, uses advisory file locks so only one process indexes a given project at a time, backs off under Linux memory pressure (PSI), and auto-shuts down idle servers after 30 minutes. Configure via serve.idle_timeout_minutes and serve.max_ort_threads.
Memory Nudges
CCE's cross-session memory depends on the agent calling record_decision and record_code_area. Memory nudges make recording ambient: after N searches without a recording, context_search results include a short reminder. At session end, the Stop hook summarizes unrecorded activity. Nudges re-arm after the first recording so they stay useful without being noisy.
HTTP Search Endpoint
cce serve --http exposes a POST /search endpoint for custom agent integrations that speak HTTP instead of MCP stdio. Same hybrid retrieval pipeline, structured JSON response with confidence scores. Input validation clamps top_k (1..100) and confidence_threshold (0.0..1.0).
Agent Plugins is an open standard (v1.0.0) backed by Amazon, Cursor, Microsoft, OpenAI, and Vercel for packaging AI skills and MCP servers into portable, zero-install bundles. CCE can generate a plugin directory that compatible editors can discover and load automatically.
VS Code, GitHub Copilot, ChatGPT, Codex, Cursor, and Kiro. The plugin uses uvx to launch CCE on demand, so users do not need to pre-install the Python package. The MCP server auto-discovers the project root by walking up from its working directory, looking for .context-engine.yaml or .git/.
When to use --plugin vs --agent
--agent (default)
--plugin
Install method
Writes editor-specific config files
Generates a portable plugin directory
Zero-install
No, CCE must be on PATH
Yes, uvx fetches CCE on demand
Instruction updates
Stale until cce init re-run
Stale until cce init --plugin re-run
Best for
Your own machine
Sharing with a team or distributing
Both can be used together. --agent handles per-editor MCP config, --plugin provides a portable alternative.
CLI at a glance
bash
cce init # Index + install hooks + register MCP
cce init --plugin # Generate Agent Plugin for VS Code, Cursor, etc.
cce # Status banner
cce savings # Token savings with dollar estimates
cce savings --all # All projects
cce dashboard # Web dashboard with live charts
cce search "auth flow"# Test a query
cce status # Index health + config
cce services # Ollama + dashboard + MCP status
cce commands add-rule '...'# Project rules for Claude
cce uninstall # Clean removal of all CCE artifacts
Run cce list for the full command reference.
Configuration
Zero-config by default. Override what you need in ~/.cce/config.yaml or .context-engine.yaml:
yaml
compression:level:standard# minimal | standard | fulloutput:standard# off | lite | standard | maxollama_url:http://localhost:11434# point at a remote Ollama if desiredretrieval:top_k:20confidence_threshold:0.5pricing:model:opus# opus | sonnet | haiku | gpt-4o | gemini-2.5-pro | ...# input: 15.0 # override $/1M input tokens# output: 75.0 # override $/1M output tokens
Remote Ollama: If you run Ollama on another machine in your network, set compression.ollama_url (e.g. http://nas.local:11434) or export CCE_OLLAMA_URL (the env var wins). CCE probes the endpoint and falls back to truncation-only compression when it's unreachable, so a flaky link won't break indexing.
Output Compression
CCE also compresses Claude's responses (same concept as Caveman):
Level
Style
Savings
off
Full output
0%
lite
No filler or hedging
~30%
standard
Fragments, drop articles
~65%
max
Telegraphic
~75%
Tell Claude: "switch to max compression" or "turn off compression". Code blocks and commands are never compressed.
Disk Footprint
Component
Size
Core install (Ollama backend)
~17 MB
With [local] extra (fastembed + ONNX)
~189 MB
Embedding model (one-time download)
~60 MB (fastembed) or managed by Ollama
Index per project (small/medium/large)
5-60 MB
No GPU required. With Ollama, embeddings are handled by the Ollama server. With the [local] extra, the embedding model runs on CPU via ONNX Runtime.
CCE replaces "dump the entire file" with "search for the relevant function." The model still gets the code it needs (0.90 Recall@10 in benchmarks). Less irrelevant context means less noise competing for attention, which can improve the model's focus on your actual question.
How does output token savings work?
CCE writes output compression rules directly into your agent's instruction files (CLAUDE.md, AGENTS.md, .cursorrules, etc.) during cce init. These rules apply to the entire session, not just CCE tool responses, so every reply from the agent follows them.
Set the level in ~/.cce/config.yaml or .context-engine.yaml:
yaml
compression:output:max# off | lite | standard | max
Then re-run cce init to update instruction files. Or change at runtime:
code
set_output_level output_level=max
Level
Savings
What it does
off
0%
No compression
lite
~25%
Removes filler/hedging/pleasantries + diff-only for code changes
standard
~70%
Drops articles, fragments, short synonyms + diff-only for code
max
~80%
Telegraphic style + diff-only for code
Default is standard. All levels include code output rules that tell the model to show only changed lines (not full file rewrites), which is where most output tokens go in coding sessions. The max level produces very terse prose (similar to "caveman mode"). Code blocks, paths, and commands are never compressed regardless of level.
Where do the savings come from?
Most savings are input tokens (what goes into the model):
Layer
Type
Typical savings
Retrieval
Input
94% (full files → relevant chunks)
Chunk compression
Input
89% (chunks → signatures)
Grammar compression
Input
13% (article/filler removal)
Turn summarization
Input
varies (session history)
Progressive disclosure
Input
varies (tool payloads)
Output compression
Output
25-80% (depends on level)
Output tokens cost 5x more per token (e.g. Opus: $15/1M input vs $75/1M output), so even a small output reduction has outsized cost impact.
Roadmap
Multi-repo benchmarks (FastAPI, chi, fiber)
More benchmarks (Django, Express)
Tree-sitter support for C, C++, Ruby, Swift, Kotlin