The open retrieval layer for AI agents — index code, docs, data. Search via MCP.
io.github.roomi-fields/rtfm (MCP Server)
The io.github.roomi-fields/rtfm MCP server provides an open retrieval layer for AI agents that indexes project materials and makes them searchable via MCP. It targets code, docs, PDFs, legal texts, research, and data to help an agent retrieve relevant context.
🛠️ Key Features
Indexes “everything in your project” (code, docs, PDFs, legal texts, research, data)
Search via MCP
Retrieval focused on context (topics include retrieval, RAG, code search, semantic search)
Uses supporting infrastructure topics such as embeddings, FTS, sqlite, and json-schema
🚀 Use Cases
Code and documentation search for agents
Retrieval-augmented generation workflows using project context
Local knowledge-base style access for agent grounding
⚡ Developer Benefits
Developer-oriented retrieval capabilities with MCP integration
Project context organization implied by topics like knowledge-base and knowledge-graph
⚠️ Limitations
The provided description does not specify tool count or exact indexing/search configuration details.
The open retrieval layer your AI agent was missing
Index everything in your project — code, docs, PDFs, legal texts, research, data — and your agent finds the right context instantly. No hallucinations. No cloud. No API costs.
It greps through thousands of files, misses the doc that answers the question, invents modules that don't exist, forgets what you decided last session. The bigger the project, the worse it gets. You've added a smarter model. It didn't help. Because the bottleneck isn't intelligence — it's retrieval.
Code indexers (Augment, Sourcegraph, Cursor) only see code. But your project isn't just code. It's specs, PRs, architecture decisions, research papers, PDFs, regulations, vault notes — the context your agent needs to stop guessing.
Why I built this
I was writing a French tax article (~50 pages of regulatory text, cross-references between code articles, case law, administrative doctrine). Claude Code kept grep-ing the same directories in loops, running out of context, and producing confidently wrong citations. I'd added more memory, better prompts, a smarter model. None of it worked, because the agent wasn't reasoning badly — it just couldn't find the right paragraph in a 2,000-file legal corpus. So I stopped trying to make the model smarter and built the layer it was missing. That's RTFM.
The solution
RTFM indexes everything. One command, one SQLite file, one retrieval layer your agent queries before grepping.
bash
pip install rtfm-ai && cd your-project && rtfm init
30 seconds. Claude Code now searches your indexed knowledge base — code and docs and PDFs and whatever else you drop in — with full-text, semantic, or hybrid search. The agent sees 300 tokens of metadata first, then expands only what's relevant. Progressive disclosure instead of context dumps.
Free. Runs locally. No API keys. No cloud. Your data stays yours.
That's it. The plugin auto-initializes each project on first use:
Creates .rtfm/library.db (one SQLite file)
Injects search instructions into CLAUDE.md
Pre-grants permission for the MCP tools (no prompt every search)
Indexes the project on the first prompt, re-indexes incrementally on every prompt
No pip install required. Pure Python, runs on Linux / macOS / Windows / WSL with Python 3.10+ already on PATH. The plugin bundles its own MCP server (no mcp SDK dep) and resolves python3 / python / py automatically.
Then say to Claude: "Find the authentication flow" — it uses rtfm_search instead of grepping.
Optional extras (semantic search, PDF parsing)
The core plugin is dependency-free. Heavier optional extras (embedding model, PDF parsers) install on demand into an isolated venv inside the plugin's data directory — no pollution of your system Python, no PEP 668 conflicts:
code
/rtfm:install-embeddings # FastEmbed ONNX (~85 MB), semantic + hybrid search
/rtfm:install-pdf # pdftext only (~50 MB), fast text extraction
/rtfm:install-pdf-full # + marker-pdf + CPU-only torch (~1.5 GB), complex layouts
The pdf-full install uses PyTorch's CPU-only index (no CUDA, no GPU needed) to stay around 1.5 GB instead of 5 GB.
Restart Claude Code after install for the extras to be picked up.
Manual install (Cursor, Codex, Claude Desktop chat, other MCP clients)
For clients without Claude Code's plugin system :
bash
pip install rtfm-ai
cd /path/to/your-project
rtfm init
Then point your MCP client at rtfm-serve (the entry exposed by the pip package). Optional extras via pip install rtfm-ai[embeddings,pdf].
How it compares
RTFM
Augment CE
Sourcegraph
Code-Index-MCP
MemPalace
Code indexing
✅ (AST-aware)
✅
✅
✅
Shallow (char-chunk)
Docs, specs, markdown
✅ (header-parsed)
Partial
❌
Limited
Verbatim chunks
Legal / regulatory
✅ (XML, BOFiP)
❌
❌
❌
❌
Research (LaTeX, PDF)
✅
❌
❌
❌
❌
Custom parsers
✅ (~50 lines)
❌
❌
❌
❌
Knowledge graph
✅ (file/code links)
❌
Partial
❌
Entity graph (people)
File version history
✅ (unlimited)
❌
❌
❌
❌ (purge-and-replace)
MCP native
✅
✅
✅
✅
✅
Runs locally
✅
Cloud
Enterprise
✅
✅
Open source
MIT
❌
Partial
✅
MIT
Price
Free
$20-200/mo
$$$/mo
Free
Free
RTFM is the only open-source option that indexes multi-domain content with structural parsing, a code-level knowledge graph, and unlimited per-file history. That's the niche.
Different from MemPalace specifically: MemPalace is an entity-level memory for conversations (who/project/decision triples in SQLite, plus verbatim chunks in ChromaDB). RTFM is a retrieval layer for artefacts — parsed by format, linked at the file level, versioned over time. The two are stackable, not competing.
For a deeper breakdown of the design choices behind any RAG (chunking, retrieval, augmentation, integration, freshness, storage), see RAG Fundamentals — the 6 axes →
Memory that survives sessions
Between sessions, most agents forget. RTFM indexes Claude Code's own memory files across every project on your machine, with full version history.
bash
rtfm memory # Manual snapshot
rtfm memory --install-hook # Auto-snapshot on every SessionEnd
Cross-project index — one DB at ~/.rtfm/memory.db sees every ~/.claude/projects/*/memory/ directory on your machine. Ask rtfm_search("OAuth auth decisions") and get hits from all 18 of your projects.
Unlimited version history — every change to a memory file is snapshotted (no prune). rtfm_history <slug> returns the full evolution.
Auto-snapshot on SessionEnd — one command installs a global Claude Code hook. Every session you close captures a new snapshot.
Curated, not verbatim — RTFM indexes the notes the agent already curated itself during the session (small, structured, signal-dense). Different philosophy from MemPalace, which indexes the full conversation transcripts in ChromaDB (large, noisy, needs aggressive semantic filtering).
Obsidian vault mode
RTFM is the retrieval layer for the Karpathy LLM Wiki pattern. Karpathy himself wrote: "at small scale the index file is enough, but as the wiki grows you want proper search." This is proper search.
bash
cd /path/to/your-obsidian-vault
rtfm vault
Detects .obsidian/, proposes a folder → corpus mapping
Resolves [[wikilinks]] following Obsidian rules → stored as graph edges
Generates _rtfm/ with Obsidian-native navigation (index, graph with Mermaid, hubs, orphans, Dataview frontmatter)
RTFM pairs naturally with notebooklm-mcp. NotebookLM caps you at 50 queries/day per notebook; RTFM removes that ceiling by indexing answers locally — ask once, retrieve forever, offline, in milliseconds.
notebooklm-mcp's /batch-to-vault endpoint writes citation-backed Q&A as {slug}.md (markdown with frontmatter) plus {slug}.json (structured nblm-answer-v1 sidecar). Both are guaranteed to coexist. Two integration paths, both ship today:
Path A — Markdown (zero config): drop the vault into RTFM and rtfm sync. The default markdown parser slices each answer into question / answer / per-citation chunks automatically. No mapping, no schema, no code.
Path B — JSON sidecar (typed metadata): drop a nblm-answer.yaml mapping into .rtfm/mappings/. Each .json answer file produces typed chunks with notebook_id, source_name, citation_marker queryable via SQL, plus cites edge candidates between answers and sources.
Use Path A unless you specifically need to filter or graph by structured citation fields.
I ran two kinds of benchmarks. The honest picture is nuanced — retrieval helps most on tasks that are actually solvable and where the agent is spending time looking for things.
Document-heavy task: French tax article generation (B10)
Writing a ~50-page regulated article from a corpus of legal code, case law, and administrative doctrine. Same agent (Claude Code + Sonnet 4), same prompt, eight configurations tested.
Configuration
Duration
Cost
Tokens
Baseline (no RTFM)
8m 16s
$22.61
8.21 M
With RTFM (FTS default)
6m 58s
$11.14
3.22 M
Δ : −51 % cost, −61 % tokens, −16 % duration — with better factual accuracy.
This is the use case RTFM was built for: navigating a large multi-domain corpus where grep misses the right paragraph.
11 tasks, 3 repos of varying size, 4 conditions (A = standard prompt with file paths; B = discovery, no paths; C = RTFM FTS; D = RTFM hybrid), 3 runs each.
Repo
Size
Where RTFM helps
metaflow
620 files
Everyone resolves — RTFM adds no measurable gain
astropy
1,119 files
All conditions 25–30 % F2P pass; none fully resolve
mlflow
8,255 files
All conditions 0–5 % F2P pass; none fully resolve
On a single smaller-scope run (test_stub_generator on metaflow), RTFM cut agent time by −37 % vs the no-paths baseline. On the larger repos, the tasks themselves were too hard for Sonnet 4 to resolve inside a 20-minute timeout regardless of retrieval.
The honest caveats
Single model (Sonnet 4), single agent (Claude Code). Not statistically bullet-proof.
On small repos (< 1k files), grep is enough and RTFM adds overhead.
FeatureBench measures code modification, not information retrieval. It's the wrong benchmark for a retrieval tool — I'm running against it because it's what exists. Better-suited benchmarks (RepoQA, SWE-QA, LocAgent) are on the roadmap.
What this says
RTFM measurably wins when the bottleneck is "find the right paragraph in a 2,000-file corpus". It doesn't magically make unsolvable tasks solvable. The model still has to do the work — RTFM just makes sure it has the right context to do it with.
Who it's for
RTFM works anywhere your project isn't just code:
LegalTech — Code + tax law + regulatory specs. Ships with Legifrance XML and BOFiP parsers.
Research — Code + LaTeX papers + datasets. Ships with LaTeX and PDF parsers.
FinTech — Code + financial regulations + XBRL reports. Write an XBRL parser in 50 lines.
HealthTech — Code + medical records (HL7/FHIR) + clinical guidelines.
Solo devs with big projects — Stop watching your agent grep the same 8,000 files every session.
Obsidian / PKM users — Make your vault actually searchable by your AI.
Any regulated industry — If your project mixes code with domain documents, RTFM is for you.
Full feature list
Search & retrieval
FTS5 full-text search — instant, zero-config, works out of the box
Semantic search — optional embeddings (FastEmbed/ONNX, no GPU needed)
Hybrid mode — combine both, rank by relevance score
@AVeryTastyRaspberry made RTFM
run on native Windows. RTFM is developed on Linux, and every command was
broken there — the CLI died at import time before it could parse an argument.
The report (#8) named the
line; the testing that followed, on a real Windows 11 machine and checked
against tasklist rather than against RTFM's own claims, found five more
defects behind it and confirmed each fix. #9
then traced the console windows that kept popping up. That is a platform this
project could not otherwise support.