Deterministic code map for AI coding agents: a tree-sitter call graph with exact file:line over MCP.
io.github.getdomovoi/osnova — Model Context Protocol (MCP) Server
This MCP server provides a deterministic code map for AI coding agents. It exposes a tree-sitter call graph that includes exact file and line locations via MCP, enabling structured navigation of code relationships with file/line precision.
🛠️ Key Features
Deterministic code map for AI coding agents
Tree-sitter call graph
Exact file:line mapping exposed over MCP
🚀 Use Cases
Finding where functions call other functions using a call graph
Navigating code paths with file and line references
Supplying AI coding agents with structured, location-accurate context
⚡ Developer Benefits
More precise context through exact file:line references
Deterministic output via a deterministic code map approach
Developer-friendly graph structure derived from tree-sitter
⚠️ Limitations
Only described capabilities are deterministic call-graph mapping with exact file:line data; additional tools, topics, and usage details are not provided.
A deterministic code map for AI coding agents. Osnova indexes a repository into a symbol and call graph with tree-sitter, then serves it to any MCP client or from the command line. Same input, same output, byte for byte. No embeddings, no telemetry, and no network connection unless you type osnova update-check. Twenty languages, seven of them (TypeScript, JavaScript, Python, Go, Rust, Java, C#) with deep adapters.
After npm install -g @getdomovoi/osnova, one command per kind of agent: claude wires Claude Code (MCP entry, session, prompt and stop hooks, and the skill); agents wires every installed harness that reads AGENTS.md (Codex, OpenCode, Kilo, Pi, Cursor), each with its hooks, plugin or extension, plus one shared skill in ~/.agents/skills/. Each previews its changes; --apply writes them, backing up every file it changes.
sh
osnova setup claude --apply
osnova setup agents --apply
Or add the MCP entry by hand; replace osnova with npx -y @getdomovoi/osnova when there is no global install.
A terminal recording: osnova warp lists the five resolved callers of click Context.invoke, each with its receiver hint; grep finds twelve .invoke( lines; osnova plumb checks those twelve and reports five confirmed, six name-only matches on other invoke methods and one line inside a docstring.
osnova warp src/api.ts#refreshWorkspace on this repository, cut to twelve lines (test/docs-readme-warp.test.ts fails when the first two no longer reproduce):
text
function src/api.ts#refreshWorkspace: 83 indexed edges
reach: d1 callers 83 in 14 files (3 dirs); d2 +16 in 2 files; unresolved same-name 10; tests 18
d1 calls src/cli/cli.ts#ensureIndex:78 [import-binding]
d1 calls src/cli/hook.ts#runHook:191,203,225,240 [import-binding]
d1 calls src/mcp/server.ts#createOsnovaMcpServer.refresh:297 [import-binding]
d1 calls test/verification-fastpath.test.ts#<module>:37,47,48,51,53,57,58,64,68,74,78,85,110,114,132,145,154,166,171,197 [re-export-binding]
via src/index.ts:65 refreshWorkspace -> src/api.ts (export refreshWorkspace)
This does not prove absence of callers or that deletion is safe.
unresolved evidence (10); not confirmed relationships
candidates for refreshWorkspace (2, unverified): src/api.ts#refreshWorkspace, test/mcp-watch.test.ts#refreshWorkspace
d1 calls refreshWorkspace test/artifact-format9.test.ts:78,85,104,120,140 [binding-blocked]
d1 calls refreshWorkspace scripts/perf.mjs:485,486,491 [import-target-unresolved]
Each confirmed line names the caller, its lines and the evidence that tied the call to this definition: an import binding, a same-file definition, a re-export chain with its hop, or an identified receiver. Calls the index could not tie to a definition are listed apart as unresolved evidence, with the same-name candidates it found and the reason it stopped, so a name match is never mistaken for a caller. The reach line gives exact counts, not scores, and when the list outgrows its budget a capped: line and an omitted: footer count what was left out.
How much of the graph is exact
A call site counts as resolved when the index ties it to one definition through evidence it can name; everything else stays unresolved with a reason. These are the shares on the pinned checkouts under benchmarks/corpora/, recorded in benchmarks/results/resolution-coverage-2026-09-21b.json; the last column leaves out calls through packages outside the repository and calls to builtins, which can never resolve locally, and the reference defines every column.
Corpus
Languages
Call sites
Resolved
Share
Excluding externals
click
all
6593
2903
44.0%
62.4%
cobra
all
4374
1980
45.3%
88.9%
gson
all
23382
8473
36.2%
56.7%
humanizer
all
30517
7798
25.6%
51.3%
pyright
all
58578
30070
51.3%
74.5%
ripgrep
all
13387
6121
45.7%
71.6%
zod
all
53417
21630
40.5%
73.2%
Per-language rows are in the reference. osnova coverage reports the same numbers for your own repository, per language and per reason.
Resolved is not the same as right, so the call edges are also scored against each language's type checker or compiler. Every call site in six pinned checkouts was sent to the checker for the callee's declarations, and each osnova edge was marked true when the definition it names contains that declaration and false when it does not. Measured at 0.12.0:
Corpus
Checker
Decided edges
False
False-edge rate
In-repo sites covered
click
pyright 1.1.414
2883
0
0%
90.4%
zod
TypeScript 5.9.3
21593
4
0.02%
78.2%
cobra
go/types, Go 1.27.1
1980
0
0%
87.5%
ripgrep
rust-analyzer 1.98.1
6066
29
0.48%
85.8%
humanizer
Roslyn 5.9.0
7363
3
0.04%
62.2%
gson
javac 27
8133
4
0.05%
72.3%
Java and C# overloads share one symbol in the index, so a call edge names the overload it binds only when the written argument count fits exactly one declaration of the type, its partial parts and its declared base classes; when it fits several or none, the edge keeps the method and names no declaration, and the oracle leaves it out of both rates (2526 edges on gson, 602 on humanizer). 0.11.0 named the last overload declared for every call, recorded a call at the first line of its call expression, and lost the rest of a C# file after a primary constructor or a raw string; it measured gson 2045 false edges (24.75%), humanizer 617 (8.25%) and ripgrep 112 (2.06%). Of the 29 left on ripgrep, 11 are functions declared once per #[cfg] branch, 6 are a test module's function named like its parent's, and 5 are calls to a local closure or parameter named like a function elsewhere. A call whose target has one declaration that fits keeps it even when a base type is outside the index, where an unseen base overload could be the one the compiler binds; this limit is kept because withdrawing those edges cost 10 recall points on humanizer for one fewer false edge. Every false edge is classified by cause in benchmarks/results/type-checker-oracle-2026-09-30.json with the method and its limits. The scripts that produce these numbers are in benchmarks/oracle/.
Grep versus the graph
The reason to keep a call graph instead of running a text search is not speed. It is that the first regex a person types is wrong more often than it looks, and nobody notices. Nine call-site sets in six languages were verified line by line; each cell shows sites found (precision / recall), and the method lists what each side got wrong.
The names play on the foundation image. The CLI uses the same names without the prefix: osnova ground and osnova_ground are the same query.
Tool
Meaning
Does
osnova_ground
the ground you stand on
Keyword search: ranked definitions with exact file:line
osnova_thread
the thread you follow through the cloth
Text search: regex or literal matches grouped by symbol
osnova_outline
the outline of one part
Signatures and line spans for one file
osnova_warp
the threads that hold the weave
Call graph: callers and callees, direct or transitive
osnova_groundwork
the groundwork under everything
Repository map: directory clusters, hubs and hotspots
osnova_footing
the footing you build on
Task context: definitions, relationships and candidate tests around a question
osnova_settle
how the ground settles after a change
Change impact: the symbols a diff touches and their indexed dependents
osnova_plumb
the plumb line dropped straight through
Check claims: which listed call sites the index confirms, and which dependents were left out
osnova_tests
the cloth pulled to see what holds
Tests: the test files that reference a symbol, or the symbols one test file reaches
osnova_unreferenced
threads left loose at the edge
Definitions with no indexed caller, each with its leads; candidates, never proof
Every MCP response opens with its index generation, says when the index is partial, and counts what its budget left out; the budgets, and which CLI subcommands carry the generation, are in the reference.
Hooks and clients
One global install serves every repository and every client: one entry in each client's global config, nothing per project, nothing written inside your repository. osnova setup claude and osnova setup agents show the diff and write nothing until --apply; agents skips a harness whose config folder is missing, --only codex,pi narrows it, and a skill or plugin file you edited is kept and reported. --uninstall reverses either family the same way, previewing first. The table lists the per-client form, for one piece at a time. What each hook prints, when it stays quiet, and what the trials measured are in the reference.
Client
How to wire
What it adds
Claude Code
osnova setup --apply --client claude-code --hooks
MCP entry plus session, prompt and stop hooks; --skill adds the skill, --nudge the opt-in grep nudge
Claude Code, as a plugin
/plugin marketplace add getdomovoi/osnova then /plugin install osnova@osnova
The same MCP entry, hooks and skill, run through npx -y @getdomovoi/osnova@<version>, pinned to the plugin's release, with no global install; updates follow the marketplace
Codex
osnova setup --apply --client codex --hooks
MCP entry plus the same three hooks; trust them in /hooks or Codex skips them silently; step by step in Osnova with Codex
Cursor
osnova setup --apply --client cursor --hooks
MCP entry plus the stop hook as a follow-up message
OpenCode
osnova setup --apply --client opencode --plugin
MCP entry plus a plugin that appends starting points to each message; the tool guidance comes from the MCP instructions
Kilo
osnova setup --apply --client kilo --plugin
MCP entry plus the same plugin
Pi
osnova setup --apply --client pi --plugin
MCP entry through pi-mcp-adapter plus an extension that does the same
In CI
The action runs osnova settle --base-ref against the pull request base and lists every indexed dependent of the symbols the pull request changed; the report is indexed structural evidence only, so a missing dependent is not proof that nothing depends on the change.
No type inference and no dynamic dispatch. Edges come from syntax: direct calls, imports, name references, declared heritage (a written superclass or interface name) and framework routes (a registration whose receiver binds to a listed framework import, with the verb and the literal path written at the site), with lexical binding and receiver hints for TypeScript, JavaScript and Python. Resolution is heuristic and says so.
No route table. A routes edge is one registration site tied to one handler; prefixes from mounts, blueprints and controllers are recorded on their own edges and never composed into a full path, a computed path records no path, and an inline closure or a wrapped handler records the route with no target. Express, NestJS, Flask and FastAPI are read; gin, axum, Django, Spring, ASP.NET, Rails and file-based routers are not.
That boundary has a measured price. Scored against a type checker on two pinned corpora, the calls osnova does not resolve are mostly calls whose receiver type is never written down: 2945 of 6359 missed sites on zod and 164 of 307 on click are a plain name carrying no annotation, and another 1478 on zod are a call result or a property chain. Only 349 missed sites on zod and 10 on click have a type written at the receiver's declaration, and 248 of those 349 are a single library idiom. Resolving every one of them would move recall from 76.9% to 78.2% on zod and from 90.4% to 90.7% on click, so the boundary costs roughly one recall point rather than ten. The full census is in benchmarks/results/receiver-boundary-census-2026-09-21.json.
No semantic search. osnova_ground is fielded lexical ranking over definitions. It is fast, deterministic and explainable, and it will not match a paraphrase.
No proof of safety. An empty caller list means the index found no caller, not that none exists.
No cost claims. Agent trials so far show correctness parity with and without the graph on small tasks. A benchmark that separates the two is in progress.
How it stays honest
Every query refreshes the index from the working tree first, uncommitted edits included, and reports its generation.
Every clipped output says how much was clipped. Every ranked list says how many candidates it dropped. Partial indexes say so on every response.
Every hit carries a file:line span, a source hash and an index generation, so you can check what the agent cites.
Every benchmark result in benchmarks/results/ is frozen with its corpus fingerprint, and rejected experiments stay on record next to accepted ones.
The cache verifies a SHA-256 over the structural core before parsing it and hash-checks each source text on read. Nothing is written outside it.
CLI
Every tool is a subcommand: osnova build <root>, osnova warp <symbol>, osnova settle --base-ref <ref>, plus coverage, check, doctor, setup and mcp. The full list is in the reference.
Library
import { buildIndex, ask, callersDetailed, renderMapCard } from "@getdomovoi/osnova". The structured APIs return complete results with omission counts; the text budgets apply to CLI and MCP presentation only. Every export is in the reference.
Contributing
Bug reports, language adapters, benchmark corpora and agent trials are all welcome. Read CONTRIBUTING.md for the gates a change must pass: lint, typecheck, tests, build, perf budgets and package smoke.