Pre-computed atlas of your codebase for Claude Code: LSP + ADRs + git. 45-72% token reduction.
ContextAtlas (io.github.traviswye/contextatlas) MCP Server
This MCP server provides a pre-computed atlas of a codebase for Claude Code. It packages LSP-grade structure plus architectural intent from Architectural Decision Records (ADRs), integrates git history, and includes test associations to build a single-call context bundle for architectural prompts.
π οΈ Key Features
Pre-computed codebase βatlasβ for Claude Code
LSP-grade structure
ADRs for architectural intent
Git history integration
Test associations included
Reported 45β72% token reduction (with stated zero quality regression on benchmark axes)
π Use Cases
Supplying architectural context to Claude Code in one call
Reducing token usage for codebase context retrieval
Supporting prompts across architectural tasks (benchmark suite referenced)
Uses structured sources: LSP data, ADRs, git history, and test links
β οΈ Limitations
Described performance and βzero quality regressionβ are based on the referenced benchmark suite (hono / httpx / cobra); no broader coverage is stated.
Stop watching Claude burn tokens grepping for context it can't possibly find.
ContextAtlas turns your codebase into a single-call context bundle for Claude Code β
fusing LSP-grade structure, architectural intent from your Architectural Decision Records (ADRs), git history, and test
associations. Measured 45-72% token reduction with zero quality regression across
benchmark axes on architectural prompts across the hono / httpx / cobra benchmark suite.
ContextAtlas ships two equivalent paths β CLI and Claude Code Skills β
both producing the same atlas.json. See Quick Start
for setup.
The Problem
Claude Code currently learns your codebase by brute force. Every session
starts fresh. Every "where is X?" triggers multiple grep calls. Every
"what depends on Y?" is another flurry of file reads. On a mid-sized
codebase, answering a single architectural question can consume 40+ tool
calls and 100,000+ tokens before Claude has enough context to reason
well.
Worse: the architectural intent that governs your code β the ADRs, the
design decisions, the "we did it this way because" β is invisible to
Claude. The rule that OrderProcessor must be idempotent lives in
docs/adr/. When Claude proposes a change, it has no way to know that
constraint exists.
Yesterday's understanding doesn't carry to today. Every conversation
starts from zero. Your ADRs, your commit history, your test coverage β
none of it is on the agent's table.
What if expensive understanding happened once, at index time, and
every query became a dictionary lookup?
That's ContextAtlas.
What ContextAtlas Is
ContextAtlas is an MCP server that gives Claude Code a curated atlas of
your codebase β fusing LSP-grade structural precision with architectural
intent extracted from your ADRs, docs, and git history, delivered to
Claude in single-call context bundles.
Every bundle Claude receives combines four independent signals about a
symbol:
Structural data from the language server β definition, references,
types, diagnostics. Compiler-grade precision.
Architectural intent from your ADRs, READMEs, and design docs β
structured claims extracted by Opus 4.7 at index time, keyed to
specific code symbols.
Historical context from git β recent commits touching the symbol,
hot/cold indicators, co-change patterns.
Test associations β which tests reference the symbol, where
coverage lives.
One MCP call returns all four, fused. No ADRs in your repo yet? You
still get LSP + git + tests in one call instead of fifteen β a
meaningful baseline improvement. Add ADRs and the bundles get richer.
The architecture is designed so any subset of signals produces value.
Given an ADR stating that OrderProcessor must be idempotent, a call
to get_symbol_context("OrderProcessor") returns:
code
SYM OrderProcessor@src/orders/processor.ts:42 class
SIG class OrderProcessor extends BaseProcessor<Order>
INTENT ADR-07 hard "must be idempotent"
RATIONALE "All order processing must be safely retryable."
INTENT ADR-12 soft "prefer async base class for new processors"
REFS 23 [billing:14 admin:9]
TOP ref:ts:src/billing/charges.ts:88
TOP ref:ts:src/admin/orders.ts:12
GIT hot last=2026-03-14
RECENT "Fix idempotency bug in retry path" a3f2c1d
TESTS src/orders/processor.test.ts (+11)
When Claude is asked to modify OrderProcessor, it sees the
idempotency constraint before proposing changes β not after a user
review catches the violation.
Who this is for. ContextAtlas is built for the average developer
using Claude Code on real codebases β not just engineers at large orgs
working on 500,000-file monorepos. Token-burn reduction scales with
codebase size β dramatic on a 200-file framework, modest on a 30-file
library. But architectural intent capture is size-invariant. A
30-file library can have meaningful architectural decisions worth
surfacing, and Claude respecting them matters just as much as on a
larger codebase.
Beyond tokens: a design-alignment case study
Efficiency and quality are necessary but not sufficient. The
substantive value of context-grounding shows up in design
choices on non-trivial code-change tasks.
A/B trial during v0.3 development. Identical 3-paragraph
prompt across two ContextAtlas clones: implement a known bug fix β
locate the bug, design and implement the fix, write tests,
document via ADR. The only setup difference: MCP availability.
Arm
MCP
Approach selected
A (vanilla)
none
Recall-first approach (broader matching; fought the project's precision-thesis with noise)
B (CA-aided)
ContextAtlas
Precision-optimization approach (aligned with the project's pre-extracted-claims-with-structural-attribution thesis)
Arm B's approach landed in main. Both arms functionally fixed
the bug at similar wall-clock and token cost. The substantive
difference was alignment with project design thesis β the
CA-aided arm could read the relevant ADR + prior architectural
work from the atlas, and made a choice that fit. The vanilla arm
couldn't see that context and chose an approach that worked but
fought the architecture.
Arm A's substantive consideration wasn't lost β captured as
future-work investigation trigger. The recall-vs-precision
tradeoff is preserved.
N=1 trial; this is anecdote, not benchmark. The systematic
benchmark suite (hono / httpx / cobra) measures efficiency and
quality (see Β§The Numbers below). This A/B trial measures the
substantively-distinct design-alignment axis β which doesn't fit
benchmark-suite methodology (every code-change task is repo-
specific) but is the substantive value proposition for cohort
developers building on real codebases.
The Numbers
We benchmark ContextAtlas against baseline Claude Code on three
repositories chosen to reflect realistic developer workloads:
Repo
Language
Source files
Role
honojs/hono
TypeScript
186
Mid-sized framework
encode/httpx
Python
23
Focused production library
spf13/cobra
Go
19
CLI framework
Methodology. 24 prompts per repo, 6 task buckets, blind manual
grading, pre-registered rubric, no cherry-picking. Full methodology in
RUBRIC.md.
Efficiency: 50-71% tool-call reduction on architectural prompts
Phase 5 reference run on hono, six pre-registered prompts:
Prompt
Bucket
Alpha calls
CA calls
Ξ
Alpha $
CA $
h1-context-runtime
win
18
9
β50%
$2.36
$1.52
h2-router-contract
win
11
5
β55%
$0.60
$0.53
h3-middleware-onion
win
5
5
0%
$0.38
$0.47
h4-validator-typeflow
win
21
6
β71%
$2.95
$0.52
h5-hono-generics
tie
11
13
+18%
$0.79
$1.17
h6-fetch-signature
trick
3
4
+33%
$0.17
$0.29
aggregate
69
42
β39%
$7.25
$4.50 (β38%)
The headline case: h4-validator-typeflow ran 7.3Γ cheaper ($2.95 β $0.52)
at equivalent answer depth. CA opens with the governing ADR by number;
the baseline reconstructs the architecture from source. Tie and trick
buckets (h5, h6) show CA net-negative as the rubric predicted β CA
over-engineers on questions where architectural intent doesn't carry
load. Bucket-aware methodology surfaces these expected cases rather
than burying them.
Cross-language replication: the same architectural-intent win mechanism
holds on Python (Phase 6 β httpx)
and Go (Phase 7 β cobra).
Phase 8 re-ran the locked prompt sets against v0.3-sharpened atlases at
the same pinned target SHAs: 45-72% token reduction on architectural-intent
prompts across all three target languages. Full synthesis at
phase-8-v0.3-reference-run.md.
Quality: blind-graded, paired-t with confidence intervals
v0.5 shipped the LLM-judge methodology under paired-mode anonymization
(per ADR-19). Cross-cell
rollup paired-t at N=27 differences per axis (5 anchor cells Γ n=5
trials Γ 2 conditions; hono h1 auto-stretch to n=8):
Quality axis
Mean Ξ (0-3 scale)
95% CI
Tier
Factual correctness
+0.370
[0.176, 0.565]
CLEAN
Hallucination
+0.296
[0.032, 0.561]
Borderline
Actionability
+0.148
[0.005, 0.291]
Borderline
Completeness
+0.037
[-0.039, 0.113]
Not distinguishable
Threshold pre-registration: the three-tier framing (β₯0.05 CLEAN;
0.001-0.05 BORDERLINE; β€0 NOT distinguishable) was locked before
precision values were computed. No goalpost-shifting after data.
76% tie rate confirms anonymization worked β the judge couldn't
tell which condition was which on three-quarters of comparisons.
A few deliberate framings β what ContextAtlas is and isn't relative to
neighboring tools:
vs. graph-based code intelligence (Graphify and similar). We're in
the same category β both build pre-computed indexes over codebases for
LLM agents via MCP. That's genuine category overlap, and we want to be
straight about it. Where we differ:
LSP-grounded vs. heuristic-extracted. ContextAtlas delegates all
structural questions to the language server (tsserver, Pyright,
gopls, ruby-lsp, csharp-ls). Graphify derives structure via parsing
and extraction.
Pre-composed bundles vs. graph primitives. ContextAtlas's MCP
tools return fused bundles in one call. Graphify exposes graph
operations (graph_query, get_neighbors, shortest_path) that
callers compose.
Narrow scope vs. broad scope. ContextAtlas indexes code + prose
git. Graphify ingests documentation, diagrams, research papers,
and more.
Claim-first vs. graph-first. ContextAtlas stores discrete claims
with severity labels, optimized for "what constrains this symbol?"
Graphify models the world as nodes and edges, optimized for "what
connects to this node?"
Whether our bets produce better results for a given workload is an
empirical question. See the numbers above.
vs. session-memory tools (claude-mem, engram, anamnesis). Those
capture accumulated session history β what Claude learned or did in
past conversations. ContextAtlas provides static architectural ground
truth extracted from your code, ADRs, and docs. Different information
sources with occasional overlap (when session discussions become ADRs
or commits), but fundamentally different problems. Session-memory
tools also can't really be committed to a repo; ContextAtlas's atlas
can.
vs. LSP-in-MCP (LSP-AI and similar). ContextAtlas uses LSP as
its source of structural truth. If you just want LSP-in-MCP, those
projects solve that well. ContextAtlas layers architectural intent
and git history on top.
vs. embedding-based search. We evaluated this and chose
symbol-keyed claims instead. Embeddings are fuzzy; LSP symbols are
exact. Embedding-based ranking is a post-MVP enhancement contingent
on benchmark evidence that it helps β see
ADR-09 for the full
rationale.
The committed-atlas pattern
ContextAtlas produces a committable team artifact β atlas.json β
that lives in the repo alongside your code and ADRs. This is the piece
that turns ContextAtlas from a personal productivity tool into a team
asset.
New team member clones the repo: they pull down atlas.json
with everything else. On first run, ContextAtlas imports the
committed atlas directly into their local cache β no extraction API
calls, no 10-minute wait. Productive from the moment they open
Claude Code.
Contributor submits a PR: if their code change affects
architectural claims, they regenerate atlas.json as part of their
commit. Reviewers see both the code change and the atlas diff in
the PR.
Developer bounces between machines: atlas state is
version-controlled, not trapped on one laptop.
Returning to a project after months away: pull the latest main,
and the atlas reflects everything the team did in your absence.
Only files changed since you last pulled need incremental reindex.
Open-source projects: casual contributors benefit immediately
without paying any setup cost. The project's accumulated
architectural knowledge flows to them automatically.
For teams that cannot commit the atlas, set atlas.committed: false
in the config. Every developer runs their own extraction. The team
artifact benefit is lost, but ContextAtlas still works as a personal
tool.
This model β committed team artifact with a local cache for query
performance β is a categorical difference from both session-memory
tools and knowledge-graph tools. Detailed in
ADR-06.
Architecture
code
INDEX TIME (once per source change)
ββββββββββββββββββββββββββββββββββββββ
ADRs βββββββββββ
Docstrings βββββ€
Git commits ββββΌβββΊ Opus 4.7 extraction
LSP symbols ββββ β
βΌ
atlas.json (committed to repo)
β
βΌ
SQLite + FTS5 BM25 (local cache)
QUERY TIME (every Claude call, zero API)
ββββββββββββββββββββββββββββββββββββββ
Claude Code: get_symbol_context("X")
β
βΌ
One fused bundle, sub-100ms
(LSP refs + intent + git + tests)
Five layers, each with one job:
MCP interface.get_symbol_context, find_by_intent, and
impact_of_change tools exposed to Claude.
Query fusion. Composes results from signal sources per query.
Extraction pipeline. Opus 4.7 reads prose docs and emits
structured claims keyed to symbols.
Storage. SQLite index, SHA-keyed for incremental reindex.
Signal fusion at query time works as a substantively cheap lookup:
when Claude calls get_symbol_context("OrderProcessor"), the MCP
handler hits the LSP for live structural facts (definition,
references, types) + reads the symbol's pre-extracted intent claims
from SQLite + folds in git heat + tests. The bundle returned to
Claude is composed, not computed β substantive joins happened at
index time. This is the substantive distinction from graph-based
alternatives that expose primitives (get_neighbors, shortest_path)
which callers compose at query time.
The architectural promise: expensive understanding happens once at
index time; queries are local dictionary lookups, zero API calls.
This bounds cost, latency, and unpredictability β and it's a hard
invariant, not an optimization.
What ContextAtlas does and doesn't send off your machine:
Sent to Anthropic's API (at index time only):
The full text of every file adrs.path and docs.include in
.contextatlas.yml match: by default ADRs, READMEs and other
markdown docs. A docs.include glob that matches source files (for
example docs/** over docs/examples/*.ts) sends those files whole,
code included; narrow the globs to documentation files to avoid it
The docstring text of exported symbols in your source files (the
comment text only, not the code), one symbol per request
Messages of commits that pass the commit filter (subject and body;
no diffs), one commit per request
This happens once per source per change β on initial index and on
incremental reindex of changed files and new commits. A commit is
sent once.
contextatlas index sends all three by default (v1.2+). Set
extraction.streams: [adr] in .contextatlas.yml to send ADRs and
docs only. The /index-atlas Skill processes ADRs, docstrings and
commit messages the same way inside your Claude Code session; it
does not extract docs.include pages.
contextatlas generate-adrs (CLI) sends a structural inventory
instead: source file paths and the symbol names the language server
lists for each file (top-level declarations plus class, interface
and namespace members, exported or not), plus README.md,
DESIGN.md and CLAUDE.md verbatim and any --reference-context
documents.
Never sent anywhere:
Your source code, apart from the docstring text above and any
source file a docs.include glob matches (sent whole, as above)
Your git history, apart from the filtered commit messages above
LSP reference and type data
Query contents at runtime
Stored locally only:
The extracted claims database (.contextatlas/index.db by default)
All runtime query resolution happens against this local SQLite file
At query time β every get_symbol_context call Claude makes during
your work β ContextAtlas performs a local SQLite lookup plus local LSP
calls. No network traffic. No model calls. Your code never leaves your
machine during normal use.
Index-time extraction uses the Anthropic API per standard API terms. If
your ADRs, docstrings or commit messages contain sensitive
architectural decisions, they'll be processed under those terms like
any other API-submitted content.
The Three Tools
The three MCP tools are not three parallel features β they're one fused
context substrate with three access patterns.
get_symbol_context β the primitive. "I know the symbol; give me
everything." Returns the full fused bundle (signature, ADR claims,
references, git heat, tests, types) in a single call. Multi-symbol
mode handles up to 10 symbols per request (per
ADR-15).
find_by_intent β the semantic-query composite. "I don't know
the symbol; find it by what it does." Ranks by BM25 against indexed
claim text in local SQLite FTS5 β no embedding service, no external
calls, deterministic results (per
ADR-09).
impact_of_change β the blast-radius composite. "I'm about to
change this; what breaks?" Adds git co-change patterns and test impact
on top of the primitive.
Refresh Discoverability
ContextAtlas atlas is a substrate you build once and refresh after code
or ADR changes. ONE canonical entry point per cohort path; behavior
adapts based on substrate state:
CLI
Skills
Cold-start
contextatlas index (full extraction)
/index-atlas (full extraction)
Refresh
contextatlas index (Phase 4 SHA-diff incremental)
/index-atlas (refresh-aware workflow)
SHA-diff incremental refresh per ADR-12
is substantively cheaper than cold-start scaffolding. Unchanged ADR
and docstring sources skip, and commits already extracted are not
extracted again; only changed sources are re-extracted. Since v1.2 the
CLI index covers all three streams (ADRs/docs, docstrings, commit
messages) like the Skill does; see extraction.streams under
Configuration.
Quick Start
Status: v0.9.0 shipped 2026-05-16. v1.0 public launch substrate
complete. Package not yet published to npm; install instructions
below describe the intended shape.
Runtime requirements:
Node.js 20 or newer.
A language server for each language you configure:
TypeScript β typescript-language-server (declared as a
peer dependency rather than a direct one, so you control the
version). Install alongside ContextAtlas
(e.g. npm i -D typescript-language-server typescript).
Python β Pyright on the PATH (also a peer dependency).
Go β gopls on the PATH (install via
go install golang.org/x/tools/gopls@latest).
Ruby β ruby-lsp 0.26.x. Recommended install via Bundler in
your project's Gemfile (gem 'ruby-lsp', '~> 0.26.0', require: false under group :development). Rails projects additionally
benefit from ruby-lsp-rails 0.4.x. Ruby 3.3+ required (4.0+
recommended).
C# / .NET β csharp-ls 0.24.x on the PATH (Roslyn LSP
wrapper). Install via dotnet tool install --global csharp-ls.
.NET SDK 8 minimum (10+ recommended; matches cohort backend
pin). On Windows, the %USERPROFILE%\.dotnet\tools directory
must be on PATH β the adapter enriches PATH automatically for
Bash/Git-Bash where the SDK installer only configures
PowerShell.
Path A β Claude Code Skills (60 seconds, no API key)
bash
npm install -g contextatlas
contextatlas init
Then in Claude Code:
code
/generate-adrs # Skip if you already have ADRs (any path; see Using existing ADRs below)
/index-atlas # Build the atlas
/prime-atlas # Verify connection (once per session)
Path B β CLI (90 seconds, API key required)
bash
npm install -g contextatlas
export ANTHROPIC_API_KEY=sk-...
contextatlas init
contextatlas generate-adrs # Skip if you already have ADRs (any path; see Using existing ADRs below)
contextatlas index
contextatlas doctor # Verify health
Using existing ADRs and docs
ContextAtlas extracts architectural intent from whatever ADRs and
documentation you already have β generate-adrs is for repos
without existing ADR substrate.
Existing ADRs at docs/adr/? Skip generate-adrs;
ContextAtlas extracts your existing ADRs automatically.
ADRs at a different path? Set adrs.path in
.contextatlas.yml (default: docs/adr/).
README, design docs, or other prose? ContextAtlas extracts
these too via docs.include (default: README.md +
docs/**/*.md). Add custom paths to extract additional
documentation surfaces.
The extraction pipeline produces structured claims from any prose
source pointed at via config β existing substrate doesn't go unused.
See Configuration below for full schema.
MCP server registration
Configure ContextAtlas as an MCP server in your Claude Code settings.
Choose based on whether contextatlas is on your PATH:
Option A β global binary on PATH (e.g., installed via
npm install -g or npm link):
If atlas.json is already committed (teammate ran it first, or it
came with the repo), ContextAtlas imports it instantly. No API
calls. You're ready in seconds.
If no atlas exists yet, ContextAtlas runs full extraction. Depending
on ADR count and size, this takes 1-10 minutes and costs a few
dollars in Opus API credits (CLI path) or session tokens (Skills
path). The resulting atlas.json can be committed so future
contributors skip this step.
Docstrings and commits add to the first run (v1.2+). The CLI
index also makes one API call per documented exported symbol and
one per filter-passing commit. Before its first call it prints an
estimate to stderr: calls and input tokens per stream, and a cost
range. For this repository's own first three-stream run, planned
without calling the API at commit 9d2bf4c, that was 479 calls (13
ADR/doc files, 461 docstrings, 5 commits), estimated at $4.42 to
$10.35.
extraction.streams: [adr] keeps extraction to ADRs and docs, and
--budget-warn <usd> warns when a run's spend passes a threshold.
Cost projection note. Until v1.2 this note said platform-billed
actuals typically run ~3x below script-reported costs because of a
prompt-cache discount (v0.4: cobra $5.44 β $1.82, httpx $5.53 β
$1.85, hono $10.89 β $3.65). That explanation is very likely wrong.
The v0.4 script costs used stale $15/$75 per-million-token prices,
exactly 3x the real $5/$25 (corrected in v0.6), and extraction
requests do not use prompt caching. cost_usd, computed at $5/$25,
is expected to track what you are billed; do not discount it or the
preview's range by 3x.
On subsequent runs, only files whose SHAs have changed since the
last index get reprocessed. Usually seconds.
Configuration
Create .contextatlas.yml in your repo root:
yaml
version:1languages:-typescript-python-go-ruby-csharpadrs:path:docs/adr/format:markdown-frontmatterdocs:include:-README.md-docs/**/*.mdgit:recent_commits:5atlas:committed:true# default; commits atlas.json to your repo# extraction:# streams: [adr] # CLI `index`: ADRs/docs only. Default: all three# # (adr, docstring, commit). Must include adr.
Credibility is built by stating what we don't claim.
Statistical methodology. All quality measurements are paired-t with
95% confidence intervals β no p-values. NHST at n=5 is
statistically void; CIs preserve effect-size visibility. Threshold
pre-registration honored verbatim (Option Ξ± strict three-tier framing
locked before precision values computed).
Single-judge model. v0.5 quality measurements use Sonnet 4.6 as
the judge with within-judge consistency β₯80% per axis (pass-1 vs
pass-2). Cross-vendor judge-panel graduation is post-v1.0 work.
Three benchmark repos. All quantitative claims are bounded to
hono (TypeScript, 186 files), httpx (Python, 23 files), and cobra
(Go, 19 files), plus our own dogfood. Generalization beyond these is
post-launch cohort work.
v0.5 substrate scope. Quality-axis measurements are 5 anchor cells
Γ nβ₯5 trials Γ 2 conditions (hono h1 auto-stretch to n=8); not
full-matrix replication. Matrix-completion graduation is post-v1.0.
v0.6 cross-cycle replication caveat. A targeted matrix-replication
subset at v0.6 (8 cells Γ n=5) showed attenuation on 2 of 4 quality
axes vs the v0.5 anchor cells (factual_correctness CLEANβBORDERLINE;
actionability BORDERLINEβNOT distinguishable). Root cause: the v0.5
measurements were against an earlier atlas version, and the cross-cycle
methodology didn't control for atlas-substrate-version. Full causal
investigation deferred to post-launch. Detail at
Phase-10 reference doc.
v0.3 single-run methodology. Phase 8 reports n=1 per cell;
blind-graded quality-axis measurement was added at v0.5. The Beta-vs-
Beta+CA reporting at Phase 8 carries the atlas-file-visibility caveat
(bias direction conservative β actual CA contribution likely larger
than published numbers indicate).
Dogfooding is not a measured benchmark. Throughout development,
ContextAtlas indexes its own ADRs and is used by Claude Code during
work on ContextAtlas itself. This is a development practice, not part
of the four-condition matrix β which runs only against the three
external targets.
Favorable and unfavorable results both published. Phase 7's
cross-harness asymmetry hypothesis was FALSIFIED on v0.3 substrate.
v0.6's atlas-substrate-version confound was surfaced and disclosed.
Tie and trick buckets routinely show CA net-negative; we report them
inline rather than burying them.
Status and Roadmap
Current: v0.9.0 (shipped 2026-05-16). v1.0 public launch substrate
complete; launch execution work folds into v1.0.0 without a separate
v0.9.1 tag.
v0.9 (May 16): Ruby adapter ship β fourth supported language
via ruby-lsp + ruby-lsp-rails per ADR-21. Repo launch substrate
(MIT license + community substrate + cycle docs migration to
docs/cycles/v0_X/). Launch positioning work in progress.
v0.8 (May 14): Substrate-equivalence + path-comparability +
BM25 activation. Closed Skill-substrate parity to CLI at 65-83%
claim ratio across hono/httpx/cobra benchmarks.
v0.7 (May 12): Launch-bearing cycle β Path-3 entry-point-
determined architecture (CLI/Skills equivalence per ADR-02
graduation) + generate-adrs feature with canonical depth-floor
enforcement via validate-adrs.
v0.6 (May 9): Pipeline mechanics + targeted matrix-replication
subset (8 cells Γ n=5 Γ 2 conditions). F1 PRIMARY atlas-substrate-
version confound surfaced; full causal investigation deferred to
post-launch.
The language adapter interface is a stable plugin surface β each new
language is an additive contribution, not a core change. See
docs/language-adapter-guide.md for
the contributor onboarding walkthrough.
Language
Adapter
LSP Server
Shipped
TypeScript
TypeScriptAdapter
typescript-language-server
v0.1
Python
PyrightAdapter
Pyright
v0.1
Go
GoAdapter
gopls
v0.2
Ruby
RubyAdapter
ruby-lsp (+ ruby-lsp-rails)
v0.9
C# / .NET
CsharpAdapter
csharp-ls (Roslyn)
v1.1
Contributing
ContextAtlas is MIT licensed and welcomes contributions. Areas where
contribution will be especially valuable:
New language adapters. The LanguageAdapter interface is small
and stable. Adding Java, .NET, Rust, Kotlin, or other language
support is a self-contained project. See
docs/language-adapter-guide.md.
Non-markdown intent sources. Currently we support markdown ADRs
with YAML frontmatter. RST, AsciiDoc, and other formats are welcome.
Benchmark repos. Additional benchmark coverage on more codebases
strengthens the eval.
Benchmarks and Methodology
Benchmarks and methodology live in a separate repository:
github.com/traviswye/ContextAtlas-benchmarks.
That repo contains the harness code, locked prompt sets, published
measurement results, and the full methodology document (RUBRIC.md).
Keeping the harness out of this repo means the benchmarks measure the
published contextatlas package's actual behavior rather than an
internal monorepo build.
Credits
Built during the "Build anything with Opus 4.7" hackathon.
ContextAtlas uses:
Claude Opus 4.7 for index-time intent extraction
typescript-language-server for TypeScript symbol resolution
Pyright for Python symbol resolution
gopls for Go symbol resolution
ruby-lsp for Ruby symbol resolution
csharp-ls for C# / .NET symbol resolution (Roslyn LSP wrapper)
better-sqlite3 for the index store
@modelcontextprotocol/server (MCP TypeScript SDK v2) for MCP server implementation