@cyanheads/gnomad-genetics-mcp-server
Look up variant allele frequencies by ancestry, gene loss-of-function constraint, gene variant lists, and sequencing coverage over gnomAD β with ClinVar significance joined in β via MCP. STDIO or Streamable HTTP.
7 Tools (+1 opt-in) β’ 2 Resources β’ 1 Prompt
Overview
Population genetics over gnomAD (Broad Institute), with ClinVar clinical significance joined in from NCBI. Look up per-ancestry allele frequencies, gene loss-of-function constraint, gene variant catalogs, and sequencing coverage, then query large result sets with SQL from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|
gnomad_get_variant | Full population record for one or more variants β AC/AN/AF overall and per genetic-ancestry group, homozygote/hemizygote counts, quality flags, transcript consequence, in-silico predictors, and joined ClinVar significance. |
gnomad_get_gene_constraint | Gene loss-of-function constraint β pLI, LOEUF (oe_lof_upper) with confidence interval, observed/expected ratios, and Z-scores. By HGNC symbol or Ensembl gene ID. |
gnomad_list_gene_variants | Every variant in a gene, transcript, or region with allele frequencies and predicted consequences, filterable by consequence class and max allele frequency. |
gnomad_get_coverage | Sequencing coverage across a gene, transcript, or region β mean/median depth and the fraction of samples over depth thresholds, per callset track. |
gnomad_search_clinvar | Gene-level ClinVar detail via NCBI E-utilities β classified variants, review status, conditions, submission counts, and gnomAD-compatible variant IDs, paged by offset. |
gnomad_dataframe_query | Run a read-only SQL SELECT across canvas tables staged by the list tools. |
gnomad_dataframe_describe | List the tables staged on a canvas and their columns before writing SQL. |
gnomad_dataframe_drop | Drop a named table from a canvas to reclaim memory. Opt-in via GNOMAD_DATAFRAME_DROP_ENABLED=true β off by default. |
Resources
| Resource | Description |
|---|
gnomad://variant/{dataset}/{variantId} | Population record for one variant β mirrors gnomad_get_variant. |
gnomad://gene/{dataset}/{gene}/constraint | Gene loss-of-function constraint β mirrors gnomad_get_gene_constraint. |
All resource data is also reachable via tools. The list tools (gnomad_list_gene_variants, gnomad_get_coverage, gnomad_search_clinvar) return analytical row sets rather than stable single-URI documents, so they are not exposed as resources β call the tools instead.
Prompts
| Prompt | Description |
|---|
gnomad_variant_triage | Guided rare-disease variant-triage workflow: population frequency β gene constraint β callability check, in order. |
Capability reference
- Batch up to 25 IDs per call (default; raise via
GNOMAD_MAX_VARIANT_BATCH), each a chrom-pos-ref-alt variantId on chromosome 1β22, X, or Y with an optional chr prefix (e.g. 1-55051215-G-GA) or an rsID (e.g. rs11591147)
- Per-item partial success β a malformed or absent ID lands in
failed[] without failing the others
- Each
failed[] item is { variant, error, reason, recovery }: reason is typed (invalid_variant_id, variant_not_found, upstream_unavailable, β¦) and recovery is the hint the tool declares for it; an ambiguous rsID adds candidates
- Mitochondrial IDs (
M, MT, chrM) are refused per item as mitochondrial_unsupported with no fetch β gnomAD models mitochondrial variants separately, and they are outside this server
- Per-ancestry frequency vector is returned in full, never collapsed to a single global AF
- Reports which callset(s) (
exome / genome) carry the variant, quality flags, transcript consequence, in-silico predictor scores, and the ClinVar significance gnomAD joins per variant
- An empty
found[] for a well-formed ID means the variant is not in the chosen dataset β pair with gnomad_get_coverage to confirm the position is callable before concluding true absence
-
Accepts an HGNC symbol (PCSK9) or an Ensembl gene ID (ENSG00000169174)
-
Returns pLI (>0.9 intolerant), LOEUF / oe_lof_upper with its lower bound, observed/expected ratios for LoF / missense / synonymous, and the three Z-scores
-
constraint_release names the release the metrics come from:
dataset | Source | constraint_release | LoF-intolerance guidance |
|---|
gnomad_r4 | GRCh38 gnomAD constraint | gnomAD v4.1.2 | LOEUF < 0.45 |
gnomad_r3 | GRCh38 gnomAD constraint (gnomAD publishes no v3 constraint) | gnomAD v4.1.2 | LOEUF < 0.45 |
gnomad_r2_1 | GRCh37 gnomAD constraint | gnomAD v2.1.1 | LOEUF < 0.35 |
exac | GRCh37 ExAC constraint | ExAC r0.3 | pLI only |
-
ExAC publishes pLI, the Z-scores, and observed/expected counts only, so on exac the ratios and LOEUF are null and constraint_flags is empty
-
Many genes have null constraint (sparse upstream) β null fields are reported as such, never fabricated
-
constraint_flags carries the caveat flags gnomAD attaches to a gene's constraint (e.g. no_exp_lof, syn_outlier)
- Supply exactly one of
gene, transcript_id, or region (chrom-start-stop, 1-based inclusive, chromosome 1β22, X, or Y with an optional chr prefix)
- A region must span less than 2,500,000 bp (stop β start) and hold at most ~30,000 variants; an unserved chromosome or out-of-range coordinate fails as
invalid_region and an over-wide span as region_too_large, both before any fetch, while a region over the variant ceiling fails as region_too_large after one request, carrying gnomAD's own message
- Mitochondrial targets (an
M/MT region, or a gene or transcript gnomAD places on chromosome M) fail as mitochondrial_unsupported instead of returning an empty list
- Optional filters: one
consequence_class (lof / missense / synonymous / other) and/or a maximum allele frequency
- A result too large to inline (the preview holds about 14,000 characters of rows, keeping a response near 24 KB) is staged on a DataCanvas table named
gene_variants, returned as canvas_id and table_name beside the preview β inspect it with gnomad_dataframe_describe, then query it with gnomad_dataframe_query to rank by AF, count by consequence, or group across every row. The response notice names the table and both tools
- A result that fits inline stages no table and uses no canvas (
canvas_id is empty) unless you pass a canvas_id
- Passing a
canvas_id always writes the result to gene_variants on that canvas, REPLACING the previous table (it never appends), even when the result fits inline; a result with no variants removes the table
- When the canvas is disabled (
CANVAS_PROVIDER_TYPE != duckdb) the tool returns the same capped inline preview (as many rows as fit about 14,000 characters) and the SQL path is unavailable
- A blank
gene or transcript_id counts as omitted
- Supply exactly one of
gene, transcript_id, or region; a blank gene or transcript_id counts as omitted
region takes the same chrom-start-stop form as gnomad_list_gene_variants β chromosome 1β22, X, or Y, optional chr prefix, a span under 2,500,000 bp β and fails with invalid_region or region_too_large before any fetch otherwise
- Mitochondrial targets fail as
mitochondrial_unsupported instead of returning empty coverage
- Returns mean and median read depth plus the mean fraction of samples covered at each threshold (1Γ through 100Γ), summarized per callset track
coverage_source narrows to one track (exome / genome); omit to return every available track
- A variant missing from a well-covered region is informative; one missing from a poorly-covered region is not
- Returns a gene's classified ClinVar variants β clinical significance, review status with a 0β4 star rating, associated conditions, molecular consequences, and submission counts
- Each row carries gnomAD-compatible identifiers:
canonical_spdi, rsids, and grch38_variant_id (chrom-pos-ref-alt, set for SNVs, MNVs, and delins), which gnomad_get_variant resolves in the GRCh38 datasets
- Optional filters:
clinical_significance (e.g. pathogenic; blank means no filter) and a minimum star rating (min_review_stars, 0β4)
- Returns one window of up to 500 ClinVar records per call:
total_found is ClinVar's candidate count for the gene and filter terms (taken before the significance and star filters narrow each window), truncated and next_offset say whether more remain, and offset / limit (1β500) page through them. limit counts records before the filters, so a window can return fewer rows
- VariationIDs ClinVar returns no summary for are listed in
unavailable_ids rather than returned as blank rows
- Accepts an HGNC symbol only β ClinVar's gene index doesn't resolve Ensembl gene IDs, unlike the other gnomAD tools. An Ensembl gene ID returns guidance to resolve its symbol and searches nothing, leaving any
canvas_id you pass untouched
- A window too large to inline (the preview holds about 11,000 characters of rows, keeping a response near 24 KB) is staged on the
clinvar_variants canvas table β inspect it with gnomad_dataframe_describe, then query it with gnomad_dataframe_query. A window that fits inline stages no table unless you pass a canvas_id; passing one always writes the window to clinvar_variants, REPLACING the previous table, and a window with no rows removes it. With the canvas disabled, the preview is the same capped preview (as many rows as fit about 11,000 characters)
- Keyless, but honors
NCBI_API_KEY for a higher rate limit (10 vs 3 req/s)
- Runs single-statement, read-only SQL
SELECTs against a canvas table staged by gnomad_list_gene_variants or gnomad_search_clinvar β writes, DDL, and file/HTTP table functions are rejected by the canvas gate
- Reference tables by the name the staging tool returned (
gene_variants or clinvar_variants)
- Returns one page of the result:
offset (default 0) and limit (default 100, max 500) select it, and a page also ends before its rows pass 10,000 characters of JSON, which keeps a full page under about 24 KB
- Output:
rows (dynamic columns per the SQL projection), columns, offset, returned, total (exact row count; null when the result exceeds the canvas row cap), truncated (rows exist after this page), and next_offset (null on the last page) β follow next_offset until it is null
- Each page re-runs the SQL: stable paging needs an
ORDER BY over a unique key (such as variant_id) and an unchanged table. Paging stops at the canvas row cap (CANVAS_DEFAULT_ROW_LIMIT, 10,000 by default); filter or aggregate in SQL to reach rows past it
- A row over the 10,000-character budget fails with
row_too_large β select fewer or narrower columns
- Requires
CANVAS_PROVIDER_TYPE=duckdb β otherwise fails with a canvas_disabled error
- Lists every table staged on a canvas with its row count and column schema (name and DuckDB type)
- Call it before writing SQL for
gnomad_dataframe_query
- Requires
CANVAS_PROVIDER_TYPE=duckdb β otherwise fails with a canvas_disabled error
- Drops a named table from a canvas to reclaim memory β a deliberate mutation (
readOnlyHint: false, destructiveHint: true) on an otherwise read-only surface
- Opt-in via
GNOMAD_DATAFRAME_DROP_ENABLED=true; absent from tools/list when off, since per-table TTL already reclaims memory automatically
- Requires
CANVAS_PROVIDER_TYPE=duckdb β otherwise fails with a canvas_disabled error
gnomad://variant/{dataset}/{variantId} resource
- Population record for one variant as
application/json β mirrors gnomad_get_variant; the dataset segment keeps the URI self-describing
variantId accepts a chrom-pos-ref-alt ID or an rsID, same grammar as the tool
- Typed errors:
invalid_variant_id (outside the coordinate/rsID grammar), mitochondrial_unsupported (an M/MT/chrM ID), variant_not_found, ambiguous_rsid (with candidates), and the gnomAD failures graphql_error, upstream_build_mismatch, upstream_unavailable, upstream_timeout, upstream_access, and invalid_upstream_response β each with its declared recovery hint
gnomad://gene/{dataset}/{gene}/constraint resource
- Gene loss-of-function constraint as
application/json β mirrors gnomad_get_gene_constraint, including the same constraint_release for each dataset segment (exac serves ExAC r0.3 constraint; gnomad_r3 serves the GRCh38 gnomAD v4.1.2 table)
gene accepts an HGNC symbol or Ensembl gene ID
- Typed errors:
gene_not_found when no gene matches in the requested build, invalid_constraint_data when gnomAD's metrics fall outside their valid ranges, and the gnomAD failures graphql_error, upstream_unavailable, upstream_timeout, upstream_access, and invalid_upstream_response β each with its declared recovery hint
gnomad_variant_triage prompt
- Arguments:
variant required (chrom-pos-ref-alt or rsID); gene and dataset optional. dataset is one of gnomad_r4, gnomad_r3, gnomad_r2_1, exac β any other value is rejected; omitted or blank, the emitted calls carry no dataset and use the server default. A blank gene counts as omitted
- Emits a three-step chain as one user message: population frequency (
gnomad_get_variant) β gene constraint (gnomad_get_gene_constraint) β callability check (gnomad_get_coverage on the exact position, not gene-level)
- Rejects a malformed
variant with a validation error before generating the chain
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
gnomAD-specific:
- Single keyless GraphQL source for the entire core surface β ClinVar significance is joined per variant inside gnomAD's own response
dataset and reference_genome are distinct, coherence-validated parameters (v4/v3 β GRCh38, v2.1/ExAC β GRCh37); both are echoed in every tool's output so a wrong-build coordinate mismatch is visible
- Polite client β conservative concurrency cap (
GNOMAD_MAX_CONCURRENCY, default 2) and exponential backoff against a community-funded, rate-limited API
- In-conversation SQL analytics:
gnomad_list_gene_variants and gnomad_search_clinvar stage results too large to inline on a DuckDB-backed canvas table β gnomad_dataframe_describe lists its columns, gnomad_dataframe_query runs SQL over every row
Agent-friendly output:
- Per-ancestry allele-frequency vector returned in full, never collapsed to a single global AF β the cross-ancestry contrast is the signal clinical interpretation needs
- Graceful partial failure β
gnomad_get_variant returns per-item failed[] rows, each with a typed reason and its recovery hint, instead of failing the whole batch
- Provenance on every response β effective
dataset and reference_genome echoed back; null upstream fields preserved as null, never fabricated
- Every error carries a typed
reason and the recovery hint its tool or resource declares for that reason, so callers know the next move: input problems (incoherent_build, invalid_target, invalid_variant_id, invalid_region, region_too_large, mitochondrial_unsupported, ambiguous_rsid), absences (gene_not_found, variant_not_found), gnomAD refusals (graphql_error, upstream_build_mismatch, invalid_constraint_data), upstream faults (upstream_unavailable, upstream_timeout, upstream_access, invalid_upstream_response), and the canvas (canvas_disabled, row_too_large). gnomad_get_variant puts the same reason and hint on each failed[] item
Getting started
Public Hosted Instance
A public instance is available at https://gnomad-genetics.caseyjhand.com/mcp β no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"gnomad-genetics-mcp-server": {
"type": "streamable-http",
"url": "https://gnomad-genetics.caseyjhand.com/mcp"
}
}
}
Self-Hosted / Local
Add the following to your MCP client configuration file. gnomAD is a free, keyless API β no credentials required.
{
"mcpServers": {
"gnomad-genetics-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/gnomad-genetics-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"gnomad-genetics-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/gnomad-genetics-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"gnomad-genetics-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"ghcr.io/cyanheads/gnomad-genetics-mcp-server:latest"
]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
To enable the SQL analytics path, also set CANVAS_PROVIDER_TYPE=duckdb β @duckdb/node-api ships as a dependency, so nothing extra to install.
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- No API key β gnomAD's GraphQL endpoint is keyless. An optional
NCBI_API_KEY raises the gnomad_search_clinvar rate limit.
Installation
- Clone the repository:
git clone https://github.com/cyanheads/gnomad-genetics-mcp-server.git
- Navigate into the directory:
cd gnomad-genetics-mcp-server
- Install dependencies:
- Configure environment:
Configuration
All variables are optional; the server runs keyless with the defaults below.
| Variable | Description | Default |
|---|
GNOMAD_API_BASE_URL | gnomAD GraphQL endpoint. Override for a private mirror or testing. | https://gnomad.broadinstitute.org/api |
GNOMAD_DEFAULT_DATASET | Dataset used when a tool call omits dataset (gnomad_r4 / gnomad_r3 / gnomad_r2_1 / exac). | gnomad_r4 |
GNOMAD_REQUEST_TIMEOUT_MS | Per-request timeout against the GraphQL endpoint, in milliseconds. | 30000 |
GNOMAD_MAX_CONCURRENCY | Cap on concurrent upstream requests β politeness against a community-funded API. | 2 |
GNOMAD_MAX_VARIANT_BATCH | Maximum variant IDs accepted per gnomad_get_variant call. | 25 |
CLINVAR_BASE_URL | NCBI E-utilities base URL for gnomad_search_clinvar. | https://eutils.ncbi.nlm.nih.gov/entrez/eutils |
NCBI_API_KEY | Optional NCBI key. Raises the E-utilities rate limit from 3 to 10 req/s. | β |
CANVAS_PROVIDER_TYPE | Set to duckdb to enable the spill/SQL path behind the list tools. When none, they return a capped inline preview. | none |
GNOMAD_DATAFRAME_DROP_ENABLED | Gate for the opt-in gnomad_dataframe_drop tool. Off by default. | false |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for the HTTP server. | 3010 |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (RFC 5424). | info |
OTEL_ENABLED | Enable OpenTelemetry instrumentation. | false |
See .env.example for the full list of optional overrides.
Running the server
Local development
-
Build and run:
bun run rebuild
bun run start:stdio
bun run start:http
-
Run checks and tests:
bun run devcheck
bun run test
bun run lint:mcp
Docker
docker build -t gnomad-genetics-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=http -p 3010:3010 gnomad-genetics-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/gnomad-genetics-mcp-server. OpenTelemetry peer dependencies are installed by default β build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
| Directory | Purpose |
|---|
src/index.ts | createApp() entry point β registers tools/resources/prompts and inits services. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts) and shared input schemas. |
src/mcp-server/resources | Resource definitions (*.resource.ts). |
src/mcp-server/prompts | Prompt definitions (*.prompt.ts). |
src/services/gnomad | gnomAD GraphQL client, query documents, and domain types. |
src/services/clinvar | NCBI E-utilities client for the optional ClinVar tool. |
src/services/canvas-accessor.ts | Module-level accessor for the framework's optional DataCanvas. |
Development guide
See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches β no
try/catch in tool logic
- Use
ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
- Register new tools and resources in the
createApp() arrays in src/index.ts
- Wrap external API calls: validate raw β normalize to domain type β return output schema; never fabricate missing fields
Data attribution
gnomAD data is provided by the Genome Aggregation Database (Broad Institute). ClinVar data is provided by NCBI.
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
License
Apache-2.0 β see LICENSE for details.