@cyanheads/ensembl-mcp-server
Look up genes, fetch sequences, predict variant consequences, find orthologs, and retrieve cross-database xrefs from Ensembl REST via MCP. STDIO or Streamable HTTP.
7 Tools • 4 Resources • 1 Prompt
Overview
Gene, sequence, and variant data for vertebrates and other model organisms from the Ensembl REST API. Look up genes, fetch sequences, predict variant consequences, find orthologs, and cross-reference external databases from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|
ensembl_list_species | List species supported by Ensembl with display name, common name, assembly, taxon ID, and division |
ensembl_lookup_gene | Resolve a gene by symbol + species or by stable ID to its Ensembl ID, genomic location, biotype, and transcript list |
ensembl_get_sequence | Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region |
ensembl_query_region | Find genomic features (genes, transcripts, variants, regulatory elements, exons) overlapping a chromosomal region |
ensembl_predict_variant | Predict functional consequences of a sequence variant using the Ensembl Variant Effect Predictor (VEP) |
ensembl_get_homology | Find orthologs and/or paralogs of a gene across species with percent identity and taxonomy level |
ensembl_get_xrefs | Retrieve cross-database references for a gene — HGNC, UniProt, EntrezGene, OMIM, RefSeq, Reactome, and others |
Resources
| Resource | Description |
|---|
ensembl://gene/{id} | Gene record by stable ID (ENSG…) — location, biotype, description, and transcript list |
ensembl://transcript/{id} | Transcript record by stable ID (ENST…) — parent gene, location, biotype, canonical flag, and length |
ensembl://species | Supported Ensembl species for the endpoint default division (vertebrates on the default endpoint) |
ensembl://species/{division} | Supported species in one division (EnsemblVertebrates, EnsemblPlants, EnsemblFungi, EnsemblMetazoa, EnsemblProtists) |
All resource data is also reachable via the ensembl_list_species tool, which additionally filters by name.
Prompts
| Prompt | Description |
|---|
ensembl_gene_dossier | Structured workflow for assembling a complete gene profile: symbol → ID + location → sequence → variants → orthologs → xrefs |
Capability reference
- Filter by division (
EnsemblVertebrates, EnsemblPlants, EnsemblFungi, EnsemblMetazoa, EnsemblProtists) or nameContains for a local substring match against name, display name, and common name
- Omit
division to return the endpoint default division (vertebrates, ~356 species on the default GRCh38 endpoint)
- Returns internal name (the value every other tool expects), display name, common name, taxon ID, assembly, and division
- Required first step — species names like
homo_sapiens are opaque to non-biologists
- Exactly one of
symbol (+ optional species, default homo_sapiens), id, ids (batch, up to 20), or symbols (batch, up to 20)
expand_transcripts (default false) adds the full transcript list with biotype and canonical flag
- Batch modes (
ids/symbols) return a succeeded/failed split with per-item error strings instead of failing the call
- Errors:
not_found, invalid_species, no_input, conflicting_input
type: genomic (default, includes introns), cdna (spliced), cds (coding only), protein
- Accepts a stable ID (
ENSG…/ENST…/ENSP…) or a region — species:chr:start-end, or bare chr:start-end with species set; a region spans at most 10,000,000 bases, with start at or below end
expand_5prime / expand_3prime (default 0) extend flanking base pairs for genomic and region queries
protein and cds require a transcript or protein ID, not a gene ID; region ids are genomic-only
- Returns a bounded window:
offset (0-based, default 0) and max_length (default 10000, 0 for the rest uncapped) index the resolved sequence, flanks included
length is always the full sequence length; truncated and nextOffset say whether more follows and where to resume, so walking nextOffset reconstructs the whole sequence
- Errors:
not_found, type_mismatch, missing_species, invalid_region
region in chr:start-end format, at most 5,000,000 bases; feature array (at least one) defaults to ["gene"], also accepts transcript, variation, regulatory, exon; optional biotype filter
- Defaults to genes only — requesting
variation on a large locus can match 44,000+ features
max_results caps the feature list (default 100, 0 uncapped); totalCount always reports the true count found
assemblyName (e.g. GRCh38) names the assembly the coordinates are on
- Exon rows carry a
parentId and rank, since one exon is reported once per parent transcript
- Errors:
invalid_region, invalid_species
variant accepts HGVS (transcript-relative or genomic), region+allele (chr:start:end:strand/allele), or a dbSNP rsID
max_transcript_consequences (default 10) and max_pubmed_ids_per_variant (default 10) cap large VEP results; set either to 0 for the full set, or include_all_colocated_pubmed: true for uncapped PubMed IDs
- Returns most severe consequence term, per-transcript impact (HIGH/MODERATE/LOW/MODIFIER), and colocated known variants with clinical significance
- Totals (
transcriptConsequencesTotal, pubmedTotal) are always reported even when capped
- Errors:
invalid_notation, not_found
- Exactly one of
symbol (+ species, default homo_sapiens) or id; optional target_species filter
type: orthologues (default), paralogues, or all
max_results caps the homolog list (default 25, 0 uncapped); totalCount always reports the true count available
- Errors:
not_found, no_input, conflicting_input
id (ENSG…/ENST…) required; optional dbname filter (e.g. HGNC, Uniprot_gn, EntrezGene, MIM_GENE, RefSeq_mRNA, Reactome, GO)
- Uses the
xrefs/id endpoint, returning the full cross-reference set (56+ entries for well-annotated genes like BRCA2)
- Errors:
not_found
ensembl://gene/{id} resource
- Returns location, biotype, description, and transcript list for a gene stable ID (
ENSG…); version suffix optional
- Errors:
not_found
ensembl://transcript/{id} resource
- Returns parent gene, location, biotype, canonical flag, and length for a transcript stable ID (
ENST…); version suffix optional
- Errors:
not_found
ensembl://species resource
- No parameters — returns the endpoint default division (vertebrates, ~356 species on the default GRCh38 endpoint)
- For a named division, read
ensembl://species/{division} instead
ensembl://species/{division} resource
division required: EnsemblVertebrates, EnsemblPlants, EnsemblFungi, EnsemblMetazoa, or EnsemblProtists
ensembl_gene_dossier prompt
- Arguments:
gene_symbol required; species optional (default homo_sapiens)
- Sequences a 7-step workflow: resolve the gene → fetch the protein sequence → find variants in the locus → predict variant consequences → find cross-species orthologs → get external database IDs → synthesize the dossier
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
Ensembl-specific:
- Keyless REST API — no API key required; Ensembl REST is fully public at 55,000 req/hr
- Rate-limit-aware service layer: retries 429 honoring
Retry-After, and retries transient 5xx and HTML error pages
- Batch POST endpoints used throughout —
POST /lookup/id and POST /lookup/symbol/{species} (up to 1,000 items each upstream) reduce N+1 round trips in multi-gene workflows
- GRCh37 legacy support via
ENSEMBL_BASE_URL — point the entire server at https://grch37.rest.ensembl.org for clinical workflows on the older assembly
- All coordinate-bearing responses echo the assembly name so agents never see a bare genomic position without assembly context
Agent-friendly output:
ensembl_get_sequence returns sequences in bounded windows (10,000 characters by default) with the full length and a nextOffset to continue, so a long gene or locus never lands in one response unasked
ensembl_list_species is explicitly the discovery step — tool descriptions call out the opaque internal-name format and direct agents to it before using species-dependent tools
- Cross-tool chaining made explicit: xref IDs from
ensembl_get_xrefs are described as inputs for protein and literature servers; the ensembl_gene_dossier prompt sequences all 6 tools into one research workflow
Getting started
Public Hosted Instance
A public instance is available at https://ensembl.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"ensembl-mcp-server": {
"type": "streamable-http",
"url": "https://ensembl.caseyjhand.com/mcp"
}
}
}
Self-Hosted / Local
Add the following to your MCP client configuration file.
{
"mcpServers": {
"ensembl-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/ensembl-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"ensembl-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/ensembl-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"ensembl-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"ghcr.io/cyanheads/ensembl-mcp-server:latest"
]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- No API key required — Ensembl REST is fully public.
Installation
- Clone the repository:
git clone https://github.com/cyanheads/ensembl-mcp-server.git
- Navigate into the directory:
- Install dependencies:
- Configure environment:
Configuration
All configuration is validated at startup via Zod schemas in src/config/server-config.ts.
| Variable | Description | Default |
|---|
ENSEMBL_BASE_URL | Ensembl REST API base URL. Override for GRCh37 (https://grch37.rest.ensembl.org) or a local mirror. | https://rest.ensembl.org |
MCP_TRANSPORT_TYPE | Transport: stdio or http | stdio |
MCP_HTTP_PORT | HTTP server port | 3010 |
MCP_HTTP_ENDPOINT_PATH | HTTP endpoint path | /mcp |
MCP_SESSION_MODE | HTTP session mode: auto, stateful, or stateless. Schema default auto resolves to stateful; this server explicitly uses stateless. | stateless |
MCP_AUTH_MODE | Authentication: none, jwt, or oauth | none |
MCP_LOG_LEVEL | Log level (debug, info, warning, error, etc.) | info |
LOGS_DIR | Directory for log files (Node.js only) | <project-root>/logs |
OTEL_ENABLED | Enable OpenTelemetry | false |
See .env.example for the full list of optional overrides.
Running the server
Local development
-
Build and run:
bun run rebuild
bun run start:stdio
bun run start:http
-
Run checks and tests:
bun run devcheck
bun run test
bun run lint:mcp
Docker
docker build -t ensembl-mcp-server .
docker run --rm -p 3010:3010 ensembl-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/ensembl-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
| Directory | Purpose |
|---|
src/index.ts | createApp() entry point — registers tools/resources/prompts and inits services |
src/config | Server-specific environment variable parsing and validation with Zod |
src/mcp-server/tools | Tool definitions (*.tool.ts) — 7 tools |
src/mcp-server/resources | Resource definitions (*.resource.ts) — gene, transcript, species |
src/mcp-server/prompts | Prompt definitions (*.prompt.ts) — gene dossier workflow |
src/services/ensembl | Ensembl REST API client — HTTP, rate-limit handling, retry, error normalization |
tests/ | Unit and integration tests mirroring src/ |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — no
try/catch in tool logic
- Use
ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
- Register new tools and resources in the
createApp() arrays in src/index.ts
- Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
License
Apache-2.0 — see LICENSE for details.