This Model Context Protocol (MCP) server provides a way to find and fetch research datasets across 17 archives, omics registries, and literature sources. It exposes a single normalized model for dataset discovery and retrieval, based on the server description.
π οΈ Key Features
Find and fetch research datasets
Coverage across 17 archives, omics registries, and literature sources
Single normalized model for access
π Use Cases
Data discovery for bioinformatics research datasets
Aggregating datasets from sources such as omics registries and literature
β‘ Developer Benefits
Supports MCP workflows for model-context integration
Tags align with common research-data tooling topics (e.g., datacite, zenodo, pubmed, ncbi)
Python-oriented project scope
β οΈ Limitations
Supported sources are described only at a high level (17 archives, registries, literature); specific archives/tools are not listed in the provided data.
search one query across 17 sources β Zenodo, DataCite (Dryad /
Figshare / Dataverse / OSF / OpenNeuro / Mendeley), NCBI omics
(GEO / SRA / BioProject), BioStudies (EBI, incl. ArrayExpress),
literature (PubMed / OpenAIRE), HuggingFace datasets, DataONE
(eco / environmental), OmicsDI (proteomics / metabolomics), DANDI
(neurophysiology), CZ CELLxGENE (single-cell), OpenML (ML datasets),
RCSB PDB (structures), UniProtKB (proteins), the GWAS Catalog,
GBIF (biodiversity), data.gov (US federal open data), and NASA CMR
(Earth science) β deduplicated, normalized, and cross-linked. resolve any hit to its file
manifest, citation, trust signals, and the data it points at. fetch it to
disk, checksum-verified where the source publishes a checksum.
mcp-name: io.github.musharna/data-aggregator-mcp
β¨ Why this
Many research-data MCP servers wrap one source each. This one unifies many
behind six tools and one DataResource model, so an agent searches once and gets
back comparable records:
Multi-domain, one model β generalist archives + raw omics + literature,
deduplicated by DOI (the fetchable record wins over bare metadata).
Taxonomy synonym expansion β organism="Orobanche aegyptiaca" also matches
Phelipanche aegyptiaca (NCBI Taxonomy), so a species rename doesn't cost you
results.
Paper β data bridge β resolve a paper and get links to the GEO / SRA /
BioProject / DataCite records it produced.
Checked fetch β streams to disk with md5 / sha-256 verification where the
source publishes a checksum (a mismatch raises), and optional archive
unpacking. Many sources publish no checksum (see the Checksum column below);
those downloads are not verified, and the only content check is an HTML sniff
on files declared as PDF or XML, which rejects a paywall page served as a
"PDF".
Citations, access & full text β render a citation in any CSL style, get
normalized access/license, and pull open-access full text β all in one
resolve.
Trust signals β usage metrics (citations / views / downloads / likes),
version status (is_latest / superseded_by), and last_updated freshness,
surfaced wherever the source exposes them.
Interop exports β resolve(format="croissant") or "ro-crate" hands a
dataset to an ML or research-packaging pipeline as standard JSON-LD.
Operate on data in place β operate reads the schema, previews rows, or
runs a read-only SQL SELECT against a remote Parquet/CSV/TSV without
downloading it (Parquet footer + DuckDB httpfs range reads). Optional
[operate] extra; base install is unchanged.
Relate across records β relate takes a handful of resolved ids and
reports how they connect β shared accession, shared cross-identifier, an
explicit link, or version lineage β naming the literal shared value as
evidence. Metadata hints only: it never reads files or executes a join.
β Full rationale and a comparison vs. single-source servers, breadth gateways, and
ML-dataset tools: docs/POSITIONING.md.
β‘ Quickstart
Run with no install:
bash
uvx data-aggregator-mcp
Register with Claude Code:
bash
claude mcp add data-aggregator -- uvx data-aggregator-mcp
this machine only; any non-loopback value requires --allow-host
--port
8000
--allow-host HOST:PORT
auto on loopback
permitted Host header, repeatable β required off loopback
--allow-origin ORIGIN
derived
permitted browser Origin header, repeatable
--stateless
off
fresh transport per request, no session affinity
--json-response
off
plain JSON responses instead of SSE streams
The endpoint is served at /mcp/ β with the trailing slash. /mcp answers
307 redirecting there, which is fine for any client that follows redirects (a
307 preserves the POST body); point one that doesn't straight at /mcp/. In
stateful mode, sessions idle for 30 minutes are reaped.
DNS-rebinding protection is always on. A loopback bind derives its own
host/origin allowlist, so the default needs no configuration. A non-loopback bind
(--host 0.0.0.0, a LAN address, a container interface) refuses to start
without at least one explicit --allow-host β guessing an allowlist there is
precisely the hole the protection exists to close, so it fails loud instead of
open:
Once running, a request whose Host header is outside the allowlist is refused
with 421 Invalid Host header.
β οΈ fetch(dest=β¦) writes to the server's filesystem, not the client's.
Over stdio those are the same disk; over HTTP they may be different machines,
and the caller gets back paths it cannot read. Treat dest on an HTTP
deployment as server-side staging, or use stdio when you need the bytes
locally. search, resolve, operate, relate, and list_sources are
unaffected β they return data, not paths.
ποΈ Sources
Source
Discover
Fetch
Checksum
Zenodo
β
β
md5
DataCite β Figshare
β
β
md5
DataCite β Dataverse
β
β
md5
DataCite β OSF
β
β
md5
DataCite β Dryad
β
manifest onlyΒΉ
sha-256 (listed)
DataCite β Mendeley & others
β
β
β
NCBI SRA
β
β (ENA FASTQ)
md5
NCBI GEO
β
β (suppl/)
noneΒ²
NCBI BioProject
β
β SRA links
β
PubMed / OpenAIRE
β
β (OA full text)
noneΒ³
HuggingFace datasets
β
β (resolve URL)
noneΒ²
DataONE (eco/env)
β
β (Member Node)
md5 / sha-256
OmicsDI β PRIDE
β
β (HTTPS FTP)
noneΒ²
OmicsDI β MetaboLights
β
β (HTTPS FTP)
sha-256
OmicsDI β other MS repos
β
β
β
DataCite β OpenNeuro
β
β (snapshot)
noneΒ²
DANDI (neurophysiology)
β
β (302βS3)
sha-256
CZ CELLxGENE (single-cell)
β
β (H5AD/RDS)
noneΒ²
OpenML (ML datasets)
β
β (ARFF)
md5
RCSB PDB (structures)
β
β (.cif/.pdb)
noneΒ²
UniProtKB (proteins)
β
β (FASTA)
noneΒ²
BioStudies (EBI)
β
β (study files)
noneΒ²
GBIF (biodiversity)
β
β (Darwin Core)β΄
noneΒ²
data.gov (DCAT-US)
β
β (file URL)β΄
noneΒ³
NASA CMR (Earth science)
β
ββ΅
β
GWAS Catalog
β
β PMID bridge
β
ΒΉ Dryad downloads are token / bot-challenge gated, so fetch fails loud;
resolve still lists the files.
Β² No upstream checksum, so fetch does not verify these bytes. It still fails
loud on an HTTP error or when the download exceeds max_bytes.
Β³ No upstream checksum. Files declared as PDF or XML (literature full text, and
data.gov distributions with that mediaType) get an HTML sniff: an HTML login or
paywall page served in their place fails loud. Other files are not checked.
β΄ Only records that carry a downloadable file (a GBIF Darwin Core Archive, a
data.gov distribution URL); metadata-only records are discovery-only.
β΅ Discovery-only: granule downloads need an Earthdata login, which is not wired.
resolve returns the DOI and a data-access portal link.
Fan out across all wired sources in parallel and return compact DataResource
records, deduped by DOI. Per-source failures land in errors{} β never silently
dropped.
organism β expand the query with NCBI-Taxonomy synonyms; the expansion is
echoed in taxon_expansion, and results carry normalized taxa[]
({taxid, name}) plus a described_in link to plant-genomics-mcp for plant
taxa.
sources β restrict the fan-out, e.g. ["omics"].
size β max results (1β50).
kind β keep only dataset / sequencing_run / study / publication /
software. A record whose upstream type none of these covers (a Zenodo image, a
DataCite Audiovisual, an untyped record) is kind other and matches no filter.
published_after / published_before β filter by publication year.
rank β relevance (default) or semantic (re-rank the fetched page by
embedding similarity to the query; needs EMBEDDING_API_BASE, degrades to
relevance order otherwise).
understand β opt into LLM query understanding (default false). A free-text
query is normalized into a focused keyword query: conversational fluff
("I'm looking forβ¦", "where can I findβ¦") is stripped while the scientific
and entity terms are kept so they still match by text. The LLM also detects
structured entities (organism/disease/tissue/chemical/assay, kind) β these are
echoed in query_understanding.extracted for transparency but not
auto-applied, because ANDing LLM-inferred facets across free-text keyword
upstreams over-constrains and hurts recall. Only the cleaned keyword_core and
explicit year scopes are applied; the ontology resolvers still run on the
facets you pass (the LLM proposes, you dispose). Needs an LLM endpoint
(LLM_API_BASE); with none configured the search runs unchanged and notes it in
errors['understand']. Effectiveness is query- and model-dependent β opt-in /
default-off; validate the recall lift on your own corpus and LLM (see the eval
harness below).understand= was measured once (v0.38.0, 2026-06-11) on a
5-query verified gold set: mean recall@20 lift β0.10 against the plain query,
with 4 of the 5 queries neutral or better. multi_query= has not been measured.
multi_query β opt into diverse multi-query recall expansion (default false).
An LLM generates up to a few deliberately-diverse reformulations of your query
(different facets/synonyms/framings, not paraphrases), each is fanned out across
every source, and the deduped union is re-ranked against your original query,
aiming to reach records a single keyword query would miss. Bounded at
MAX_QUERY_VARIANTS (4, incl. the original), so it costs at most NΓ the upstream
calls. The original query's results are always among the candidates, but only the
top size of the re-ranked union are returned, so a result the plain query would
have returned can be displaced by one from a variant. Composes with
understand= (which structures variant 0). The variants used are echoed in
query_expansion. Needs an LLM endpoint (LLM_API_BASE); with none configured
the search runs as a normal single query and notes it in errors['multi_query'].
cursor β opaque token from a prior result's next_cursor; pages forward
across every source. In cursor mode the other params are read from the
token, so query is optional.
resolve(id, cite?, format?, trust?, fair?, use?)
Full record + files manifest. Routes by id shape β zenodo:7654321, a bare DOI,
datacite:10.5061/dryad.x, an omics id (sra:SRX079566, geo:GSE332789,
bioproject:PRJNA1468572), a literature id (pubmed:34320281, openaire:<id>),
a HuggingFace id (hf:owner/name), a DataONE id (dataone:doi:10.5063/F1HT2M7Q),
or an OmicsDI id (omicsdi:pride:PXD000001). Attaches, where available:
files[] β ENA FASTQ manifest (SRA), GEO suppl/, or the host repo's
native manifest (Figshare / Dataverse / OSF / Dryad).
access / license β normalized status
(open / embargoed / restricted / closed / unknown) and license where
the source exposes it.
identifiers β normalized {pmid, pmcid, doi}, plus an open-access
full-text FileEntry (EuropePMC XML, or an Unpaywall PDF fallback) for papers.
citation β pass cite=<format>: bibtex, ris, csl-json, or any CSL
style name (apa, mla, vancouver, β¦). DOI records use content
negotiation; others render CSL-JSON from metadata. Off by default; failures
degrade quietly.
trust signals β metrics (citations / views / downloads / likes),
is_latest / superseded_by (derived from version links), and last_updated
freshness, where the source provides them.
errors β {step: message} when an enrichment step failed on this record
(e.g. taxonomy during an NCBI rate limit); the rest of the record stands. Such a
record is not cached, so the next resolve retries the step.
truncated β {field: note} when a list on this record is deliberately partial,
e.g. a BioProject's links past 100 SRA runs: first 100 of 891 SRA runs; β¦. Empty
when every list is complete.
trust=true β attach retraction status (via Crossref) under trust{}.
One extra Crossref call; meaningful for DOI-bearing records only.
fair=true β attach an RDA-grounded FAIRness score (0β100 + F/A/I/R
sub-scores + actionable gaps) computed from the record metadata under fair{}.
Pure/local β no extra network call.
use=<intent> β attach a licence-compatibility advisory under
license_compat{} for the intended use (commercial / redistribute /
modify / ml-training). Returns ALLOW/REVIEW/DENY with the governing clause.
Metadata-derived advisory, not legal advice; an absent/unrecognized licence
yields REVIEW.
format β pass format="croissant" (file-level Croissant JSON-LD),
"ro-crate" (minimal RO-Crate 1.1), or "provenance" (one-call RO-Crate 1.1
data-availability dossier bundling version-currency, licence+SPDX, FAIR score,
and retraction status) to attach a standard manifest under the matching field.
Download files to disk and return their paths. Streams under a max_bytes guard
(force to override) with md5 / sha-256 verification wherever the source
publishes a checksum.
files β restrict to a subset of the resolved manifest.
extract β unpack downloaded zip / tar archives in place, guarded against
path traversal and runaway extracted size. Off by default.
Sources without a checksum are downloaded unverified. The one content check
there is an HTML sniff on files declared as PDF or XML (literature full text,
some data.gov distributions): it fails loud if the body is actually an HTML page.
Checksum-verified: Zenodo, SRA (ENA FASTQ), DataONE (Member-Node
objects), DataCite-hosted Figshare / Dataverse / OSF, OpenML
(ARFF), MetaboLights (via OmicsDI; sha-256 from the study's HASHES/) and
DANDI (sha-256; an asset whose hash DANDI has not computed yet is unverified).
Fetchable but unverified: GEOsuppl/, HuggingFace datasets,
PRIDE (via OmicsDI), DataCite-hosted OpenNeuro,
CZ CELLxGENE, RCSB PDB, UniProtKB, BioStudies,
GBIF (Darwin Core Archives), data.gov distributions, and literature
open-access full text.
Dryad, other DataCite repos, other OmicsDI repos (MassIVE / GNPS / ...),
BioProject, NASA CMR, and the GWAS Catalog are discovery-only and
raise FetchNotSupportedError.
list_sources()
Wired sources with their capabilities β layer, kinds, supported filters,
fetchability, operable flag, id examples, auth, and rate limits.
operate(op, id, file?, query?, n?, columns?)
Inspect or query a remote tabular file (Parquet / CSV / TSV) without
downloading it. Addresses a file by catalog id + file name (defaults to the
first tabular file on the resolved record). Ops:
schema β column names + types (reads the Parquet footer / sniffs the CSV
header; no full load).
preview β a small sample of rows.
head β the first n rows (default 20), optionally restricted to columns.
sql β a read-only SELECT (the file is the view data), e.g.
SELECT col, count(*) FROM data GROUP BY 1.
peek β per-column profile via DuckDB SUMMARIZE (type, null-rate,
approximate distinct count, min/max, numeric quartiles) without
downloading the file. Like head/sql, reads the whole file and honors
the source-size ceiling.
Backed by the Parquet footer reader + DuckDB httpfs range reads. sql runs in
a locked-down DuckDB (read-only, local filesystem disabled, single-SELECT
validation, row / wall-clock caps). Requires the optional [operate] extra
(pip install data-aggregator-mcp[operate]); without it, operate returns a
clear install-the-extra message and the other five tools are unaffected.
Any HuggingFace dataset with a datasets-server converted view is operable
(schema / preview / head / sql): resolve surfaces the auto-converted
Parquet files (source="hf-datasets-server") even for datasets stored as
JSON/JSONL/arrow, so pass file=<config>/<split>/...parquet to pick a split when
there are several. A split HF converted only in part (its first 5 GB) is named
<config>/partial-<split>/..., so a query on it covers that part, not the whole split.
relate(ids)
Cross-resource join/harmonization hints. Given 2β10 resource ids, relate resolves
each (TTL-cached) and reports how they relate and on what key they could be joined:
shared_accession β same BioProject/SRA/GEO accession on β₯2 records β joinable key.
shared_identifier β same doi/pmid/pmcid across records β same work / paperβdata link.
explicit_link β one record's links[] points at another input record.
version_lineage β one record supersedes another (dedupe, don't join, those).
Hints only.relate never reads file columns, fetches files, or executes a
join/merge/conversion β every hint names the shared value as evidence. Per-id resolve
failures are reported in errors, not fatal; an empty result carries an explanatory
note.
Prompts
Three workflow prompts surface in clients (e.g. /mcp__data_aggregator__* in
Claude Code):
find_data β find datasets for a topic, optionally scoped to an organism.
data_behind_paper β find the datasets / accessions behind a paper.
search_resolve_fetch β walk the end-to-end search β resolve β fetch flow.
βοΈ Configuration
All optional, set via environment variables:
NCBI_API_KEY β raises the NCBI E-utilities rate limit (3 β 10 req/s) used by
the omics, literature, and taxonomy lookups.
DATA_GOV_API_KEY β optional; data.gov works without it through the keyless
catalog API (catalog.data.gov). With a free
api.data.gov key set, data.gov requests go
through the api.data.gov gateway instead (1,000 requests/hour per key).
UNPAYWALL_EMAIL β enables the Unpaywall fallback leg of literature full-text
retrieval (the EuropePMC leg works without it).
NCBI_EMAIL β contact address sent to NCBI's ID converter; falls back to
UNPAYWALL_EMAIL when unset.
DATAVERSE_BASE_URL β resolve Dataverse DOIs against a different installation
(default https://dataverse.harvard.edu).
CACHE_TTL_SECONDS β resolve-cache lifetime in seconds (default 3600; an
unparseable value falls back to that default).
EMBEDDING_API_BASE / EMBEDDING_API_KEY / EMBEDDING_MODEL β an
OpenAI-compatible embeddings endpoint enabling rank=semantic. Absent β
semantic re-rank degrades to relevance order. Key is optional (keyless local
servers supported); model defaults to text-embedding-3-small.
LLM_API_BASE / LLM_API_KEY / LLM_MODEL β an OpenAI-compatible
/chat/completions endpoint enabling search(understand=true) (NLβstructured
query rewriting) andsearch(multi_query=true) (diverse multi-query recall
expansion). Absent β both run the raw query unchanged and note it in
errors['understand'] / errors['multi_query']. Key is optional (keyless local
servers supported); model defaults to gpt-4o-mini (a passthrough string β set
it to whatever your endpoint serves). multi_query fans out at most
MAX_QUERY_VARIANTS (4, incl. the original) variants, bounding the NΓ cost.
To measure the recall lift of understand=true / multi_query=true on a small
labeled set, run the gated eval harnesses (need a live LLM endpoint):
They print per-query and mean recall@20 (understand / multi-query off vs. on). See
the fixtures at scripts/eval_understand_fixture.json and
scripts/eval_multi_query_fixture.json.
π§ͺ Develop
bash
uv venv && uv pip install -e ".[dev]"
uv run pytest -q
uv run ruff check src tests
DATA_AGGREGATOR_MCP_LIVE=1 uv run pytest -k live -q # real-API probes
The README demo (examples/assets/demo.svg) is recorded network-free from
examples/_demo_stdio.py β see the header of that file to re-record.