ENCODE Toolkit MCP Server (io.github.ammawla/encode-toolkit)
The io.github.ammawla/encode-toolkit MCP server provides 20 MCP tools and 48 skills focused on ENCODE Project genomics workflows. It supports searching and downloading data and running pipelines for genomics research infrastructure, with topics spanning bioinformatics and MCP integrations.
🛠️ Key Features
20 MCP tools for ENCODE-related operations
48 skills for additional ENCODE workflow coverage
Genomics search and download capabilities
Pipeline support for ENCODE Project use cases
🚀 Use Cases
Searching ENCODE genomics resources
Downloading ENCODE data
Executing ENCODE-focused genomics pipelines
⚡ Developer Benefits
Topics and tooling aligned to MCP (includes “mcp-server”, “claude-code-mcp”, and “claude-plugin” context)
Versioned beta release: 0.3.0
Python 3.10+ compatible
AGPL-3.0 licensed
⚠️ Limitations
Readme excerpt provided does not describe other constraints beyond beta status and supported Python version
Search ENCODE, cross-reference 14 databases, run 7 analysis pipelines, and generate publication-ready methods — all from natural language in Claude Code.
Start from ENCODE but go everywhere: discover histone peaks, cross-reference with GWAS variants, check ClinVar pathogenicity, pull GTEx expression, analyze TF binding motifs from JASPAR, run pipelines, and generate publication-ready methods with full provenance — in one conversation.
Citation Notes
If you use ENCODE-Toolkit, please cite:
Alex M. Mawla. (2026). ENCODE-Toolkit: an MCP server, Claude plugin, and skills suite for ENCODE genomic data access and analysis. Zenodo. https://doi.org/10.5281/zenodo.18917511
BibTeX
bibtex
@software{mawla_2026_encode_toolkit,
author = {Mawla, Alex M.},
title = {ENCODE-Toolkit: an MCP server, Claude plugin, and skills suite for ENCODE genomic data access and analysis},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.18917511},
url = {https://doi.org/10.5281/zenodo.18917511}
}
"Find all histone ChIP-seq experiments for human pancreas tissue"
"What ATAC-seq data is available for mouse brain?"
"Search for RNA-seq on GM12878 cell line"
"What histone marks have ChIP-seq data for pancreas?"
Download and track
"Download all BED files from ENCSR133RZO to ~/data/encode"
"Track experiment ENCSR133RZO with its publications"
"Export citations for my tracked experiments as BibTeX"
Cross-reference databases
"What GWAS variants overlap islet enhancers?"
"Check ClinVar pathogenicity for rs7903146"
"Pull GTEx expression for TCF7L2 across tissues"
"Find JASPAR motifs for HNF4A binding sites"
Run pipelines
"Set up a ChIP-seq pipeline for my H3K27ac experiments"
"Run ATAC-seq analysis with ENCODE-standard QC thresholds"
Generate methods and provenance
"Log that I created filtered_peaks.bed from ENCSR133RZO using bedtools"
"Generate a methods section for my analysis with citations"
More example prompts
Experiment details
"Show me the full details for experiment ENCSR133RZO"
"What files are available for ENCSR133RZO?"
"List only the BED files from ENCSR133RZO"
Bulk downloads
"Download all FASTQs from human pancreas ChIP-seq to /data/fastqs"
"Get the IDR thresholded peaks from these experiments"
"Download the bigWig signal tracks for H3K27me3 in GRCh38"
Compatibility analysis
"Are experiments ENCSR133RZO and ENCSR000AKS compatible for combined analysis?"
"Compare these two ChIP-seq experiments"
Provenance chains
"Show me the provenance chain for my derived files"
"What files have I derived from ENCSR133RZO?"
The Problem
Using genomics databases today means:
Navigate web portals, click through dozens of filters
Manually find the right experiments and files across multiple databases
Write custom scripts to batch download
Lose track of which files came from where
With ENCODE Toolkit, just tell Claude what you need:
"Find all histone ChIP-seq data for human pancreas tissue"
Claude searches ENCODE, returns a structured table of 66 experiments with targets, replicates, and file counts. Downloads are organized by experiment with MD5 verification and full provenance tracking.
Available Tools (20)
Five core tools are shown below. The remaining 15 are collapsed for readability.
encode_search_experiments
Search ENCODE experiments with 20+ filters.
Parameter
Type
Description
assay_title
string
Assay type: "Histone ChIP-seq", "ATAC-seq", "total RNA-seq", "Hi-C", etc.
organism
string
Species (default: "Homo sapiens")
organ
string
Organ: "pancreas", "brain", "liver", "heart", "kidney", etc.
biosample_type
string
"tissue", "cell line", "primary cell", "organoid"
target
string
ChIP target: "H3K27me3", "H3K4me3", "CTCF", etc.
biosample_term_name
string
Specific biosample: "GM12878", "HepG2", etc.
limit
int
Max results (default: 25)
encode_get_experiment
Get full details for a single experiment including all files, quality metrics, and audit info.
Parameter
Type
Description
accession
string
Experiment ID (e.g., "ENCSR133RZO")
encode_download_files
Download specific files by accession to a local directory.
Get external references linked to tracked experiments for cross-server workflows.
Parameter
Type
Description
experiment_accession
string
Filter by experiment (optional)
reference_type
string
Filter by type (optional)
Authentication
Most ENCODE data is public and requires no authentication. Just install and use.
For restricted/unreleased data, ask Claude: "Store my ENCODE credentials"
Credentials are encrypted using your OS keyring (macOS Keychain, Linux Secret Service, Windows Credential Locker) and never stored in plaintext. Get your access keys from your ENCODE profile.
Plugin Skills (47)
When installed as a Claude Code plugin, ENCODE Toolkit includes 47 literature-backed workflow skills that guide Claude through complex genomics tasks. Each analysis skill includes evidence-based quality thresholds, assay-specific metrics, and citations to primary literature.
Core Skills
Skill
Description
setup
Install and configure the ENCODE Toolkit server
search-encode
Search and explore ENCODE experiments and files
download-encode
Download files with organization and verification
track-experiments
Track experiments, citations, and provenance locally
cross-reference
Connect ENCODE data to PubMed, bioRxiv, ClinicalTrials.gov
Analysis skills (9)
Skill
Description
quality-assessment
Evaluate experiment quality using ENCODE metrics — assay-specific thresholds for ChIP-seq (FRiP, NSC, RSC, NRF, IDR), ATAC-seq (TSS enrichment, NFR ratio), RNA-seq (mapping rate, gene body coverage), WGBS (bisulfite conversion, CpG coverage), Hi-C (cis/trans ratio), and CUT&RUN/CUT&Tag. Backed by Landt 2012, Buenrostro 2013, ENCODE Phase 3 (2020), Li 2011
integrative-analysis
Combine multiple experiments with batch effect awareness — integration strategies (peak overlap, signal correlation, DiffBind, DESeq2, ChromHMM, ABC model). Backed by Ernst & Kellis 2012, Ross-Innes 2012, Love 2014, Fulco 2019
regulatory-elements
Discover enhancers, promoters, insulators from combinatorial histone marks — ENCODE cCRE classification (926,535 elements), ChromHMM state interpretation. Backed by ENCODE Phase 3 (2020), Roadmap Epigenomics (2015), Whyte 2013
epigenome-profiling
Build comprehensive chromatin state profiles — three-tiered histone panels, ChromHMM 15-state model, bivalent chromatin analysis. References the chromatin biology catalog
compare-biosamples
Compare experiments across tissues and cell types — biosample hierarchy, tissue-specific regulation, batch effect detection. Backed by Roadmap Epigenomics (2015), Leek 2010
visualization-workflow
Generate publication-quality visualizations: genome browser tracks, heatmaps, and signal profiles
motif-analysis
Discover and analyze TF binding motifs in regulatory regions using HOMER, MEME, and JASPAR
peak-annotation
Annotate genomic peaks with features (promoter/enhancer/intergenic), nearest genes, and functional categories
batch-analysis
Batch processing and QC screening across multiple ENCODE experiments with systematic quality filtering
Functional genomics skills (1)
Skill
Description
functional-screen-analysis
Analyze CRISPR screens, MPRA, and STARR-seq data from ENCODE — MAGeCK, BAGEL2, MPRAflow integration
Data aggregation skills (4)
Skill
Description
histone-aggregation
Union merge of histone ChIP-seq peaks across studies — signalValue-based noise filtering, sample-of-origin tagging, ENCODE blacklist removal. Backed by ChIP-Atlas (Oki 2018), Amemiya 2019, Perna 2024
accessibility-aggregation
Union merge of ATAC-seq and DNase-seq peaks — cross-platform integration, peak summit preservation. Backed by Corces 2017, Amemiya 2019, Zhao 2020
hic-aggregation
Union catalog of Hi-C chromatin loops (BEDPE) — resolution-aware anchor matching, loop caller concordance tracking. Backed by Loop Catalog (Reyna 2025), Mustache (Roayaei Ardakany 2020)
Work with scRNA-seq and scATAC-seq data — platform comparison, cross-study integration, WNN multimodal analysis. Backed by Hao 2021, Stuart 2019
disease-research
Disease-focused workflows — GWAS variant interpretation, disease-tissue mapping, heritability enrichment, drug target identification via Open Targets. Backed by Buniello 2019, Finucane 2015
publication-trust
Publication integrity assessment — 5-level trust scoring, retraction/erratum detection, citation analysis. Integrates with PubMed, bioRxiv, and Consensus
bioinformatics-installer
Install all bioinformatics tools for ENCODE analyses — 7 conda environment YAMLs, 3 install scripts, 134+ tools across ChIP-seq, ATAC-seq, RNA-seq, WGBS, Hi-C, DNase-seq, CUT&RUN
scientific-writing
Generate publication-ready methods sections, figure legends, supplementary tables, and data availability statements with full tool citations
liftover-coordinates
Convert genomic coordinates between assembly versions (hg19/hg38, mm9/mm10) using UCSC liftOver, CrossMap, Ensembl REST API, and rtracklayer
External database skills (9)
Skill
Description
gtex-expression
Query GTEx tissue expression data via REST API for gene expression context across 54 tissues
clinvar-annotation
Annotate variants with ClinVar clinical significance, pathogenicity, and review status
cellxgene-context
Query CellxGene single-cell atlas for cell type expression context across tissues
gwas-catalog
Search NHGRI-EBI GWAS Catalog for trait associations, risk alleles, and study metadata
jaspar-motifs
Query JASPAR database for transcription factor binding motifs and matrix profiles
ensembl-annotation
Ensembl VEP variant annotation, Regulatory Build, coordinate liftover, gene lookup via REST API
geo-connector
Search NCBI GEO for complementary datasets, cross-reference with ENCODE, FTP downloads
gnomad-variants
gnomAD population allele frequencies, gene constraint (LOEUF/pLI), structural variants via GraphQL
ucsc-browser
UCSC Genome Browser REST API for cCRE tracks, TF binding clusters, and sequence retrieval
Pipeline execution skills (7)
Pipeline
Assay
Aligner
Caller
pipeline-chipseq
ChIP-seq
BWA-MEM
MACS2 + IDR
pipeline-atacseq
ATAC-seq
Bowtie2
MACS2 (Tn5-adjusted)
pipeline-rnaseq
RNA-seq
STAR
RSEM + Kallisto
pipeline-wgbs
WGBS
Bismark
MethylDackel
pipeline-hic
Hi-C
BWA
Juicer + HiCCUPS
pipeline-dnaseseq
DNase-seq
BWA
Hotspot2
pipeline-cutandrun
CUT&RUN
Bowtie2
SEACR
Each pipeline includes a SKILL.md overview, 5-stage reference files (preprocessing through QC), a complete Nextflow DSL2 pipeline, a Dockerfile, and deployment configurations for local, SLURM, GCP, and AWS.