@cyanheads/cern-opendata-mcp-server
Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths via MCP. STDIO or Streamable HTTP.
7 Tools โข 1 Resource
Overview
Particle-physics data from the CERN Open Data Portal: collision, simulated and derived datasets, analysis software, environments and documentation from ALICE, ATLAS, CMS, LHCb and other experiments. Search with exact-vocabulary filters and live facet counts, open records with license and citation, list data files, assemble a record's analysis environment, and look up CMS good-run lists and trigger paths. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|
cern_opendata_search_records | Search datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and live facet counts |
cern_opendata_get_records | Fetch full metadata for 1โ20 records by recid, DOI, CMS dataset path or documentation slug, with license and citation |
cern_opendata_list_files | Page through a record's file indexes and files: XRootD URIs, HTTPS URLs, sizes, checksums, tape availability |
cern_opendata_get_analysis_env | Assemble a record's analysis environment: container images, CMSSW release, global tag, linked environment and software records, guide sections |
cern_opendata_get_validated_runs | Get a CMS validated-run (good-run) list for a dataset, a list or a run period, with luminosity-section ranges |
cern_opendata_search_trigger_paths | Look up CMS High-Level Trigger paths by name or prefix, parsed into run ranges, versions and L1 seeds |
cern_opendata_list_reference | Decode the vocabulary the other tools accept: experiments, record types, energies, formats, physics categories, LHCb stripping, identifiers, query syntax, licensing, run periods |
Resources
| Resource | Description |
|---|
cern-opendata://record/{recid} | One record's metadata, license and citation, in the cern_opendata_get_records record shape |
Tool-only clients get the same data from cern_opendata_get_records.
Capability reference
- Optional
query (an OpenSearch query_string, up to 500 characters) plus OR-list filters type, experiment, category (physics category, Higgs Physics::Standard Model), keywords, collision_energy, collision_type, file_type, availability, collection, and the LHCb magnet_polarity, stripping_stream and stripping_version, each an array or a comma-separated string; year_from/year_to and min_events/max_events bound the data-taking year and event count
sort (bestmatch, mostrecent for newest date_published first, title, title_desc), limit 1โ50 (default 10) and page from 1; page ร limit past 10,000 fails as page_window_exceeded
- Compact
hits with recids, plus thirteen live facets that each ignore their own filter (type and category with their secondary values); applied_filters echoes what ran, and values outside the verified vocabulary appear under unrecognized_values
ids: 1โ20 recids, DOIs, CMS dataset paths (/Primary/Era/TIER) or documentation slugs, mixed in one array or comma-separated string
- Each record carries a
license with its basis (record, cern_terms_default, not_stated) and, when it has a DOI, a ready citation
- Documentation and news bodies come in slices of up to 30,000 characters:
body_offset with that one id reads on from the body_next_offset the last slice returned
- Each response stays within 64,000 bytes: records past the budget are left out whole and listed in
deferred, to pass back as ids (the first record always comes back whole)
- Records also return what they state of a
variables dictionary (name, type, unit, description), a physics category, pile-up (pileup_html, with the pile-up datasets under links), keywords, and the LHCb magnet_polarity and stripping stream and version
- Unresolved identifiers land in
missing with interpreted_as and guidance instead of failing the call; file lists come from cern_opendata_list_files
recid required; without index, returns the record's file indexes and regular files, and with an index key, that index's files, read without the rest of the record
limit 1โ500 (default 50), continued with next_cursor; each file carries xrootd_uri, size_in_bytes, checksum and availability, an https_url when one can be built from its key or EOS path, and key or filename as the portal states them, and each index a uri_list_url listing every XRootD URI in it
- Files marked
on demand sit on tape and must be requested on the record's portal page first; an umbrella record with no files of its own returns its child recids under children
recid required; software carries the record's own container images, CMSSW release, global tag and environment recid
environment_records (condition, VM, validation) for the record's run periods and example_software that declares it works with the record, up to 50 between them; guides quotes the linked section of the first two portal guides, each capped at 12,000 characters, and a cut names the cern_opendata_get_records body_offset that reads on from it
- Always
separately_licensed: true; linked records or guides that can't be read leave a notice instead of failing the call
- Exactly one of
recid (a CMS collision dataset or a validated-run list) or run_period (Run2012B; 2012B also matches); variant full or muons_only; run_min/run_max; limit 1โ2000 (default 200)
- A dataset
recid bounds the runs to the dataset's first and last listed run, echoed in run_bounds; when several lists match, matched_lists names them and no runs are read
- Each run carries
lumi_sections and lumi_ranges; list.https_url downloads the whole list file. CMS only: other records fail as no_validated_runs
path: an exact name (HLT_IsoMu24, AlCa_EcalPi0) or a prefix with one trailing * (HLT_IsoMu*); a name without the HLT_ prefix is matched against record path names in the case given and also searched with HLT_ added, and a _v<n> version suffix is dropped; optional year, limit 1โ50 (default 10) and page
- Each per-year record is parsed into the primary
datasets its title names, first_seen, last_seen, per-version run ranges with their l1_seed, and HLT menu record links; parsed: false marks a record to read from its abstract_html
- CMS open data from 2011โ2016 only
- Optional
topic: experiments, record_types, collision_energies, collision_types, file_types, availability, categories, lhcb, identifiers, query_syntax, licensing or run_periods; omit it for every table
- Static and offline;
categories, lhcb and run_periods are dated snapshots, while the search facets and cern_opendata_get_validated_runs read the live portal
cern-opendata://record/{recid} resource
- One record by
recid (6004, or a prefixed recid such as atlas-160006; leading zeros ignored) as application/json, in the cern_opendata_get_records record shape: metadata (variable dictionary, physics category, pile-up and LHCb run conditions included), license and citation, without file lists
recid comes from cern_opendata_search_records; reads carry a 15-minute public cache hint
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
CERN Open Data-specific:
- Keyless and read-only; it never stages tape files or writes anything
- One shared pacer at 50 requests a minute, under the portal's published 60 per client IP, and one 50-second deadline per call across queue wait and retries
- Filter values canonicalized against the portal's verified vocabulary (
13 tev โ 13TeV, lhcb โ LHCb, Pb-Pb โ PbPb, dataset/collision โ Dataset::Collision); unknown values are sent as given and flagged
- Tape-resident (
ondemand) records are included in every search and lookup, where the portal otherwise drops them silently
- File manifests, and file indexes read on their own, are cached for 15 minutes, so paging through a record's or an index's files costs one read
Agent-friendly output:
- Provenance on every response:
portal_url on each hit and record, a license with its basis, a DOI citation, and applied_filters or effectiveQuery echoing what ran
- Graceful partial results: unresolved ids land under
missing with guidance, and unreadable linked records or guides in cern_opendata_get_analysis_env surface as a notice rather than a failure
- Discriminated outputs:
kind, license.basis, interpreted_as, scope, variant, run_bounds.source and parsed let callers branch on data, not string parsing
- Portal text kept as data: titles, descriptions, guide sections and file names are fenced or escaped in
content[] and relayed as received (HTML in _html fields) in structuredContent
Data and licensing
Portal metadata and datasets are CC0 under the CERN Open Data Terms of Use. Software, container images, documentation and guide code are licensed separately, per record (software is commonly GPL). cern_opendata_get_records reports each record's license and its basis: record when the record states one, cern_terms_default for a dataset that states none (CC0), and not_stated otherwise.
CERN asks reusers to cite each dataset's DOI in applications and publications; cern_opendata_get_records returns a ready citation for every record with a DOI.
This server is an independent project and is not affiliated with or endorsed by CERN.
Known limitations
- 60 requests a minute per client IP. The portal publishes this limit; the server paces itself to 50 a minute, and a call that cannot start within its deadline fails as
rate_limited with retryAfter. A hosted deployment shares one budget across every user behind its egress IP. cern_opendata_get_analysis_env and cern_opendata_get_validated_runs cost 2โ4 requests each.
- 10,000-result window. Search and trigger-path paging reach only the first 10,000 matches; narrow deeper result sets with filters.
- Facet lists are partial. Terms facets return the first 10 values alphabetically (
file_type up to 100), with the rest counted in other_count.
- Tape-resident files. Files with availability
on demand must be requested on the record's portal page before download; staging them is a write and out of scope. A record whose availability is ondemand lists none of its files through the API, so cern_opendata_list_files reports only the count and size its metadata states.
- CMS-only run lists and triggers. Trigger records cover 2011โ2016, and no prescale tables are published, so trigger detail is limited to what each record's abstract states. Muons-only run lists do not exist for Commissioning2010, Run2010B or the 2011 ReReco list.
- Glossary entries are not served. The portal's glossary links answer 404, so search excludes them.
Getting started
Public Hosted Instance
A public instance is available at https://cern-opendata.caseyjhand.com/mcp โ no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "streamable-http",
"url": "https://cern-opendata.caseyjhand.com/mcp"
}
}
}
Every caller of the hosted instance shares one portal request budget of 50 requests a minute. For sustained use, run your own instance.
Self-Hosted / Local
Add the following to your MCP client configuration file.
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cern-opendata-mcp-server:latest"]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- No API key or account: the CERN Open Data Portal is public.
Installation
- Clone the repository:
git clone https://github.com/cyanheads/cern-opendata-mcp-server.git
- Navigate into the directory:
cd cern-opendata-mcp-server
- Install dependencies:
- Configure environment:
Configuration
The server has no settings of its own; these framework variables apply.
| Variable | Description | Default |
|---|
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | HTTP server port. | 3010 |
MCP_HTTP_HOST | HTTP server host. | 127.0.0.1 |
MCP_SESSION_MODE | HTTP session mode: stateless, stateful, or auto. | stateless |
MCP_AUTH_MODE | Authentication: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (debug, info, warning, error, etc.). | info |
LOGS_DIR | Directory for log files (Node.js only). | <app-root>/logs |
OTEL_ENABLED | Enable OpenTelemetry. | false |
See .env.example for the common framework overrides.
Running the server
Local development
Project structure
| Directory | Purpose |
|---|
src/index.ts | createApp() entry point: registers the tools and resource, sets the server instructions, starts the portal client. |
src/mcp-server/tools | Tool definitions (*.tool.ts), shared input helpers and list enrichment. |
src/mcp-server/resources | The cern-opendata://record/{recid} resource definition. |
src/mcp-server/record-schema.ts | Record output schema shared by cern_opendata_get_records and the resource. |
src/services/cern-opendata | Portal client (pacing, retries, deadline, caches), normalization, vocabulary tables, text rendering, trigger parsing. |
tests/ | Unit and tool tests, mirroring the src/ structure, run against fixture portal responses. |
docs/design.md | Tool surface, verified portal behavior, and design decisions. |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches โ no
try/catch in tool logic
- Use
ctx.log for logging and ctx.enrich for notices and paging context
- Register new tools and resources in the barrels in
src/mcp-server/*/definitions/index.ts
- Wrap external API calls: validate raw โ normalize to domain type โ return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
License
This project is licensed under the Apache 2.0 License. See the LICENSE file for details.