@cyanheads/devops-status-mcp-server
Check vendor status pages, inspect SSL/TLS certificates, verify DNS propagation, and get incident-response playbooks via MCP. STDIO or Streamable HTTP.
7 Tools โข 1 Resource
Overview
Vendor status pages, SSL/TLS certificates, and DNS propagation โ normalized across Atlassian Statuspage, Status.io, Slack, AWS Health, Google Cloud Service Health, Azure status, and Firehydrant backends, plus direct TLS/DNS checks for any domain. List and check 52 built-in vendors, fetch incident timelines, watch a persisted stack, and get a tailored incident-response playbook, all without API keys. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|
devops_list_vendors | List vendors in the built-in registry, optionally filtered by name or category. Returns slug, display name, category, and status page URL. |
devops_status_check | Check the current health status for one or more vendors. Returns per-vendor indicator (none / minor / major / critical / maintenance), degraded components, and active incident summaries. |
devops_get_incidents | Fetch incident history for a vendor โ active, resolved, or scheduled maintenance. Returns the full incident timeline with per-update bodies and affected components. |
devops_watch_stack | Check the health of a named vendor stack persisted in session state. Pass vendors once to save the list; subsequent calls reuse it. Returns an aggregate health rollup plus per-vendor detail. |
devops_check_certs | Inspect SSL/TLS certificate health for one or more domains via a real TLS handshake. Reports expiry, chain depth, protocol version, cipher suite, and HSTS presence. Pure TypeScript โ no external API. |
devops_check_dns | Resolve DNS records and verify propagation for one or more domains across Google (8.8.8.8), Cloudflare (1.1.1.1), and Quad9 (9.9.9.9). Reports per-resolver latency and resolver discrepancies. Pure TypeScript โ no external API. |
devops_suggest_action | Instruction tool โ returns a tailored incident-response playbook and pre-filled follow-up tool calls given a vendor name and optional incident context. No external calls; fully deterministic. |
Resources
| Resource | Description |
|---|
devops-status://vendors/{name} | Full registry entry for a vendor by slug โ status page URL, category, and API type. |
All resource data is also reachable via tools. Tool-only agents are fully supported.
Capability reference
- Accepts an optional free-text
query (matches name and slug, case-insensitive) and an optional category filter โ eight categories: cloud, cdn-edge, dev-platform, data, comms, auth, monitoring, ai
- Returns slug (what to pass to other tools), display name, category, and status page URL
- 52 built-in entries, most on Atlassian Statuspage;
aws, gcp, azure, gitlab, neon, slack, and redis-cloud route through native-API adapters normalized to the same shape
azure reads Microsoft's Azure status RSS feed, which carries no severity or lifecycle: each posted item is an open minor incident with its services and regions as affected components, and the feed is empty while nothing is posted
- Statuspage-compatible pages not in the registry are still reachable by passing a raw base URL to other tools
Built-in vendor registry:
| Category | Vendors |
|---|
cloud | digitalocean, linode, aws, gcp, azure |
cdn-edge | cloudflare, akamai |
dev-platform | gitlab, github, npm, vercel, netlify, render, fly-io, circleci, travis-ci, snyk, atlassian, figma, launchdarkly |
data | mongodb-atlas, planetscale, supabase, neon, redis-cloud, elastic, influxdb, upstash, cloudinary, segment |
comms | slack, discord, twilio, sendgrid, mailgun, hubspot, brevo, courier, loops |
auth | auth0, clerk, workos |
monitoring | datadog, sentry, new-relic, grafana-cloud, honeycomb |
ai | openai, anthropic, elevenlabs, pinecone, cohere |
- Accepts registered vendor slugs (e.g.,
github, aws) or raw Atlassian Statuspage base URLs, mixed freely โ up to 20 per call
mode: "summary" (default): indicator + degraded components + active incidents; mode: "detailed" adds the full component list (capped at component_limit, default 50, max 500) and scheduled maintenance windows
Promise.allSettled fan-out โ one failing vendor never blocks the rest; failures surface as a per-vendor error field
- Results served from a 60-second in-memory cache;
cached: true on each result
summary partitions the batch into operational / degraded / down / maintenance / unavailable counts
nextToolSuggestions carries one pre-filled devops_suggest_action call per vendor with an active problem โ indicator minor/major/critical, or an open incident of that impact โ with the vendor slug (or normalized URL), indicator, affected components, and latest incident title filled in; a vendor that could not be checked never gets one, so the list is empty when no checked vendor has a problem
filter: all (default, incidents + scheduled maintenances), active (investigating/identified/monitoring), resolved (fully resolved), or scheduled (maintenance windows only)
- Returns per-update bodies in chronological order, affected component names, duration in minutes for resolved incidents, and a direct shortlink to the incident page
limit (1โ50) with offset for paging; a truncated result discloses the total and the next offset to fetch
- Some vendor feeds cap their own history (
upstreamCeiling). Atlassian Statuspage's API stops at the newest 50 incidents, which on a busy page is about a week
since (YYYY-MM-DD, up to 24 months back, with filter: "all" or "resolved") leaves out incidents that started before that date. On Statuspage vendors it also reads the status page's quarterly history archive back to that date, reaching past the 50-record ceiling
- Every incident carries
source: api for the status API, or history for an archive record, which has the title, impact, start and end times, and final update message but no components or update timeline. When a record is in both, the API version is kept
- If the archive can't be read (the page doesn't publish one, it times out, or it comes back in an unexpected shape), the call still returns the status API result, with a
notice saying how far history reached
- AWS keeps a resolved event listed for hours after it ends, so
filter: "resolved" returns only those still listed; Azure's feed lists open items only, so filter: "resolved" is always empty for it; AWS, Azure, Google Cloud, and Slack publish no maintenance windows, so filter: "scheduled" is always empty for them
- On the first call, provide
vendors to define the stack โ it is saved to tenant-scoped session state under stack_name
- Subsequent calls can omit
vendors; the saved list is reused automatically
- Multiple stacks coexist via distinct
stack_name values (e.g., "production", "data-layer") โ letters, digits, hyphens, and underscores, optionally separated by single dots or slashes, 1-64 characters
- Aggregate
health rollup: all_operational / maintenance (a vendor in a scheduled window, nothing worse open) / degraded / partial_outage / major_outage / unknown (a vendor could not be reached) โ never all_operational when any vendor errored or is in a window
nextToolSuggestions pre-fills a devops_suggest_action call for each vendor with an active problem, same as devops_status_check
- Note: stack state is in-memory; it does not persist across server restarts
- Accepts bare hostnames (no
https:// prefix) โ up to 10 per call
- Reports: days to expiry (flagged
warning at < 30 days, critical at < 7), certificate subject and SANs, issuer common name, chain depth, negotiated TLS version (flags 1.0 and 1.1 as insecure), cipher suite
- HSTS detection: sends a minimal HTTP/1.1 GET over the same TLS socket, reads the
Strict-Transport-Security response header
status: "critical" distinguishes a hostname mismatch (hostname_verification_error) from an untrusted chain (authorization_error) โ both would be rejected by ordinary clients
- Per-domain failures are reported inline (
status: "error", reason in error) rather than throwing โ useful partial results when checking multiple domains. flags is empty unless a handshake completed; a server that completes one without presenting a certificate still reports its TLS session
- Configurable port (default 443) and timeout per domain
- Queries Google (8.8.8.8), Cloudflare (1.1.1.1), and Quad9 (9.9.9.9) in parallel per domain
- Supported record types: A, AAAA, CNAME, MX, TXT, NS (defaults to A, AAAA, MX, TXT)
- Reports per-resolver latency, propagation discrepancies, and human-readable flags
- Discrepancies are typed:
partial_resolution (some resolvers answered, others didn't) signals a real problem; value_variation (all answered, different values) is normal for anycast/geo-steered domains
- Custom resolver list supported โ pass public resolver IP literals to test resolver-specific behavior (private and loopback resolvers need
DEVOPS_STATUS_ALLOW_PRIVATE_TARGETS=true); an empty resolvers or record_types array uses the defaults
- Up to 10 domains per call; per-domain timeouts configurable
- Category-tailored markdown playbook (cloud, CDN, dev-platform, data, comms, auth, monitoring, AI); falls back to generic guidance for unrecognized vendors
- Accepts a vendor slug or display name, resolved to the canonical slug so pre-filled follow-up arguments stay valid
- Optional
incident_summary / affected_components prepend a targeted subsystem section (e.g. GitHub Actions โ CI/CD steps, Cloudflare DNS โ DNS/TTL guidance); optional vendor_indicator leads with severity-tailored urgency framing
nextToolSuggestions pre-fills follow-up tool calls, including cert/DNS checks when your_domain is given โ execute in sequence
- When
DEVOPS_STATUS_DISABLE_ACTIVE_PROBES=true, guidance swaps the unregistered probe tools for equivalent manual commands (dig, openssl s_client)
devops-status://vendors/{name} resource
- Returns the full registry entry for a vendor slug โ status page URL, category, and API type (
statuspage, statusio, slack, aws, gcp, azure, firehydrant)
- Cached publicly for 1 hour โ the registry is compiled in and identical for every caller
- Same data is reachable via
devops_list_vendors โ tool-only agents are fully supported
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
DevOps-status-specific:
- No API keys required โ every status backend is a public API; TLS and DNS use Node.js stdlib (
node:tls, node:dns)
- 52-vendor built-in registry covering cloud, CDN, dev-platform, data, comms, auth, monitoring, and AI categories; adapter layer normalizes Status.io, Slack, AWS Health, Google Cloud Service Health, Azure status, and Firehydrant backends into the Statuspage shapes; extendable via raw Statuspage URL passthrough
- 60-second in-memory cache on status reads shared across all tenants โ prevents thundering-herd on batch calls
devops_watch_stack persists named vendor lists in tenant-scoped state for repeat morning checks or pre-deploy sweeps
devops_suggest_action dispatches category-specific playbooks deterministically โ no LLM sampling dependency, works in all clients
Agent-friendly output:
- Batch tools (
devops_status_check, devops_watch_stack, devops_check_certs, devops_check_dns) use Promise.allSettled โ one failing target never blocks the rest; errors surface as inline error fields
cached: true / checked_at on every status result โ agents know when data was fetched
- Discriminated indicator and status enums (
none / minor / major / critical / maintenance; operational / degraded_performance / partial_outage / major_outage / under_maintenance) โ callers branch on data, not string parsing
nextToolSuggestions in devops_suggest_action pre-fills tool arguments from incident context โ agents can execute the playbook mechanically
Getting started
Public Hosted Instance
A public instance is available at https://devops-status.caseyjhand.com/mcp โ no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"devops-status-mcp-server": {
"type": "streamable-http",
"url": "https://devops-status.caseyjhand.com/mcp"
}
}
}
Self-Hosted / Local
No API key required. Add the following to your MCP client configuration file:
{
"mcpServers": {
"devops-status-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/devops-status-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"devops-status-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/devops-status-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"devops-status-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"ghcr.io/cyanheads/devops-status-mcp-server:latest"
]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
Prerequisites
- Bun v1.4.0 or higher (or Node.js v24+).
- No API keys or external accounts required.
Installation
- Clone the repository:
git clone https://github.com/cyanheads/devops-status-mcp-server.git
- Navigate into the directory:
cd devops-status-mcp-server
- Install dependencies:
- Configure environment:
Configuration
No API keys required. All environment variables are optional.
| Variable | Description | Default |
|---|
DEVOPS_STATUS_CACHE_TTL_MS | In-memory cache TTL for vendor status reads (all backends) in milliseconds. | 60000 |
DEVOPS_STATUS_FETCH_TIMEOUT_MS | Per-request timeout for vendor status API calls (all backends) in milliseconds. | 8000 |
DEVOPS_STATUS_CERT_TIMEOUT_MS | Default timeout_ms for devops_check_certs (per-domain TLS handshake, milliseconds). A caller-passed timeout_ms overrides it. | 5000 |
DEVOPS_STATUS_DNS_TIMEOUT_MS | Default timeout_ms for devops_check_dns (per domain+resolver query, milliseconds). A caller-passed timeout_ms overrides it. | 3000 |
DEVOPS_STATUS_ALLOW_PRIVATE_TARGETS | When true, disables SSRF guards for user-supplied URLs and domains. For trusted local/intranet deployments only. | false |
DEVOPS_STATUS_DISABLE_ACTIVE_PROBES | When true, omits the arbitrary-target probe tools (devops_check_dns, devops_check_certs) from the registered tool surface; the five vendor-registry/incident tools remain. For shared/public multi-tenant instances. | false |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for HTTP server. | 3010 |
MCP_SESSION_MODE | HTTP session handling: stateless, stateful, or auto. The server declares stateless in source โ it holds no per-session state โ so setting this is only needed to override that. | stateless |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (RFC 5424). | info |
LOGS_DIR | Directory for log files (Node.js only). | <project-root>/logs |
OTEL_ENABLED | Enable OpenTelemetry instrumentation. | false |
See .env.example for the full list of optional overrides.
Running the server
Local development
-
Build and run:
bun run rebuild
bun run start:stdio
bun run start:http
-
Run checks and tests:
bun run devcheck
bun run test
bun run lint:mcp
Docker
docker build -t devops-status-mcp-server .
docker run --rm -p 3010:3010 devops-status-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/devops-status-mcp-server. OpenTelemetry peer dependencies are installed by default โ build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
| Path | Purpose |
|---|
src/index.ts | createApp() entry point โ registers tools, resources, and inits services. |
src/config/ | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools/ | Tool definitions (*.tool.ts). |
src/mcp-server/resources/ | Resource definitions (*.resource.ts). |
src/services/cert/ | node:tls โ TLS handshake, X.509 parsing, expiry and protocol flagging. |
src/services/dns/ | node:dns โ multi-resolver DNS fan-out, propagation discrepancy detection. |
src/services/statuspage/ | Statuspage public API client with 60-second in-memory cache. |
src/services/status-adapters/ | Native-API adapters (Status.io, Slack, AWS Health, Google Cloud Service Health, Azure status, Firehydrant) + api_type dispatch, normalizing into the Statuspage shapes. |
src/services/vendor-registry/ | In-memory vendor registry loaded from src/data/vendor-registry.ts. |
src/data/ | Static vendor registry data file (vendor-registry.ts). |
tests/ | Vitest tests mirroring src/. |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches โ no
try/catch in tool logic
- Use
ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
- Register new tools and resources via the barrels in
src/mcp-server/*/definitions/index.ts
devops_check_certs and devops_check_dns use only Node.js stdlib โ add no external deps for these paths
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
License
Apache-2.0 โ see LICENSE for details.