Unified MCP server for Kubernetes, ArgoCD, Prometheus, PagerDuty, and Loki
The io.github.NotHarshhaa/devops-mcp server is a unified Model Context Protocol (MCP) server for DevOps engineers. It enables any MCP-compatible AI agent to query and manage Kubernetes, ArgoCD, Prometheus, and PagerDuty. It is described as open source and positioned as MCP-compatible.
π οΈ Key Features
Unified MCP server for DevOps
Kubernetes support
ArgoCD support
Prometheus support
PagerDuty support
π Use Cases
Query Kubernetes from an MCP-compatible AI agent
Manage ArgoCD via MCP-compatible tooling
Work with Prometheus through MCP
Integrate PagerDuty via MCP
β‘ Developer Benefits
MCP-compatible for use with βany MCP-compatible AI agentβ
Developer-facing categories include topics like argocd, kubectl, kubernetes, and model-context-protocol
Supports commonly used DevOps tooling in a single server
β οΈ Limitations
Provided description does not mention tool count, Loki, or Loki-specific capabilities (only topics include loki)
Unified MCP server for DevOps engineers β query and manage Kubernetes, ArgoCD, Prometheus, and PagerDuty from any MCP-compatible AI agent.
What is this?
devops-mcp is an open source Model Context Protocol server that gives AI agents (Claude, etc.) real-time read and write access to your infrastructure stack β all from a single install.
Instead of copy-pasting kubectl output into a chat window, you can ask:
"Why is the payments deployment in CrashLoopBackOff?""What changed in the last ArgoCD sync for the auth app?""Show me the p99 latency for the API gateway over the last hour.""Who's on call right now and what incidents are open?""Debug the payments service - what's wrong with it?"
...and get live answers, sourced directly from your cluster and tooling.
Providers included:
Prefix
Provider
Transport
k8s__*
Kubernetes (via kubeconfig or in-cluster SA)
client-go
argo__*
ArgoCD
REST API
prom__*
Prometheus
HTTP API (PromQL)
pd__*
PagerDuty
REST API v2
helm__*
Helm
CLI (helm binary)
devops__*
Cross-provider incident debugging
Aggregates all providers
logs__*
Loki
HTTP API (LogQL)
Quick start
Claude Desktop (stdio β recommended)
Add this to ~/.config/claude/claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json):
npx @notharshhaa/devops-mcp
# or clone and run:
git clone https://github.com/NotHarshhaa/devops-mcp
cd devops-mcp
npm install
cp .env.example .env# fill in your values
npm run dev
Configuration
All config is via environment variables. Only set the ones for providers you actually use β providers with missing config are silently skipped.
env
# ββ Kubernetes ββββββββββββββββββββββββββββββββββββββββββββββββ
KUBECONFIG=/home/user/.kube/config # omit to use in-cluster service account
K8S_CONTEXT=my-prod-context # optional: pin a specific context
K8S_ALLOWED_NAMESPACES=default,backend # optional: restrict namespace access
# ββ ArgoCD βββββββββββββββββββββββββββββββββββββββββββββββββββ
ARGOCD_SERVER=https://argocd.company.com
ARGOCD_TOKEN=eyJhbGci... # argocd account generate-token
# ββ Prometheus βββββββββββββββββββββββββββββββββββββββββββββββ
PROMETHEUS_URL=http://prometheus:9090
PROMETHEUS_BEARER_TOKEN= # optional: for authenticated Prometheus
# ββ PagerDuty ββββββββββββββββββββββββββββββββββββββββββββββββ
PAGERDUTY_TOKEN=your-api-v2-token
# ββ Loki βββββββββββββββββββββββββββββββββββββββββββββββββββ
LOKI_URL=http://loki.monitoring:3100
LOKI_TOKEN=your-loki-token
# ββ Stateless Streamable HTTP ββββββββββββββββββββββββββββββββ
# For stdio mode (default): no transport config needed
MCP_HTTP_HOST=127.0.0.1 # use 0.0.0.0 inside a container
PORT=3000
MCP_AUTH_TOKEN=shared-secret # optional static Bearer token
MCP_REQUEST_STATE_SECRET=32+-byte-secret # optional MRTR signing key shared by all replicas
MCP_ALLOWED_HOSTS=localhost,127.0.0.1 # required with non-loopback binding
MCP_ALLOWED_ORIGINS= # optional browser Origin hostname allowlist
MCP_CACHE_TTL_MS=60000 # discovery/tools catalog TTL; 0 disables caching
# ββ Safety βββββββββββββββββββββββββββββββββββββββββββββββββββ
DEVOPS_MCP_DRY_RUN=false # true = block all mutations globally
DEVOPS_MCP_AUDIT_LOG=/var/log/devops-mcp-audit.jsonl
Tool reference
All tools follow a three-tier safety model:
Read β safe, no side effects, no confirmation needed
Mutate β defaults to dry_run: true; set dry_run: false to execute
Destructive β requires confirm: true, or a 2026-07-28 client can complete the server's interactive MRTR confirmation
Kubernetes (k8s__*)
Tool
Tier
Description
k8s__list_pods
read
List pods with status, restarts, node, age
k8s__get_pod_logs
read
Tail or stream logs from a pod container
k8s__describe_resource
read
Full describe for any resource type
k8s__get_events
read
Cluster or namespace events, filterable by reason
k8s__list_deployments
read
Deployments with replica counts and rollout health
Overall assessment: Summary of issues and positive indicators
Why this matters:
Instead of raw PromQL numbers that require interpretation, this tool provides actionable insights that AI agents can use directly in responses, making monitoring data actually useful for incident investigation and communication.
Loki (logs__*)
Tool
Tier
Description
logs__get_recent_errors
read
Get recent error logs from Loki for debugging incidents
logs__search
read
Search logs in Loki with custom query for root cause analysis
Example usage:
bash
# Get recent error logs
logs__get_recent_errors(service="payments", namespace="default", minutes=30, limit=50)
# Search logs with custom query
logs__search(query='{service="payments"} |= level="error"', limit=100)
Why this matters:
Metrics tell what: Prometheus shows you that latency increased or error rate crossed SLO
Logs tell why: Loki shows you the actual error messages, stack traces, and context around failures
Complete debugging: Without logs, you can see that something is broken but not understand the root cause
Output format:
Structured log entries with timestamp, message, service, namespace, and extracted log levels
Error count summaries and filtering
Raw LogQL results for detailed analysis
This makes incident investigation complete by combining the "what" (metrics) with the "why" (logs).
PagerDuty (pd__*)
Tool
Tier
Description
pd__list_incidents
read
Open incidents with severity, status, assignee
pd__get_incident
read
Full detail with alerts, notes, timeline
pd__who_is_oncall
read
Current on-call per schedule or escalation policy
pd__list_services
read
All services with integration keys and status
pd__get_log_entries
read
Audit log for an incident (all state changes)
pd__acknowledge_incident
mutate
Preview acknowledgement; set dry_run: false to execute
pd__add_note
mutate
Preview appending a note; set dry_run: false to execute
pd__escalate_incident
destructive
Escalate to a different policy β requires direct or interactive confirmation
pd__summarize_incident
read
π¨ Incident auto-summary - what happened, affected services, probable root cause, current status
pd__summarize_incident
Example usage:
bash
# Get an auto-summary of an incident
pd__summarize_incident(id="ABC123")
What it outputs:
What happened: Incident title, description, severity, urgency, status, creation time, and duration
Affected services: Service name, ID, and current status
Probable root cause: Analysis of trigger alerts and log entries to identify likely causes
Current status: Current incident state, assignees, acknowledgements, and notes count
Output format:
json
{"what_happened":{"title":"API Gateway High Error Rate","description":"5xx error rate exceeded 5% threshold","severity":"high","urgency":"high","status":"acknowledged","createdAt":"2025-01-15T10:30:00Z","updatedAt":"2025-01-15T11:45:00Z","duration":"1h 15m"},"affected_services":[{"id":"P123456","name":"API Gateway","status":"critical"}],"probable_root_cause":"Triggered by: High 5xx error rate from API Gateway pods","current_status":{"status":"acknowledged","lastUpdated":"2025-01-15T11:45:00Z","assignees":["john.doe@company.com"],"acknowledgements":2,"notes":3}}
Why this matters:
Instead of manually piecing together incident details from multiple API calls, this tool provides a comprehensive, human-readable summary perfect for:
Demos: Shows AI's ability to understand and summarize complex incident data
Real-world use: Quickly understand incident impact without digging through raw data
Communication: Share concise incident summaries with stakeholders
Helm (helm__*)
Tool
Tier
Description
helm__list_releases
read
List Helm releases with status, chart, app version
helm__get_status
read
Full status of a Helm release
helm__get_values
read
User-supplied or computed values for a release
helm__get_history
read
Revision history of a release
helm__rollback
mutate
Rollback to a previous revision (dry-run by default)
Requirements: Helm CLI binary must be available in PATH.
Example usage:
bash
# List all releases in a namespace
helm__list_releases(namespace="production")
# Check what values a release is using
helm__get_values(name="api-gateway", all_values=true)
# Rollback after a bad deploy
helm__rollback(name="api-gateway", revision=5, dry_run=false)
Cross-Provider Debugging (devops__*)
Tool
Tier
Description
devops__debug_service
read
π₯ Cross-provider incident debugging - aggregates Kubernetes, ArgoCD, Prometheus, and PagerDuty data to diagnose service issues in one command
devops__explain_change
read
π§ Explain what changed - combines ArgoCD history, Kubernetes rollout history, and Prometheus anomaly window to identify cause of issues
Correlation analysis that links deployments to metric changes
Summary with root cause hypothesis
Problem it solves:"Everything was working yesterday⦠what changed?"
This tool answers that question by correlating deployment events with metric anomalies, helping you quickly identify whether a recent deployment, config change, or external factor caused the issue.
devops__runbook
Example usage:
bash
# Diagnose a crashlooping service
devops__runbook(symptom="crashloop", service="payments", namespace="default")
# Investigate high latency
devops__runbook(symptom="high-latency", service="api-gateway")
Supported symptoms:
Symptom
What it checks
crashloop
Pod status β logs (tail 50) β BackOff events β deployment health
Output: Structured JSON with steps_executed[], findings[], and recommended_actions[].
devops__health_report
Example usage:
bash
# Get a full cluster health assessment
devops__health_report(namespace="production")
What it gathers:
Kubernetes: Unhealthy pods, deployments not at desired replicas
Prometheus: Count of firing alerts
ArgoCD: Out-of-sync and unhealthy applications
PagerDuty: Open incident count
Output: Overall status (healthy / degraded / critical), per-provider sections, and summary. Perfect for morning standup checks or shift handoffs.
Deployment options
stdio (recommended for local use)
The MCP host launches devops-mcp as a subprocess and communicates over stdin/stdout. Zero network config. Auth comes from the local environment (kubeconfig, env vars). Process lifecycle tied to Claude Desktop.
bash
npx @notharshhaa/devops-mcp
# or with env vars
KUBECONFIG=~/.kube/config npx @notharshhaa/devops-mcp
Stateless Streamable HTTP (shared deployments)
The HTTP entry serves MCP at POST /mcp using the 2026-07-28 stateless protocol. Each request gets a fresh MCP server instance, so requests can land on any replica without session affinity or shared protocol state. The same endpoint also accepts 2025-era Streamable HTTP clients in stateless compatibility mode.
Connect clients to http://127.0.0.1:3000/mcp. For a container or remote service, set MCP_HTTP_HOST=0.0.0.0 and configure MCP_ALLOWED_HOSTS with the public/proxy hostnames. Put the service behind TLS for team use.
The deprecated devops-mcp-sse binary remains as a temporary alias for the HTTP entry, but /sse, /message, and /ws now return HTTP 410. Legacy HTTP+SSE and non-standard WebSocket clients must migrate to Streamable HTTP.
2026-07-28 behavior
No initialize requirement or Mcp-Session-Id on modern requests.
server/discover, per-request client metadata, and MCP-Protocol-Version are handled by the official SDK.
Mcp-Method and Mcp-Name headers are validated against the JSON-RPC body for gateway routing and authorization.
server/discover and tools/list advertise deterministic, public cache hints using MCP_CACHE_TTL_MS (default 60 seconds).
Destructive tools accept confirm: true; modern clients may instead complete an MRTR interactive confirmation. Confirmation state is HMAC-signed, expires after five minutes, and is bound to the exact tool arguments. Global dry-run blocks before prompting.
MCP_REQUEST_STATE_SECRET must be the same on every replica for MRTR retries to land anywhere. When omitted, the server derives the key from MCP_AUTH_TOKEN, or uses a process-local random key when no token is configured.
2025-era Streamable HTTP and stdio clients remain supported. Sessionful HTTP and legacy HTTP+SSE are not.
The built-in MCP_AUTH_TOKEN is a static bearer-token gate, not an OAuth authorization server. For Internet-facing deployments, terminate TLS and enforce your organizationβs OAuth/OIDC policy at a gateway or integrate a dedicated identity provider; do not use deprecated Dynamic Client Registration for new deployments.
A minimal docker-compose.yml is available in examples/.
Security model
devops-mcp is designed for internal use inside a trusted network. That said:
Kubernetes: Uses standard kubeconfig via @kubernetes/client-node. Supports exec plugins (AWS EKS, GKE). In-cluster: auto-mounts SA token. Add RBAC rules scoped to your desired permissions β run devops-mcp under a dedicated ServiceAccount with minimal verbs. Context selection is fixed by K8S_CONTEXT at startup; the tool only previews changes because runtime switching would affect other callers.
ArgoCD: Generate a long-lived token: argocd account generate-token --account devops-mcp. Create a dedicated account in argocd-cm with apiKey capability and a role limited to read + sync.
Prometheus: Usually unauthenticated inside a cluster. If using Grafana Mimir or Thanos with auth, pass a Bearer token. All tools are read-only so minimal permissions are needed.
PagerDuty: Create a dedicated API key in PagerDuty β API Access β Create New API Key. Use Full Access if you want acknowledge/escalate tools; Read-only if you want a safe-only mode.
Mutations are dry-run by default. Every mutating tool defaults dry_run: true. The AI must explicitly pass dry_run: false β it won't do this unless the user clearly requests an action.
Destructive tools require confirmation. Pass confirm: true directly, or use a 2026-07-28 client that supports the server's MRTR confirmation request. DEVOPS_MCP_DRY_RUN=true blocks execution even after confirmation.
Audit log. Set DEVOPS_MCP_AUDIT_LOG to a file path. Every tool call is written as a JSONL line with timestamp, tool name, parameters, and outcome. Mutations and destructive calls are flagged.
Global dry-run mode. Set DEVOPS_MCP_DRY_RUN=true to block every executing mutation, even when a caller passes dry_run: false. Safe previews remain available β useful for read-only team deployments.