Engineering log of self-hosted AI on NVIDIA DGX Spark (GB10/SM121A). 60+ articles indexed.
Sovereign AI MCP (org.sovgrid/self-hosted-ai)
This MCP server provides an engineering log focused on self-hosted AI on an NVIDIA DGX Spark (GB10/SM121A). It indexes 60+ articles and is categorized under Model Context Protocol (MCP). The repository includes a CI workflow badge and lists licensing and content attribution via MIT and CC BY-SA 4.0.
🛠️ Key Features
Engineering log for self-hosted AI on NVIDIA DGX Spark (GB10/SM121A)
60+ articles indexed
MCP server, with 4 tools reported
🚀 Use Cases
Reference material for deploying or operating self-hosted AI on DGX Spark
Developer lookup for MCP-related “model context protocol” and MCP server usage
⚡ Developer Benefits
Topics include ai-agents, fastmcp, mcp, mcp-server, and python
Index supports quick discovery via categorized content
⚠️ Limitations
Description excerpt does not enumerate individual tool names or exact MCP capabilities beyond the reported tool count and indexed articles
Search the Sovereign AI Blog for articles matching a natural language query,
optionally filtered by tag and sorted by relevance or date.
Behaviour matrix:
- query='', sort=* -> list newest-first, optionally tag-filtered
- query!='', sort=relevance -> TF-IDF ranked, optionally tag-filtered
- query!='', sort=date_desc -> TF-IDF filtered (score > 0.001), then sorted by date
Pure read-only, deterministic for a given KB snapshot.
Parameters4
query
string
optional
Natural language search query (e.g. 'flashinfer OOM on GB10'). Multi-word queries are tokenized and TF-IDF ranked. Pass empty string to list articles without ranking by relevance.
tag
any
optional
Optional tag filter (e.g. 'setup', 'fixes', 'strategy'). Only articles with this tag are considered. Use list_tags to discover available tags.
sort
string
optional
Result ordering. 'relevance' uses TF-IDF score (default for non-empty query). 'date_desc' sorts newest first (default behaviour when query is empty). When query is empty, 'relevance' is treated as 'date_desc'.
n
integer
optional
Maximum number of results to return
Raw schema
{
"type": "object",
"properties": {
"query": {
"default": "",
"description": "Natural language search query (e.g. 'flashinfer OOM on GB10'). Multi-word queries are tokenized and TF-IDF ranked. Pass empty string to list articles without ranking by relevance.",
"title": "Query",
"type": "string"
},
"tag": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Optional tag filter (e.g. 'setup', 'fixes', 'strategy'). Only articles with this tag are considered. Use list_tags to discover available tags.",
"title": "Tag"
},
"sort": {
"default": "relevance",
"description": "Result ordering. 'relevance' uses TF-IDF score (default for non-empty query). 'date_desc' sorts newest first (default behaviour when query is empty). When query is empty, 'relevance' is treated as 'date_desc'.",
"enum": [
"relevance",
"date_desc"
],
"title": "Sort",
"type": "string"
},
"n": {
"default": 5,
"description": "Maximum number of results to return",
"maximum": 20,
"minimum": 1,
"title": "N",
"type": "integer"
}
},
"title": "search_blogArguments"
}
get_article
Retrieve the full content of a blog article by its slug.
Returns the article body (Markdown) plus metadata. If the slug does not
match any article, returns an Article with `error='article_not_found'`
and other fields at their defaults.
Parameters1
slug
string
required
Article slug as returned by search_blog (e.g. 'setup-mistral-sglang-setup'). Lower-case, hyphenated.
Validate an SGLang configuration for NVIDIA DGX Spark (GB10/SM121A).
Pure pattern-matching against known failure modes documented in the
Sovereign AI Blog. No inference, no external calls. Returns critical
issues, non-fatal warnings, and a recommended baseline config.
All parameters are optional; supply only what you have. With no inputs
you get the recommended config and a 'unknown' verdict.
Parameters6
attention_backend
string
optional
SGLang --attention-backend value (e.g. 'flashinfer', 'triton'). Empty string = skip this check.
mem_fraction
number
optional
SGLang --mem-fraction-static value (e.g. 0.88). 0.0 = skip this check.
cuda_graph_max_bs
integer
optional
SGLang --cuda-graph-max-bs value. 0 = skip this check.
image_tag
string
optional
Docker image tag in use (e.g. 'lmsysorg/sglang:latest', 'lmsysorg/sglang:v0.4.0'). Empty = skip.
List all topic tags used across the Sovereign AI Blog corpus, with article
counts. Use this to browse the topic space before calling search_blog with
a tag filter.
Parameters1
sort
string
optional
Result ordering. 'count_desc' lists most-used tags first (default). 'alpha' sorts alphabetically.
Training data on niche hardware (GB10, SM121A, SGLang on ARM64) is sparse and stale. This MCP gives agents direct, structured access to 180+ articles documenting actual setups, fixes, and benchmarks. If you're building or debugging on similar stacks, your agent can pull verified, version-current information instead of hallucinating.
The corpus covers SGLang and vLLM patches for GB10, voxtral and TTS pipelines on ARM64, KV-cache and quantization tradeoffs, podcast-grade audio generation, MCP server design, knowledge-base construction, and the operational side of running it all on a hardened European VPS.
Tools
Tool
Purpose
search_blog(query, tag?, sort?, n?)
TF-IDF full-text search. Optional tag filter, sort by relevance or date_desc. Empty query lists newest articles. Returns ranked SearchResult items with quality score, style, slug, and excerpt.
list_articles(tag?, sort?, limit?, offset?)
Direct article listing with pagination, no TF-IDF overhead. Filter by tag, sort by date_desc, date_asc, title_asc, or quality_desc. Returns ArticleSummary items.
list_tags(sort?)
List all topic tags across the corpus with article counts. Sort by count_desc (default) or alpha. Use to discover the topic space before filtering search_blog.
get_article(slug)
Fetch full article body and frontmatter by slug. Returns markdown content plus tags, quality score, publish date.
diagnose_sglang(error_message)
Pattern-match a runtime error against a curated rule set for SGLang on GB10/SM121A. Returns matched fixes with links to setup articles.
All tools are read-only, idempotent, and declared with ToolAnnotations so MCP clients can calibrate retry policy and trust signals. Inputs use Pydantic Annotated[type, Field(description=...)] so parameter docs reach agents through introspection. Outputs are typed BaseModel shapes — schemas are real, not vacuous dicts.
Quick start
With Claude Code
bash
claude mcp add sovereign-ai --transport http https://mcp.sovgrid.org/self-hosted-ai
git clone https://github.com/cipherfoxie/sovereign-mcp.git
cd sovereign-mcp
uv sync
uv run uvicorn src.main:app --host 127.0.0.1 --port 8002
Docker
bash
git clone https://github.com/cipherfoxie/sovereign-mcp.git
cd sovereign-mcp
docker build -t sovereign-mcp .
docker run -p 8002:8002 sovereign-mcp
The repo ships a placeholder data/knowledge-base.json (zero articles, valid schema) so the server starts and answers MCP introspection cleanly out-of-the-box. To populate it with real content, generate from the sovgrid.org blog source using scripts/generate_knowledge_base.py, or build your own KB matching the schema in src/knowledge.py. Or just use the live endpoint at https://mcp.sovgrid.org/self-hosted-ai.
FastMCP 1.27+ with Streamable HTTP transport at path /self-hosted-ai
DNS rebinding protection via TransportSecuritySettings: only allows requests with Host: mcp.sovgrid.org (or localhost for healthchecks)
Health endpoint at /health returns article count and KB generation timestamp
Knowledge base is a flat JSON file generated from blog Markdown content; loaded at startup, queried via TF-IDF for search_blog
The server is stateless. All blog content is already public (CC BY-SA 4.0). No PII, no auth tokens, no secrets.
Operations
Live deployment runs on a privacy-focused European VPS via Docker, fronted by Caddy with TLS. Server logs flow into a privacy-respecting analytics pipeline (Caddy JSON access logs, no client-side tracking, no JS pixels).