@cyanheads/wikipedia-mcp-server
Search Wikipedia articles, read summaries and full text, target sections, find nearby pages, and list language editions via MCP. STDIO or Streamable HTTP.
6 Tools
Overview
Wikipedia content via the MediaWiki REST API and Action API. Search articles, read summaries or targeted sections, find geotagged pages near a coordinate, and list language editions from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|
wikipedia_search_articles | Full-text search across Wikipedia, returning ranked results with short descriptions, Wikidata QIDs, plain-text snippets, and page IDs, plus Wikipedia's spelling suggestion. |
wikipedia_get_summary | Short summary for any article β plain text, Wikidata QID, description, thumbnail URL, page type, canonical URL, revision, and coordinates. |
wikipedia_get_article | Full article or a targeted section as clean plain text, with section markers preserved, the canonical URL, and the revision it was read from. |
wikipedia_get_sections | Table of contents with section_index values for targeted section reads. |
wikipedia_search_nearby | Geotagged Wikipedia articles within a radius of a WGS 84 coordinate, sorted by distance, with short descriptions and Wikidata QIDs. |
wikipedia_get_languages | All language editions available for an article, with titles and URLs, or just the editions you ask for. |
Capability reference
wikipedia_search_articles tool
- Free-text query, ranked by relevance; returns plain-text snippets (HTML stripped), page IDs, and word counts
- Each result carries the article's short
description and Wikidata QID (wikibase_item) when it has them, from one follow-up lookup per page of results. That lookup is best-effort: if it fails, the results still come back without the two fields and the notice says so
- Enrichment
suggestion carries Wikipedia's spelling correction whenever it has one (einstien β einstein), and a zero-hit first page names it in the notice β re-run with it as query
limit is 1β50; offset pages further results β enrichment nextOffset signals more remain and is passed back as offset
- Wikipedia serves no result past the 10,000th for a query: an
offset at or beyond it fails with offset_too_large, and a page ending on the window carries enrichment truncated naming the matches no offset reaches β narrow the query to bring them into range
- An empty
query fails with empty_query; a whitespace-only query is a real search that simply matches nothing
language selects any Wikipedia edition (default en)
- Best when the exact article title is unknown, or to discover multiple articles on a topic
- Returns the REST summary extract β a truncated fragment from the start of the lead, not the whole lead β plus the Wikidata QID (
wikibase_item), short description, and thumbnail URL
url is the canonical article URL, and revision_id / last_modified name the revision the extract was read from β ?oldid=<revision_id> is a permanent link to it
- Superscripts and subscripts in the extract stay distinct from the digits beside them (
10Β²Β³, HβO), rendered the same way as on wikipedia_get_article
latitude / longitude are present for a geotagged article and pass straight to wikipedia_search_nearby, whose inputs carry those names; both are absent otherwise
- For the lead section in full, call
wikipedia_get_article with section_index: 0
page_type discriminates standard / disambiguation / no-extract β on disambiguation, re-query with wikipedia_search_articles for a more specific title
- Redirect pages are followed automatically
- Right tool for most encyclopedic "what is X?" lookups; use
wikipedia_get_article for full depth
wikipedia_get_article tool
- Without
section_index: full article with == Section == markers, unless it exceeds WIKIPEDIA_ARTICLE_OVERFLOW_BYTES (default 80,000 bytes) β then returns a section outline (truncated: true) pointing to wikipedia_get_sections plus a targeted section_index read
- With
section_index (from wikipedia_get_sections): returns that section plus every nested subsection, each heading above its own body
section_index: 0 is the lead section β the text above the first heading, returned under the title Introduction
- Both paths render code samples as fenced blocks with indentation intact and formulas as their TeX. Section reads also render data tables as pipe-delimited rows (header row,
| --- |, then one line per row; rowspan cells repeated) and infoboxes as label: value lines; a table over 40,000 rendered bytes leaves a [table omitted: N rows] marker. The full-article path carries no tables or infoboxes β upstream extracts strip them β so read the section for those
- Layout-only tables (multi-column lists, succession boxes) keep their content as ordinary text
- Superscripts and subscripts stay distinct from the digits beside them:
10Β²Β³, molβ»ΒΉ, HβO, or ^x / _x where a character has no Unicode form. An abbreviation's superscript stays joined, as the edition writes it in plain text (French XIXe siΓ¨cle, 1er, Mme)
- Every read returns
url (the canonical article URL) and revision_id (the revision the text was read from β ?oldid=<revision_id> is a permanent link to it); a full read, outline included, also returns last_modified, that revision's timestamp. For a redirect, all three name the target article. With WIKIPEDIA_BASE_URL set, a section read omits url, since the mirror's article path is unknown
- Page furniture β maintenance banners, sister-project and library-resource boxes, portal bars, spoken-article notices β is stripped, as are the editor-only preview warnings a section render emits; hatnotes are kept
- Redirect pages are followed automatically
- Returns section titles, heading levels, hierarchical numbering (e.g.
"2.1"), and section_index values
- The first entry is the lead:
index: 0, titled Introduction β Wikipedia's own table of contents starts at the first heading
section_index is the integer to pass to wikipedia_get_article for a targeted read
- Fails with
no_sections on a stub or very short article β read it with wikipedia_get_article instead
- Redirect pages are followed automatically
- Returns geotagged articles sorted ascending by distance, with coordinates,
distance_meters, and each article's short description and Wikidata QID (wikibase_item) when it has them
radius_meters: 10β10,000 (default 1000); limit: 1β500 (default 10) β no pagination past limit, so raise it or sweep narrower radii for full coverage
- Only articles carrying their own coordinate tag (GeoData, set on the article β not the Wikidata item's coordinate) are returned, and that tag places and measures each result. An article with a wrong tag appears where the tag puts it; its
description usually gives it away (Palazzo Bernardo Nani β "Palace on the Grand Canal, Venice" β 161 m from the Eiffel Tower)
- Enrichment
truncated flags when more articles matched than limit allowed; at limit: 500, Wikipedia's ceiling, a full page reports truncated and the notice points to narrower sweeps rather than a higher limit
- Returns each edition's
language_code, tool-usable edition_code (can differ, e.g. gsw vs als), article title, and URL
- Pass
edition_code β not language_code β as the language parameter on other tools
editions narrows the answer to the codes asked for, matched against both edition_code and language_code; requested codes with no article come back under missing, and total_languages stays the unfiltered count. A popular article lists hundreds of editions, so the filter is the difference between a 40 KB reply and a 1 KB one
- Fails with
no_other_languages when the article has no translations β a filter that matches nothing is a normal response with an empty list, not a failure
- Redirect pages are followed automatically;
source_title reports the resolved title
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
Wikipedia-specific:
- Dual API integration β MediaWiki REST API (
/api/rest_v1/) for summaries, Action API (/w/api.php) for search, full text, sections, geo search, and language links
- Retry and backoff on every required request, including the transient refusals (search too busy, rate-limited, read-only) the Action API returns inside an HTTP 200 body, with an upstream
Retry-After honored and each request's retries bounded at 30 s. The best-effort description lookup on search results gets one short attempt; User-Agent header per Wikimedia API policy
- Every read path renders HTML through one renderer to the same plain-text shape β
== Heading == markers, one list item per line, superscripts kept, code fenced: the full article from the Action API's HTML extract, a section from the parser's own HTML for that section, the summary from the REST extract_html. A section read additionally carries data tables, infoboxes, and the lists inside layout tables, none of which the extract carries
- Per-call
language parameter on every tool β all Wikipedia language editions accessible in a single session
- Language validation against a live edition registry built from the MediaWiki
action=sitematrix endpoint (cached 24h) β catches structurally valid but nonexistent editions before they cause timeouts
Agent-friendly output:
page_type on summaries discriminates standard / disambiguation / no-extract β no string parsing needed
wikibase_item (Wikidata QID) on summaries and on search and nearby results enables direct cross-referencing with wikidata-mcp-server
- Article text, snippets, and titles are backslash-escaped on the way into the markdown
content[] render, so an article that writes about markup or markdown syntax reads as itself instead of being interpreted by the client; structuredContent carries the same text unescaped. Fenced code blocks and table-row delimiters pass through unescaped, while table cells stay escaped
section_index on table-of-contents entries links directly to the targeted-read parameter on wikipedia_get_article, index 0 included
- Titles MediaWiki cannot name a page with β
< > [ ] { }, the | multi-title separator, percent escapes, magic tildes, relative paths β are refused before any request, with invalid_title; a trailing #fragment is accepted and resolves normally
- Recovery hints on every error type β callers get actionable next steps (e.g., "use
wikipedia_search_articles to find the correct title")
Getting started
Public Hosted Instance
A public instance is available at https://wikipedia.caseyjhand.com/mcp β no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"wikipedia-mcp-server": {
"type": "streamable-http",
"url": "https://wikipedia.caseyjhand.com/mcp"
}
}
}
Self-Hosted / Local
Add the following to your MCP client configuration file.
{
"mcpServers": {
"wikipedia-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/wikipedia-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"wikipedia-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/wikipedia-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"wikipedia-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"ghcr.io/cyanheads/wikipedia-mcp-server:latest"
]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
Prerequisites
- Bun v1.3.0 or higher (or Node.js v24+).
- No API keys required β Wikipedia's API is public.
Installation
- Clone the repository:
git clone https://github.com/cyanheads/wikipedia-mcp-server.git
- Navigate into the directory:
- Install dependencies:
- Configure environment (optional):
Configuration
| Variable | Description | Default |
|---|
WIKIPEDIA_USER_AGENT | User-Agent header sent with every Wikimedia API request. Customize for your deployment. | wikipedia-mcp-server/0.2.5 (https://github.com/cyanheads/wikipedia-mcp-server) |
WIKIPEDIA_BASE_URL | Optional single-instance override. Unset (default): compose per-language hosts, language selects the edition per call. Set to a full base URL (e.g. a private MediaWiki mirror): route every call at that one fixed host β language no longer varies it. | (unset) |
WIKIPEDIA_ARTICLE_OVERFLOW_BYTES | Byte budget above which a full-article read (wikipedia_get_article without section_index) returns a section outline instead of the full text. Tuned for this domain β ordinary articles stay whole; only genuine mega-articles (World War II ~86 KB, United States ~94 KB) outline. Section-targeted reads are never affected. | 80000 |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for HTTP server. | 3010 |
MCP_SESSION_MODE | HTTP session mode: stateless, stateful, or auto (which resolves to stateful). The Docker image ships stateless. | auto |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (RFC 5424). | info |
LOGS_DIR | Directory for log files (Node.js only). | <project-root>/logs |
OTEL_ENABLED | Enable OpenTelemetry instrumentation (spans, metrics, completion logs). | false |
See .env.example for the full list of optional overrides.
Running the server
Local development
-
Build and run:
bun run rebuild
bun run start:stdio
bun run start:http
-
Run checks and tests:
bun run devcheck
bun run test
bun run lint:mcp
Docker
docker build -t wikipedia-mcp-server .
docker run --rm -p 3010:3010 wikipedia-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/wikipedia-mcp-server. OpenTelemetry peer dependencies are installed by default β build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
| Directory | Purpose |
|---|
src/index.ts | createApp() entry point β registers tools and inits the Wikipedia service. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts) β one file per tool. |
src/services/wikipedia | WikipediaService β REST API + Action API client with retry/backoff and language validation. |
tests/ | Unit and integration tests mirroring src/. |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches β no
try/catch in tool logic
- Use
ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
- Register new tools in
src/mcp-server/tools/definitions/index.ts
- Wrap external API calls: validate raw β normalize to domain type β return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
License
Apache-2.0 β see LICENSE for details.