NetLens

An MCP server for unobstructed web reading. It fetches any URL directly with
browser-like headers β past robots.txt and naive bot blocks β and returns the
full page as clean, ad-stripped Markdown, not a summary. Plus web search that
returns real links. Zero dependencies: pure Python standard library.
Built for AI Agents
AI agents constantly hit pages their built-in tools can't read. NetLens fixes the
three usual reasons a fetch comes back empty or useless:
| Native web tools | NetLens |
|---|
Honor robots.txt, so crawler-disallowed pages return nothing | Reads like the browser you'd open yourself β doesn't consult robots.txt |
Blocked by header/User-Agent bot filters (403/202 to non-browser clients) | Sends real browser headers via the system curl; commonly turns 403 β 200 |
| Return a summary of the page | Returns the full page content as Markdown |
| Leave ads, cookie banners, nav, and related-links chrome in the output | Strips boilerplate locally so only the content reaches your context |
It does not try to defeat JavaScript/Cloudflare challenge pages or CAPTCHAs β
that's out of scope by design. When a page is a hard block, the HTTP status is
surfaced honestly rather than faked.
Installation
npm (via npx):
{
"mcpServers": {
"netlens": {
"command": "npx",
"args": ["-y", "netlens-mcp"]
}
}
}
PyPI (via uvx):
{
"mcpServers": {
"netlens": {
"command": "uvx",
"args": ["netlens-mcp"]
}
}
}
Add either to your MCP client config (e.g. .mcp.json for Claude Code), then
restart the session so the tools load.
web_search
Search the web and return real result links (title, URL, snippet), parsed locally β
links, not summaries. Follow up with web_fetch to read a result.
| Argument | Type | Description |
|---|
query | string (required) | The search query |
limit | integer | Optional cap; default returns the full first page (~10) |
engine | string | auto (default), duckduckgo, bing, mojeek, searxng |
page | integer | Result page, 1-based. SearXNG only |
time_range | string | day, week, month, year. SearXNG only |
categories | string | e.g. it, science, news. SearXNG only |
The result reports which engine answered and what happened to any that were
skipped, so falling through to a different backend is visible rather than silent:
{
"query": "β¦",
"engine": "mojeek",
"results": [ β¦ ],
"engines_skipped": [ { "engine": "duckduckgo", "outcome": "HTTP 202" } ]
}
Outcomes are observations, not conclusions β HTTP 202 is what the server sent;
whether that is throttling, changed markup or genuinely no matches cannot be
determined from the response. With SearXNG configured, direct answers,
infoboxes and suggestions appear alongside the results.
A search fetches a single result page (~10 results), returned in full by default so
nothing at position 9/10 is dropped. There's no deep pagination β if the answer isn't in
the first page, refine the query.
web_fetch
Fetch any page and return its full content as clean Markdown.
| Argument | Type | Description |
|---|
url | string (required) | URL to fetch (scheme optional; https assumed) |
mode | string | article (main content only, default), full (whole body), raw (unconverted HTML), outline (heading structure) |
section | string | Return only this heading's content, plus anything nested under it |
links | string | inline (default) keeps link targets; none keeps link text but drops URLs |
max_chars | integer | Optional cap on returned characters (truncates with a note) |
Reading part of a long page
A large article can be tens of thousands of characters when you want one part of
it. mode="outline" returns its shape, and section returns just that piece β
on a large encyclopedia article that is 153,000 characters full, 1,000 as an
outline, and 7,400 for the section actually wanted.
web_fetch(url=β¦, mode="outline") β headings with each section's size
web_fetch(url=β¦, section="Gameplay") β that section and its subsections
On link-heavy pages the URLs themselves are a large share of the output β around 40%
of a big encyclopedia article β so links="none" roughly halves it when you only need
the prose. Links pointing back into the same page are always rendered as plain text.
If a page turns out to be a client-rendered shell, web_fetch says so rather than
returning an empty result as a success:
(note: Only 10 characters of readable text were found in 6,856 characters of HTML.
This page appears to be rendered client-side by JavaScriptβ¦)
Workflow: web_search to find pages, then web_fetch to read them.
Search engines
Search is a pluggable, selectable registry. In auto mode NetLens tries engines in
order and returns the first with results, so a rate-limit/challenge page on one
falls through to the next.
| Engine | Notes |
|---|
duckduckgo | Default; html.duckduckgo.com endpoint |
bing | Automatic fallback |
mojeek | Independent index; automatic fallback |
searxng | Self-hosted/public SearXNG JSON API β set NETLENS_SEARXNG_URL |
Pick per call with the engine argument, or set a default with
NETLENS_SEARCH_ENGINE.
Configuration
| Environment Variable | Default | Description |
|---|
NETLENS_SEARCH_ENGINE | auto | Default search backend |
NETLENS_SEARXNG_URL | β | SearXNG base URL, e.g. http://192.168.1.10:8888 |
NETLENS_SEARXNG_TIMEOUT | 8 | Seconds before an unreachable SearXNG is skipped |
NETLENS_USER_AGENT | Chrome UA | Override the request User-Agent |
NETLENS_MAX_BYTES | 10485760 | Cap on a single response; larger ones are truncated |
NETLENS_CACHE_TTL | 300 | Seconds to reuse a fetched page; 0 disables caching |
NETLENS_HOST_DELAY | 0.5 | Minimum seconds between requests to the same host |
NETLENS_REQUEST_TIMEOUT | 120 | Ceiling on a single tool call |
Self-hosted SearXNG
Point NETLENS_SEARXNG_URL at an instance with the JSON API enabled (search.formats
must include json in its settings.yml). It is then tried first in auto mode,
which removes the HTML scraping β and the rate limiting that comes with it β from the
common path.
Because it is tried first, it must fail fast when the box is off: the connect timeout
is bounded separately so an unreachable instance is skipped in a few seconds rather
than stalling every search, and the skip is reported in the result.
How it works
- Direct fetch. Requests go straight to the target site via the system
curl
(better TLS/HTTP-2/compression, so it looks like a real browser), falling back to
urllib. No third-party proxy or reader is involved.
- Local conversion. HTML β Markdown happens in-process with a hand-rolled
html.parser converter β headings, lists, links (relative URLs resolved), code
blocks, and GFM tables with colspan/rowspan.
- Content selection, not deletion. NetLens picks the page's main content region
β the HTML5 landmark (
<main> / <article> / [role=main]) when one exists,
otherwise the subtree holding the most prose relative to its link density β and
converts only that. Because it selects a winner rather than deleting anything
that matches a name pattern, extraction cannot silently return an empty page.
Nothing inspects CSS class or id names to decide what is content.
- Pruning by measurement. Within that region, blocks that are overwhelmingly
link anchors (navboxes, tag clouds, "more from this site" grids) are dropped
based on their link density. Non-rendering elements (
<script>, <style>, β¦),
explicitly hidden elements, and third-party ad-network slots (identified by
vendor names like adsbygoogle, which cannot collide with real prose) are
removed outright.
- Response charset is honored (from
Content-Type or <meta>), so non-UTF-8
pages don't come back garbled.
Usage from the CLI
The server is also a plain script β handy for testing before a client loads it:
python -m netlens_mcp.server search "http caching best practices"
python -m netlens_mcp.server fetch https://example.com/article
python -m netlens_mcp.server full https://example.com
python -m netlens_mcp.server raw https://example.com
python -m netlens_mcp runs the stdio MCP server; python -m netlens_mcp.server <cmd> runs the CLI.
Development
pip install -e ".[dev]"
python -m pytest
ruff check .
Requirements
- Python 3.10+ (and the system
curl, which ships with modern Windows/macOS/Linux;
falls back to urllib if absent)
License
Apache License 2.0 β see LICENSE and NOTICE.