Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
io.github.bamchi/scrapi MCP Server
This MCP server provides web scraping for AI agents. It converts input URLs into clean Markdown/Text that is formatted for LLM consumption, and it includes an anti-bot bypass. The server is identified as io.github.bamchi/scrapi and exposes 5 tools.
π οΈ Key Features
Web scraping for AI agents
URL-to-LLM-ready output
Converts URLs to clean Markdown/Text
Anti-bot bypass
Toolset size: 5 tools
π Use Cases
Feeding external webpages into LLM agents as Markdown/Text
Extracting and normalizing website content for downstream AI workflows
Scraping sources where anti-bot measures apply
β‘ Developer Benefits
Output suitable for LLM agents (clean Markdown/Text)
Direct transformation from URL to agent-consumable format
Consistent interface via 5 provided tools
β οΈ Limitations
The description does not specify supported sites, rate limits, or output coverage beyond βclean Markdown/Text.β
Captured live from the server via tools/list.
scrape_url
Scrapes a webpage and returns the content in AI-readable Markdown format. Can access blocked sites through browser rendering.
Parameters2
url
string
required
The URL of the webpage to scrape
format
string
optional
Output format: markdown (default) or text
Raw schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"format": "uri",
"description": "The URL of the webpage to scrape"
},
"format": {
"type": "string",
"enum": [
"markdown",
"text"
],
"default": "markdown",
"description": "Output format: markdown (default) or text"
}
},
"required": [
"url"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}
scrape_urls
Scrapes multiple webpages in parallel and returns the content in AI-readable Markdown format. Can access blocked sites through browser rendering.
β‘ Fast & Reliable β Built on 8+ years of web scraping expertise, 1,900+ production crawlers, and battle-tested anti-bot handling.
What is this?
An MCP (Model Context Protocol) server that lets AI agents fetch and read web pages. Simply give it a URL, and it returns clean, LLM-ready content β fast.
Before: AI can't read web pages directly After: "Summarize this article" just works β¨
# Clone the repository
git clone https://github.com/bamchi/scrapi-mcp-server.git
cd scrapi-mcp-server
# Install dependencies and build
npm install && npm run build
Note: Claude Desktop requires the mcp-remote proxy for HTTP connections.
Self-host the HTTP server (advanced)
Run your own instance instead of using the hosted endpoint:
bash
SCRAPI_API_KEY=your-api-key npx -y -p @scrapi.ai/mcp-server scrapi-http
# or from source:
SCRAPI_API_KEY=your-api-key node dist/http.js
The server starts at http://localhost:3000 with the MCP endpoint at /mcp. Configure with PORT and HOST environment variables. Replace the URL in the client configurations above with your self-hosted URL (e.g. http://localhost:3000/mcp).
Health check:GET http://localhost:3000/health
Step 3: Restart Your AI Client
Claude Desktop: Fully quit (Cmd+Q on macOS, Alt+F4 on Windows) and reopen
Claude Code: Restart the session
Cline: Restart VS Code
Cursor: Restart the editor
You should see the MCP server connection indicator.
Available Tools
scrape_url
Scrapes a webpage and returns AI-readable content.
# Article Title> Author: John Doe | Published: 2024-01-15## Introduction
This is the main content of the article, converted to clean markdown...
## Key Points- Point 1: Important detail
- Point 2: Another insight
- [Related Link](https://example.com/related)
Text Output:
text
Article Title
Author: John Doe | Published: 2024-01-15
Introduction
This is the main content of the article, converted to plain text...
Key Points
- Point 1: Important detail
- Point 2: Another insight
scrape_urls
Scrapes multiple webpages in parallel and returns AI-readable content.
[{"url":"https://example.com/page1","content":"Page 1 Title\n\nThis is the content of page 1..."},{"url":"https://example.com/page2","content":"Page 2 Title\n\nThis is the content of page 2..."}]
scraper_server_status
Check the status of all ScraperServer instances. Shows server health, circuit breaker state, failure counts, and timing info.
Parameters: None
Example:
json
{}
Output:
markdown
## ScraperServer Status
Total: 3 | Available: 2
| Name | OS | Status | Failures | Last Success | Last Failure |
|------|----|--------|----------|--------------|--------------|
| pluto | linux | OK | 0 | 01/30 14:23:05 | - |
| mars | mac | FAIL | 2 | 01/29 10:00:00 | 01/30 13:55:12 |
| venus | linux | OPEN | 3 | 01/28 09:00:00 | 01/30 12:00:00 |
### Issues-**mars**: Connection refused - connect(2)
-**venus**: Circuit breaker open until 01/30 12:30:00
-**venus**: Net::ReadTimeout
Status values:
Status
Description
OK
Server is healthy
FAIL
Server is unhealthy
OPEN
Circuit breaker open (isolated for 30 min)
N/A
Not yet checked
get_usage
Check your API usage and remaining credits.
Parameters: None
Example:
json
{}
Output:
markdown
## MCP Credits
| Item | Value |
|------|-------|
| Plan | starter |
| Subscription Credits | 1,500 |
| Purchased Credits | 200 |
| Total Remaining | 1,700 |
| Period End | 2026-03-01 |
get_billing
Retrieve detailed billing information including subscription, plans, daily usage, and spending limits.
Parameters:
Name
Type
Required
Description
action
string
Yes
subscription, plans, daily_usage, or spending_limits
start_date
string
Start date for daily_usage (YYYY-MM-DD, default: 30 days ago)
end_date
string
End date for daily_usage (YYYY-MM-DD, default: today)
Example β Current subscription:
json
{"action":"subscription"}
markdown
## MCP Subscription
| Item | Value |
|------|-------|
| Plan | starter (Starter) |
| Status | active |
| Monthly Credits | 2,000 |
| Price | $19.00/mo |
| Rate Limit | 30 RPM |
| Burst Limit | 5 concurrent |
| Period End | 2026-03-01 |
User: Summarize this article: https://news.example.com/article/12345
Claude: [calls scrape_url]
Here's a summary of the article:
## Key Points
- Point 1: ...
- Point 2: ...
- Point 3: ...
Example 2: Fetch Page Content
code
User: Get the content from https://example.com/data
Claude: [calls scrape_url]
# Page Title
> Source: https://example.com/data
The page content is returned in clean Markdown format...
Example 3: Research Competitor Pricing
code
User: What's the pricing on https://competitor.com/product/abc
Claude: [calls scrape_url]
Here's the pricing information:
- **Product**: ABC Premium
- **Regular Price**: $99.00
- **Sale Price**: $79.00 (20% off)
Example 4: Read API Documentation
code
User: Read https://docs.example.com/api/v2 and write integration code
Claude: [calls scrape_url]
I've analyzed the API documentation. Here's the integration code:
// api-client.ts
export class ExampleApiClient {
private baseUrl = 'https://api.example.com/v2';
async getData(): Promise<Response> {
// ...
}
}
How It Works
code
βββββββββββββββββββ
β User β
β "Summarize this β
β URL for me" β
ββββββββββ¬βββββββββ
β
βΌ
βββββββββββββββββββ
β Claude Desktop β
β / Cursor β
ββββββββββ¬βββββββββ
β
βΌ
βββββββββββββββββββ βββββββββββββββββββ
β MCP Server ββββββΊβ Scrapi API β
β (scrape_url) β β (format param) β
ββββββββββ¬βββββββββ ββββββββββ¬βββββββββ
β β
βββββββββββββββββββββββββ
β Markdown/Text Response
βΌ
βββββββββββββββββββ
β AI Response β
β (Summary, etc.) β
βββββββββββββββββββ
Why Scrapi?
Built by the team behind Scrapi, with 8+ years of web scraping experience:
β 1,900+ production crawlers
β JavaScript rendering support
β Anti-bot handling
β 99.9% uptime
Troubleshooting
"API key is required"
Make sure your API key is provided via one of these methods:
Environment variable: Set SCRAPI_API_KEY in your configuration
CLI argument: Pass --api-key your-key in the args
"Invalid API key"
Verify that your API key is correct and active in your Scrapi dashboard.
npx using an old cached version
If you upgraded but still see old behavior, clear the npx cache:
bash
npx clear-npx-cache
MCP Server not connecting
Ensure Node.js 20+ is installed
Try running node /absolute/path/to/scrapi-mcp-server/dist/index.js manually to check for errors
Fully quit Claude Desktop (Cmd+Q on macOS, Alt+F4 on Windows) and restart
Check Settings > Developer to verify the server is listed
Developer tab not visible
Update Claude Desktop to the latest version: Claude menu β "Check for Updates..."