Agentโ™ฅ๏ธŽAge
Catalog

TheCrawler

Official

by manchittlab ยท TypeScript

Universal web scraper with LLM-ready markdown, RAG chunking, PDF/DOCX support.

io.github.manchittlab/thecrawler โ€” MCP Server for TheCrawler

This MCP server provides a universal web scraper that outputs LLM-ready markdown, supports RAG chunking, and can handle PDF/DOCX inputs. It targets structured extraction workflows and includes a way to validate whether URLs are ready for an extraction contract before using LLM tokens.

๐Ÿ› ๏ธ Key Features

  • Universal web scraping
  • LLM-ready markdown output
  • RAG chunking
  • PDF/DOCX support
  • Validated extraction contracts readiness

๐Ÿš€ Use Cases

  • Scrape web pages for AI/LLM pipelines
  • Run LLM-powered structured extraction
  • Diagnose URL readiness for extraction contracts
  • Produce chunked content for retrieval systems

โšก Developer Benefits

  • Avoid spending LLM tokens on non-ready URLs via readiness checks
  • Use structured extraction with validated contracts
  • Start with a dryRun: true test on Apify

โš ๏ธ Limitations

  • Requires successful execution for any Apify per-page cost ($0.005 per successfully scraped page)

Topics

agplapifycheeriocrawlerllmmarkdownmcpmcp-servermodel-context-protocolnodejsplaywrightragscrapertypescriptweb-scraping

Related servers

More in Search & Web