ESLint for AI search: audits AI-crawler access, llms.txt, schema and citability.
io.github.iliasabk/geolint β Model Context Protocol (MCP) Server
This MCP server provides an ESLint-based linter for AI search. It audits AI-crawler access, including llms.txt and schema, and addresses citability. The project is positioned for generative-engine optimization workflows and integrates with TypeScript/CLI and GitHub Actions environments.
π οΈ Key Features
ESLint for AI search
Audits AI-crawler access
Checks llms.txt
Audits schema and citability
π Use Cases
Generative-engine optimization quality checks
Validating AI crawler access configuration
Reviewing llms.txt and schema-related settings
Supporting SEO/crawl readiness for AI search
β‘ Developer Benefits
Works within TypeScript-oriented tooling
Fits CLI and GitHub Actions pipelines
Outputs results in a structured format (SARIF referenced by topics)
β οΈ Limitations
No explicit tool count or MCP tool capabilities are provided in the available data
geolint fetches the page, its robots.txt and llms.txt, evaluates 51 known AI
crawler tokens against your robots.txt, runs 52 audit rules, and prints a
scored report with a concrete fix for every finding.
Why
AI answers are the new front page. ChatGPT, Perplexity, Claude, Copilot and
Google AI Overviews send traffic β or don't β based on whether their crawlers
can fetch and quote your pages.
Most sites accidentally block or confuse AI crawlers. A stale
Disallow: /, a noindex left over from staging, a client-rendered page that
looks empty to a bot that doesn't run JavaScript.
Existing tools are blocklists or score-only web apps. They tell you to
block everything, or give you a number with no path to improve it. geolint is
the linter: concrete findings, concrete fixes, runnable in CI on every PR.
What it checks
52 rules across 5 categories β geolint rules lists them all, and
docs/rules.md documents what each rule checks, why it matters
and how to fix violations.
The action produces score/grade step outputs, a SARIF report for GitHub code
scanning, and a markdown report for job summaries and PR comments. Full recipes
β SARIF upload, updating a single PR comment, baseline drift detection β in
docs/github-action.md.
Exit code is 1 when the score drops below the gate (or findings regress
against --baseline), 0 otherwise β works in GitLab CI, CircleCI, npm
scripts, pre-deploy hooks.
Show your score as a README badge
bash
npx @iliasabk/geolint check https://example.com --badge
# β writes geolint-badge.svg + prints the markdown snippet to paste
Commit the SVG, or regenerate a shields endpoint JSON in CI
(--badge-endpoint) for a badge that never goes stale.
Output formats
-f pretty (default) renders the terminal report above. The machine formats:
-f json β the full ScanReport: findings, per-category scores, bot access matrix
-f html β a self-contained interactive report (score ring, findings filter,
bot matrix) you can share or host anywhere
Add -o report.json to write to a file; stdout stays clean for piping.
geolint on the real web
The repo dogfoods itself: a nightly workflow re-audits eight
well-known sites and commits the scores back, and the showcase
site publishes the full interactive
reports β github.com, anthropic.com, stripe.com and more, regenerated on every
push to main.
scan(url, options) returns a typed ScanReport. Also exported: the bot
registry (AI_BOTS, botsByPurpose), the rule registry (allRules,
ruleById), robots.txt/llms.txt parsers, badge generators, scorers and all
four reporters.
Use it from AI assistants (MCP)
geolint mcp speaks the Model Context Protocol
over stdio β Claude Desktop, Cursor, VS Code and Windsurf can audit sites,
generate llms.txt and compare URLs as native tools:
Five tools: audit_url, generate_llms_txt, compare_urls, list_rules,
list_ai_bots β all read-only, with structured output and per-call timeouts.
Setup for every client: docs/mcp.md.
The bot registry is the point
geolint bots lists 51 AI crawler tokens with a purpose-aware impact
assessment β because "should I block this bot?" has a different answer for each:
Purpose
Examples
If you block it
training
GPTBot, ClaudeBot, CCBot
absent from future training data
search
OAI-SearchBot, PerplexityBot, Claude-SearchBot
invisible in AI answers now
user-fetch
ChatGPT-User, Claude-User
invisible in AI answers now
mixed
Bytespider, Amazonbot, Diffbot
both
And two nuances other tools miss:
Some fetchers ignore robots.txt. OpenAI, Perplexity and Meta document that
their user-triggered fetchers (ChatGPT-User, Perplexity-User,
Meta-ExternalFetcher) may not honor robots.txt. ai-crawler/user-fetch-bypass
tells you when a Disallow won't work β enforce at the WAF/auth layer instead.
Stale tokens.anthropic-ai, Claude-Web, FacebookBot are retired.
ai-crawler/stale-tokens flags them and names the replacement token β a
User-agent: anthropic-ai rule does nothing today.
Control-only tokens like Google-Extended and Applebot-Extended never fetch
at all β they only set a preference β and geolint treats them accordingly.
What geolint is honest about
llms.txt is a proposal, not a standard. No major AI vendor has committed
to reading it β so llms-txt/* findings are weighted as warnings and hints,
not errors. geolint still checks it (and geolint init generates it) because
adoption is growing and the cost is one file.
Correlation β causation. The citability rules are grounded in published
GEO research (quotations/statistics/citations measurably lift share-of-answer;
AI crawlers other than Googlebot and Applebot don't execute JavaScript), but
signals like question-shaped headings are hints, not facts β they're info
severity and geolint says so.
Every rule shows its reasoning.docs/rules.md documents
why each rule exists; the research sources are in
docs/research-notes.md, including the vendor docs
behind every bot's robots.txt posture.
The bot registry is a standalone reference.docs/ai-crawlers.md lists every tracked token with
purpose, per-vendor robots.txt posture and vendor docs β the same data
geolint bots and the list_ai_bots MCP tool expose.
Compared to the alternatives
Purpose-aware bot registry
Per-vendor robots.txt posture
Runs in CI
Fix per finding
Generates llms.txt
Free / OSS
geolint
β
β
β
β
β
β
ai.robots.txt-style blocklists
β
β
n/a
β
β
β
GEO-optimizer skills / prompt packs
β
β
β
β
β
varies
llms.txt validators
β
β
some
partial
some
β
Hosted GEO audit web apps
partial
β
β
partial
β
β
Details and the reasoning behind each column: docs/comparison.md.
geolint also ships an MCP server, a score badge and regression baselines.
Roadmap
Planned for v0.4+:
geolint watch β re-audit on deploys/file changes
Custom rule API for project-specific checks
Deeper schema coverage (more @type validators)
Homebrew formula
Report localization beyond English
Contributing
Issues and PRs welcome β see CONTRIBUTING.md. New rules are
the best contribution: each needs a check(ctx), findings with fix, a test
and a docs entry.