Pull every URL out of the current file in one keystroke
Any text file — Markdown, HTML, CSS, JavaScript, TypeScript, JSON, YAML, Properties, TOML, INI and XML know what to exclude
Useful? A star or rating is how other developers find it —
★ GitHub ·
★ Open VSX ·
★ Marketplace
What it does
Open a file, run URLs-LE: Extract URLs, and every URL in the document lands in a new editor — deduplicate and sort it from there. Works in VS Code and in VS Code–based editors like Cursor and VSCodium (installable from Open VSX).
- Link auditing — every link, autolink, and plain URL in Markdown and HTML (code blocks and comments excluded)
- Source review — URLs in string literals, template literals, and comments across JS/TS
- Config sweep — URLs in JSON strings, YAML values, Java properties, TOML/INI values, and XML (Maven POMs, feeds)
Install
| Where | What you get | Install |
|---|
| VS Code | The extraction, in your editor, on a keystroke | Marketplace |
| Cursor, VSCodium, Windsurf | The same extension | Open VSX |
| A terminal or a CI step | The same run over a whole tree, with exit codes | cargo install urls-le · crates.io |
| Any MCP agent, via Node | extract_urls over stdio | npx urls-le-mcp · npm |
Use it from an AI agent
The same extraction engine runs as an MCP server, so an agent can call it directly instead of you running a command.
| Editor | How |
|---|
| VS Code 1.101+ | Nothing to install — the extension registers extract_urls with agent mode |
| Claude Code | claude mcp add urls-le -- npx -y urls-le-mcp |
| Cursor, Windsurf, anything else | point it at npx urls-le-mcp |
extract_urls(content, format?, filename?, dedupe?, maxResults?)
Returns every URL with its protocol and 1-based line and column, capped at 500 by default with meta.truncated so a large file cannot flood the agent's context window.
The server takes content and returns data — it reads no files and makes no network requests of its own. Published as urls-le-mcp on npm and as io.github.nolindnaidoo/urls-le in the MCP registry.
Configuring it by hand — any host with an MCP config file
Most hosts read a JSON config. Add one entry:
{
"mcpServers": {
"urls-le": {
"command": "npx",
"args": ["-y", "urls-le-mcp"]
}
}
}
-y skips the install prompt on first run. Pin a version if you would rather not track releases — urls-le-mcp@2.4.0.
Prefer not to go through npx on every launch? Install it once and point at the binary instead:
npm install -g urls-le-mcp
{
"mcpServers": {
"urls-le": { "command": "urls-le-mcp" }
}
}
It speaks MCP over stdio and needs no environment variables, no API key and no configuration of its own. To check it before wiring it into anything:
echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | npx -y urls-le-mcp
That prints the tool list and exits — if you see extract_urls, the server works.
Across a folder or a workspace
Extract reads the document you have open. A scan reads many files from disk and gives one report.
- The whole workspace: run
URLs-LE: Extract URLs from Workspace from the command palette.
- One folder: right-click it in the Explorer and choose
Extract URLs from Folder, or run URLs-LE: Extract URLs from Folder and pick one.
A project writes the same URL in many places, so the report is the distinct URLs, the most widely used first, with how often each is written and where:
# URLs-LE workspace report
`my-project` · 3 file(s) read · 2 distinct URL(s), 5 occurrence(s) in 3 file(s)
| URL | Occurrences | Files |
|---|---|---|
| `https://example.com/docs` | 3 | 2 |
| `https://api.example.com/v1` | 2 | 2 |
## `https://example.com/docs` (3)
- `config.json` · **2:12**
- `docs/a.md` · **1:5**, **1:34**
## `https://api.example.com/v1` (2)
- `docs/a.md` · **2:6**
- `src/b.ts` · **1:14**
That is with urls-le.showPositions on. It is off by default, and then each line is the file and how many times the URL is in it: docs/a.md (2). The copy on the clipboard follows urls-le.clipboardIncludesPositions, as it does for Extract.
What a scan reads. Files come from disk, so an unsaved edit is not seen. A file over the safety size, or one that is not UTF-8 text, is left unread. It stops at 5,000 files or 10,000 listed occurrences. The report ends with a line for each thing it left out, so a short report is never mistaken for a clean project.
What it skips, and how to change that. Three switches are on by default, and each can be turned off on its own in Settings:
| Switch | Skips |
|---|
scanUseDefaultExcludes | Dependency folders, build output, tool caches and lockfiles. The full list is below |
scanRespectGitignore | Whatever the project's .gitignore files skip |
scanSkipBinaryFiles | Images, fonts, archives and other files that are not text |
Two lists adjust the result without turning a switch off. To skip more, add a pattern to scanExcludes. To read something a switch would skip, add it to scanAlwaysInclude:
{
"urls-le.workspace.scanExcludes": ["**/fixtures/**"],
"urls-le.workspace.scanAlwaysInclude": ["**/vendor/**"]
}
URLs-LE: Open Settings opens all of these in the Settings editor.
The built-in list
Folders, wherever they appear:
.git, .hg, .svn, node_modules, bower_components, jspm_packages, .pnpm-store, .yarn, vendor, site-packages, Pods, Carthage, dist, build, out, target, _build, _site, dist-newstyle, zig-out, storybook-static, cdk.out, DerivedData, CMakeFiles, .next, .nuxt, .output, .svelte-kit, .angular, .astro, .docusaurus, .vuepress, .expo, .turbo, .parcel-cache, .cache, .sass-cache, .jekyll-cache, .dart_tool, .pub-cache, .gradle, .kotlin, .cxx, .externalNativeBuild, captures, ephemeral, .symlinks, .swiftpm, .build, .bundle, .stack-work, .zig-cache, .godot, elm-stuff, .vercel, .netlify, .serverless, .aws-sam, .terraform, .venv, venv, __pycache__, .tox, .nox, .mypy_cache, .pytest_cache, .ruff_cache, .ipynb_checkpoints, .eggs, coverage, htmlcov, .nyc_output, .vscode-test, .idea, .vs, xcuserdata, *.egg-info
Files, wherever they appear:
*.min.js, *.min.css, *.map, *.snap, *.lock, package-lock.json, pnpm-lock.yaml, npm-shrinkwrap.json, go.sum, *.pbxproj, *.iml, local.properties, output-metadata.json, .flutter-plugins, .flutter-plugins-dependencies, .packages, Generated.xcconfig, flutter_export_environment.sh, GeneratedPluginRegistrant.*, fastlane/report.xml, fastlane/test_output/**, doc/api/**
Not on the list, because they are ordinary folders in many projects: bin, obj, tmp, logs, public, generated. A project that generates those ignores them in git, and the scan reads .gitignore.
The settings that shape a scan are under Settings.
The CLI
The same extraction runs from a terminal or a shell pipeline: a Rust CLI
in crate/, sharing one corpus with the extension —
crate/fixtures/ — so the two can never read a
document differently.
urls-le .
urls-le --dedupe docs/
urls-le mcp
urls-le . | jq -r '.urls[].value' | sort -u | lychee -
Exit codes follow grep — 0 URLs found, 1 none found, 2 the question
was malformed — so if urls-le src/; then … works and finding nothing
is an answer rather than an error.
It has no opinions, deliberately. No link checking, no
insecure-scheme flag, no filtering. An http:// URL is wrong in a
production config and right in a test fixture; a tool that decides for
you is one you configure, then argue with, then mute — and the muting
takes the extraction with it. Pipe it to lychee and let that have the
opinions.
| Format | Language IDs | Notes |
|---|
| Markdown | markdown | Fenced code blocks and inline code excluded |
| HTML | html | <!-- --> comments excluded (multi-line supported) |
| CSS | css | Quoted or bare url(...), @import |
| JavaScript / TypeScript | javascript, typescript | Strings, template literals, comments |
| JSON | json | String literals only, via jsonc-parser token offsets; comments are trivia |
| YAML | yaml | Whole content, comments included |
| Properties | properties | #/! comment lines excluded |
| TOML | toml | Parsed values only; comments excluded |
| INI | ini | Whole content; ;/# comment lines excluded |
| XML | xml | Attributes and text content |
| Anything else | any | Whole document scanned; fileType reports unknown |
No document is refused. The eleven above know what to exclude; every
other language id — python, go, shellscript, csv, plaintext,
log, whatever your editor calls it — is scanned whole, because a URL is
unambiguous in any text and there is nothing in those worth excluding.
Extracted protocols: http, https, ftp, file, mailto (requires an @), tel. Every occurrence is reported with its real line and column — TOML positions are forward-located in the source and can be approximate for repeated identical values. A URL ends at whitespace or a quote/bracket delimiter, so relative links (/docs) and bare domains (example.com) are never extracted, and URLs containing raw spaces extract as space-terminated partials. Trailing ./, are kept — they are legal URL characters.
Commands
| Command | Description |
|---|
URLs-LE: Extract URLs | Extract all URLs from the active document |
URLs-LE: Extract URLs from Workspace | The distinct URLs in every file in the workspace, and where each one is |
URLs-LE: Extract URLs from Folder | The same for one folder. Also on a folder in the Explorer |
URLs-LE: Deduplicate URLs | Remove duplicate lines from the results |
URLs-LE: Sort URLs | Sort results alphabetically, by domain, or by length |
URLs-LE: Open Settings | Open URLs-LE settings |
URLs-LE: Help | Built-in documentation |
No command is bound to a key by default. Give any of them one under Keyboard Shortcuts in the editor.
Settings
| Setting | Default | Description |
|---|
urls-le.openResultsSideBySide | true | Open results beside the current editor |
urls-le.postProcess.openInNewFile | true | Open results in a new file (when not side-by-side) |
urls-le.showPositions | false | Show the line and column of each URL |
urls-le.copyToClipboardEnabled | false | Also copy results to the clipboard |
urls-le.clipboardIncludesPositions | false | Include the line and column in that copy |
urls-le.dedupeEnabled | false | Deduplicate extraction results automatically |
urls-le.notificationsLevel | silent | all = every notification, important = warnings + errors, silent = errors only |
urls-le.workspace.scanPatterns | ["**/*"] | The files a folder or workspace scan reads |
urls-le.workspace.scanUseDefaultExcludes | true | Skip dependency folders, build output, caches and lockfiles |
urls-le.workspace.scanRespectGitignore | true | Skip what the project's .gitignore files skip |
urls-le.workspace.scanSkipBinaryFiles | true | Skip images, fonts, archives and other files that are not text |
urls-le.workspace.scanExcludes | [] | More files to skip, as glob patterns |
urls-le.workspace.scanAlwaysInclude | [] | Files to read even when one of the three above would skip them |
urls-le.workspace.scanMaxFiles | 5000 | The most files one scan reads |
urls-le.workspace.scanMaxResults | 10000 | The most occurrences one scan lists before it stops reading |
urls-le.safety.enabled | true | Guardrails for very large files |
urls-le.safety.fileSizeWarnBytes | 1000000 | Refuse extraction above this file size |
urls-le.safety.largeOutputLinesThreshold | 50000 | Warn above this line count |
urls-le.statusBar.enabled | true | Show the status bar item |
urls-le.telemetryEnabled | false | Local-only event log (see Privacy) |
Languages
Twelve languages besides English:
German · Spanish · French · Indonesian · Italian · Japanese · Korean ·
Portuguese (Brazil) · Russian · Ukrainian · Vietnamese · Chinese (Simplified)
Both halves are covered — the manifest (command titles, setting names and
descriptions) and everything shown while the extension runs (notifications,
the status bar, the sort quick-pick). The extension follows VS Code's display
language, so it matches whatever the editor is already set to; no setting of
its own.
Privacy & security
- No network access. The extension never fetches, validates, or pings the URLs it extracts — it only reads the text of the active document. The
telemetryEnabled setting writes events to a local Output Channel you can inspect (URLs-LE Telemetry); nothing leaves your machine.
- The MCP server holds the same line. It takes content as an argument and returns data: no filesystem access, no network calls, no telemetry. Your agent already has file-read tools, so duplicating them inside the server would add a path-traversal surface for no capability.
check:mcp-bundle fails the build if the server ever imports something that could reach either.
- Error notifications redact home directories and credential-shaped fragments.
- One rating prompt, at most twice. On the 3rd successful use the extension asks once whether you would rate it, and once more on the 20th if you chose Later or dismissed it. Don't Ask Again ends it. Setting
notificationsLevel to important or silent yourself turns it off. The counts are kept in VS Code's extension storage and nothing is sent anywhere; Rate opens the listing you installed from — the VS Code Marketplace or Open VSX — in your browser.
Documentation
| Input | Size | Found | Time | Rate | Scan speed |
|---|
| Markdown docs | 2.63 MB | 50,000 | 61.14 ms | 817,802/sec | 43 MB/s |
| HTML page | 1.25 MB | 30,000 | 21.88 ms | 1,370,972/sec | 57 MB/s |
| JSON config | 1.52 MB | 40,000 | 39.01 ms | 1,025,358/sec | 38.8 MB/s |
Median of 7 runs after warmup, on Apple M5 Pro, 24 GB RAM, Node 24.3.0. Inputs are generated
by scripts/benchmark.ts rather than checked in, so the sizes above are
exactly what was measured. Reproduce with bun run benchmark.
These are machine-specific and are not asserted in CI — a benchmark that gates
a build only tells you how busy the runner was.
Testing
| Metric | Coverage |
|---|
| Statements | 95.05% |
| Branches | 86.59% |
| Functions | 96.20% |
| Lines | 95.79% |
431 test cases across 31 files, plus an integration suite that runs
in a real VS Code extension host and an end-to-end test that installs the
built .vsix into a clean profile.
Generated from a real run — coverage/coverage-summary.json and
coverage/test-results.json — by scripts/coverage-readme.js; CI fails if
this section drifts. Reproduce with bun run test:coverage, and the case
count is the one vitest prints.
More from the LE family
Sixteen single-purpose tools for the work in front of every model. Each ships
a Rust CLI and an MCP server. One page: letools.dev
Get it out
- String-LE — Extract every string in a codebase, with its position, so a person can read them
- Numbers-LE — Extract every hardcoded number in a codebase, so a person can check them
- Units-LE — Extract every quantity with its unit, normalized, and refuse the ambiguous ones by name
- Dates-LE — Extract every date and timestamp, and the exact instant each one resolves to
- IDs-LE — Extract every UUID, ULID, NanoID, ObjectId and Snowflake, and decode the time inside
- IPs-LE — Extract every IP address, CIDR block and MAC, normalized and classified by scope
- URLs-LE — Extract every URL in a codebase, with its protocol and exact position
- Paths-LE — Extract every file path in a codebase, and say whether it still points at anything
- Colors-LE — Extract every color in a codebase, and say which ones are not in your palette
Check it
- Regex-LE — Find every regex in a codebase, and report which can be driven into catastrophic backtracking
- Versions-LE — Find where one dependency is constrained differently across a repository's manifests
- i18n-LE — Identify the i18n library a project uses, then audit its catalogs by that library's rules
- Scrape-LE — Check whether a page is scrapeable before the scraper is written, and say when it cannot tell
Guard it
- Secrets-LE — Find hardcoded credentials in a codebase, and never print one into the report
- EnvSync-LE — Compare the dotenv files in a tree, and say which keys are missing from which
- Unicode-LE — Find the Unicode that hides meaning — bidi controls, invisibles, homoglyphs, mixed scripts
Each stands on its own: no shared crate, no published core. Where two of them
agree, it is because the same answer was right twice.
Contact — nolindnaidoo.com · GitHub · LinkedIn
Also by nolindnaidoo
Rust — pixelcoords and pixelactions are one loop: pixelcoords answers
where, pixelactions acts there. Their own tools, their own voice — not
part of the LE family.
License
MIT © nolindnaidoo