@mindstone/mcp-server-browser-automation

Browser control you can watch: open pages, sign in, click around, fill forms, take screenshots, and keep a reusable browser session.
Best for practical web tasks where the user needs to see, approve, or reuse browser state instead of running a full browser-testing stack.
Status
- Version: 0.2.2 ยท npm
- Auth: None (
server.json)
- Tools: 21 (navigation, observation, interaction, sessions, files)
- Surface: browser-automation
- Machine-readable:
STATUS.json
Why this exists
Microsoft's Playwright MCP is a strong choice for broad browser automation and testing. This connector is deliberately smaller and more visible.
Use it when an assistant needs to work through ordinary websites in a way a person can follow: open a real browser, let the user complete a login, click through admin screens, fill forms, take screenshots, and come back to the same session later. The point is trust and day-to-day usefulness, not exposing every browser-testing capability.
Example interaction
"Open https://example.com, tell me the page title, and take a screenshot."
Tools the host calls:
browser_navigate โ opens the URL in the configured browser session.
browser_get_page_info โ returns the current URL and page title.
browser_screenshot โ captures a PNG screenshot.
Response (trimmed):
{
"ok": true,
"url": "https://example.com/",
"title": "Example Domain"
}
Requirements
- Node.js 20+
- npm
- The
agent-browser CLI on PATH, or npx available so the server can install it automatically.
One-click install

After clicking the button, your host will prompt you to fill: AGENT_BROWSER_SESSION_NAME, AGENT_BROWSER_SHOW_WINDOW.
Manual config for Claude Desktop / Claude Code / Goose / Continue.dev (Browser Automation)
{
"mcpServers": {
"Browser Automation": {
"command": "npx",
"args": [
"-y",
"@mindstone/mcp-server-browser-automation"
],
"env": {
"AGENT_BROWSER_SESSION_NAME": "mcp",
"AGENT_BROWSER_SHOW_WINDOW": "true"
}
}
}
}
Quick Start
npx
npx -y @mindstone/mcp-server-browser-automation
Or install globally:
npm install -g @mindstone/mcp-server-browser-automation
mcp-server-browser-automation
This server requires the agent-browser CLI binary to control the browser.
Binary Resolution
- PATH lookup (preferred): If
agent-browser is on your PATH, it is used directly.
- npx fallback: If the binary is not found, the server automatically falls back to
npx -y agent-browser@0.33.2.
Installing agent-browser
npm install -g agent-browser
Or let the npx fallback handle it automatically (slower on first use due to download).
Configuration
No API keys or credentials are required. The server communicates with the browser via the agent-browser CLI.
| Variable | Required | Description |
|---|
AGENT_BROWSER_SESSION_NAME | No | Session name for browser persistence (default: mcp) |
AGENT_BROWSER_SHOW_WINDOW | No | Set to false to run without a visible browser window. Default is visible (true). |
MCP_WORKSPACE_PATH | No | Workspace directory that browser_pdf writes into and browser_upload reads from. Defaults to the system temp directory. See Security notes. |
MCP Host Configuration
{
"mcpServers": {
"browser-automation": {
"command": "npx",
"args": ["-y", "@mindstone/mcp-server-browser-automation"]
}
}
}
Navigation
- browser_navigate โ Navigate to a URL
- browser_back โ Navigate back in browser history
- browser_forward โ Navigate forward in browser history
- browser_wait โ Wait for an element to appear or a specified time
Observation
- browser_snapshot โ Get the page accessibility tree with interactive element references
- browser_screenshot โ Take a screenshot of the current page
- browser_get_page_info โ Get the current page URL and title
- browser_get_text โ Get the text content of the page or a single element
- browser_pdf โ Save the current page as a PDF inside the workspace directory (refuses to overwrite existing files unless
overwrite: true)
Interaction
- browser_click โ Click an element using @ref or CSS selector
- browser_fill โ Clear a field and fill it with text
- browser_type โ Type text character by character (real keystrokes)
- browser_press_key โ Press a keyboard key
- browser_scroll โ Scroll the page in a direction
- browser_select โ Select an option from a dropdown
- browser_hover โ Hover over an element
- browser_upload โ Upload workspace files to a file input (regular files only, staged privately before upload)
- browser_evaluate โ Execute JavaScript in the page context (on by default;
destructiveHint: true so hosts can require confirmation โ see Security notes)
Session Management
- browser_tabs โ List open tabs or switch to a tab
- browser_close โ Close the browser session
- browser_authenticate โ Open a visible browser for manual login
Workflow
The typical workflow uses accessibility snapshots for reliable element targeting:
browser_navigate โ open a page
browser_snapshot โ see interactive elements with @ref IDs
browser_click / browser_fill โ interact using @ref references
browser_screenshot โ visual verification
Security notes
Browser automation has a large attack surface: the agent-browser CLI controls a real headless browser that loads URLs you pass it, runs page-side JavaScript, and persists cookies and session state across runs. Read this section before deploying.
browser_evaluate runs by default โ host confirmation is the gate
browser_evaluate lets the model execute arbitrary JavaScript inside the page context โ the security equivalent of giving the model a shell on whatever site it has just navigated to. The tool is registered unconditionally (capability-first); the safeguard is the host's tool-approval layer. The tool is marked destructiveHint: true, so MCP hosts SHOULD require explicit user confirmation before each invocation. Do not configure a host to auto-approve this tool.
URL scheme deny-list
browser_navigate and browser_authenticate accept only http: and https: URLs (plus the special about:blank). Other URL schemes are refused before the underlying agent-browser CLI is invoked:
file: โ would let pages read local filesystem paths
chrome: and chrome-extension: โ internal browser pages and installed extensions
javascript: โ equivalent to eval() against the current document
data: โ inlined attacker-controlled HTML/JS payloads
view-source: โ defeats the same-origin policy on rendered content
about: โ privileged internal pages (about:config, about:cache, about:debugging, โฆ); only about:blank is permitted
Cookie and session persistence
The connector tells agent-browser to use a named, persistent session via AGENT_BROWSER_SESSION_NAME (default value: mcp). All cookies, localStorage data, and any logins performed via browser_authenticate are stored on disk under that session name and reused across runs. Anyone who can read the session storage โ the local user, other tools running as the same user, or backups โ can also use those logged-in sessions.
To override the session name (for example, to keep separate profiles per project) set AGENT_BROWSER_SESSION_NAME explicitly in the host's MCP server config. To wipe state, close the browser via browser_close and remove the session directory managed by agent-browser.
Recommended deployment posture
- Run the connector against a separate browser profile โ a dedicated
AGENT_BROWSER_SESSION_NAME per MCP host. Do not reuse your daily browser profile: the connector reads and overwrites cookies in whichever profile it is pointed at, and a malicious page can ride the existing session of any site you are logged into.
- Require host confirmation for
browser_evaluate โ it runs by default; every call executes arbitrary JavaScript in the page context.
- Require host confirmation for
browser_authenticate and any flow that may navigate to authenticated sites โ otherwise prompt injection in fetched content can drive the browser at sites the user is logged into.
- Returned page content is enveloped as untrusted โ accessibility snapshots, page text, titles, URLs, tab lists, and
browser_evaluate outputs come from arbitrary websites and may contain prompt-injection attempts. The connector wraps them in <untrusted-content source="โฆ"> envelopes (with close-tag breakout escaping) so hosts and models treat them as data, not instructions. Keep that treatment on your side: don't strip the envelopes, and don't let page text alone trigger irreversible actions.
browser_pdf writes PDF files and browser_upload reads files, and both are constrained to a workspace directory: MCP_WORKSPACE_PATH when set, otherwise the system temp directory. Paths are canonicalised before use, so .. traversal, absolute paths outside the workspace, and in-workspace symlinks pointing outside it are all refused before the agent-browser CLI runs. Set MCP_WORKSPACE_PATH explicitly to control exactly where page captures land and which files the model can attach to a page. A MCP_WORKSPACE_PATH that cannot be resolved fails closed: file tools refuse to run rather than fall back to weaker checks.
Both tools are marked destructiveHint: true, so hosts can require user confirmation: browser_upload can trigger an immediate remote upload on pages that submit when the file input changes, and browser_pdf writes a local file.
Additional hardening:
- Staged file hand-off. Validated files are never re-opened by pathname.
browser_upload sources are opened once (O_NOFOLLOW + O_NONBLOCK, so a post-validation leaf-symlink swap fails and a planted FIFO cannot wedge the open), verified to be regular files (directories, devices, and FIFOs are refused), bound to a fresh confined resolution by device+inode (an intermediate-directory swap after validation redirects the resolution and is refused with UPLOAD_SOURCE_CHANGED), and copied into a fresh private staging directory that the CLI consumes. browser_pdf has the CLI write into a fresh private staging directory, then installs the PDF at the requested path: the destination directory's canonical identity is pinned before the CLI runs and re-verified before installing, so an intermediate directory swapped to a symlink mid-call is refused, and exclusive-create semantics refuse a file or symlink planted at the destination leaf instead of writing through it.
- No silent overwrite.
browser_pdf refuses an existing file_path with a FILE_EXISTS error unless the caller explicitly passes overwrite: true. Even then, a destination that is a directory is refused (DESTINATION_IS_DIRECTORY) โ the overwrite delete is a bare unlink and never recurses.
- Enveloped errors. Error output from the
agent-browser CLI can contain page-authored text; it is wrapped in <untrusted-content> envelopes (with close-tag breakout escaping) before reaching the model, and timeout errors do not echo the command's argument values.
Licence
FSL-1.1-MIT โ Functional Source License, Version 1.1, with MIT future licence. The software converts to MIT licence on 2030-04-08.