Scout
Browser automation, one binary. The simpler alternative to Playwright β no Node, no Python, no runtime. Drive a real browser from Go, any shell, any AI agent (built-in MCP server), or a chat UI.
A single statically-linked scout binary gives you a CLI, an 87-tool MCP server (so any MCP-aware agent β Claude Desktop, Cursor, Cline, custom β has a browser), a conversational chat UI, and a Go library with Gin-like middleware composition. Same engine, four access points.
brew install --cask klarlabs-studio/tap/scout-mcp
vs. Playwright
| Scout | Playwright |
|---|
| Install | one ~15 MB binary | npm + ~600 MB browser cache |
| Runtime dep | none (static) | Node.js always; Python/Java/.NET as wrappers |
| Drive from | Go, any shell, MCP, chat UI | TS/JS first-class; others lag |
| AI-agent native | built-in scout mcp serve | separate playwright-mcp project |
| Token-aware extraction | DOM diff, distillation, observation budgets (50β80% fewer tokens) | not provided |
| Action playbooks | record & replay deterministic JSON | codegen produces a script you maintain |
| Container deploy | drop into scratch or distroless | carry Node + browser binaries |
| CDP access | direct WebSocket, zero abstraction | internal protocol over CDP |
Quick Start
scout observe https://example.com
scout markdown https://news.ycombinator.com
scout screenshot https://github.com
scout extract https://example.com h1
scout frameworks https://react.dev
claude mcp add scout -- scout mcp serve
scout ui serve --provider=ollama --model=mistral
cd ui && npm install && npm run dev
Install
brew install --cask klarlabs-studio/tap/scout-mcp
curl -fsSL https://raw.githubusercontent.com/klarlabs-studio/scout/main/install.sh | bash
go install go.klarlabs.de/scout/cmd/scout@latest
go get go.klarlabs.de/scout
Run scout mcp serve and any MCP-aware agent has a browser. No second project to install, no Node runtime, no Python interpreter β the binary is the server. Configure in any MCP client:
claude mcp add scout -- scout mcp serve
{"mcpServers": {"scout": {"command": "scout", "args": ["mcp", "serve"]}}}
| Category | Tools |
|---|
| Navigation | navigate, observe, observe_diff, observe_with_budget |
| Interaction | click, click_label, click_text, type, hover, double_click, right_click, select_option, scroll_to, scroll_by, focus, drag_drop, dispatch_event |
| Forms | fill_form, fill_form_semantic (checkbox/radio + state echo), discover_form |
| Extraction | extract, extract_all, extract_table, auto_extract, scroll_and_collect, markdown, readable_text, accessibility_tree |
| Capture | screenshot, annotated_screenshot, pdf |
| Network | enable_network_capture, network_requests |
| Tabs | open_tab, switch_tab, close_tab, list_tabs |
| Frameworks | wait_spa, detect_frameworks, component_state, app_state |
| Playback | start_recording, stop_recording, save_playbook, replay_playbook |
| Video | start_screen_recording, stop_screen_recording |
| Smart Helpers | check_readiness, suggest_selectors, session_history |
| Vision | hybrid_observe, find_by_coordinates |
| Batch | execute_batch |
| Iframe | switch_to_frame, switch_to_main_frame |
| Trace | start_trace, stop_trace |
| Cookies | cookies_list, cookies_clear, cookies_set, dismiss_cookies |
| Diagnostics | detect_dialog, detect_auth_wall, console_errors (incl. network 4xx/5xx), failed_requests, compare_tabs, upload_file |
| Utility | has_element, wait_for, configure, set_viewport, web_vitals, select_by_prompt |
All tools have MCP annotations (ReadOnly, OpenWorld, ClosedWorld, Idempotent) for smart auto-approval. Read-only tools like observe, extract, and screenshot run without permission prompts.
Runtime Configuration
Switch between headless and visible browser without restarting, and opt into local-dev workflows (loopback, private IPs):
Agent: configure(headless: false) β browser window appears
Agent: navigate("https://...") β watch it work
Agent: configure(headless: true) β back to headless
Agent: configure(allow_private_ips: true) β unlock localhost / 192.168.* / 10.*
Agent: navigate("http://127.0.0.1:4173/") β drive your local dev server
The MCP server also reads SCOUT_ALLOW_PRIVATE_IPS=1 at startup as a one-shot toggle for trusted environments.
Screen Recording (video)
Record the active page as a video. Pure CDP β works in headless, no Playwright needed. Recording survives navigate, open_tab, and switch_tab calls in between, so a multi-page demo lands as one continuous clip:
Agent: start_screen_recording({ width: 1280, height: 800, fps: 15, format: "webm" })
Agent: navigate("https://example.com")
Agent: click("#cta")
Agent: navigate("https://example.com/dashboard") # recording continues across pages
Agent: stop_screen_recording()
β { path: "/tmp/scout-rec-XXX.webm", format: "webm", encoder: "ffmpeg",
frame_count: N, duration_ms: N }
If ffmpeg is on PATH, the result is encoded to WebM (libvpx-vp9) or MP4 (libx264). If not, scout returns the raw JPEG frames directory plus an ffmpeg concat list so you can encode offline. The result is always a file path, never base64 β never enters your LLM token budget.
Realistic FPS: ~10β15 on typical pages, capped at 30. Implementation polls Page.captureScreenshot (CDP Page.startScreencast events are silently dropped under --headless=new Chrome).
Browser UI
A conversational browser automation interface. Type natural language, watch the browser respond in real-time.
scout ui serve --provider=ollama --model=mistral
scout ui serve --provider=claude
scout ui serve --provider=openai --model=gpt-4o
scout ui serve --provider=groq --base-url=https://api.groq.com/openai --model=llama-3.3-70b-versatile
cd ui && npm install && npm run dev
The UI streams AG-UI protocol events over SSE:
- Chat panel with markdown rendering and quick-action pills
- Live browser viewport with screenshot streaming and URL bar
- Activity timeline showing tool calls in real-time
- Stop button to cancel mid-stream
The Go server handles the agentic loop: LLM decides which scout tools to call, executes them, streams browser state deltas back to the frontend. Supports any OpenAI-compatible endpoint via --base-url.
Agent Package (Go)
High-level Go API for callers that want to embed scout in a program. Structured output, auto-wait, goroutine-safe. Most users reach scout through the CLI or MCP server above β this section is for the Go-library path.
session, _ := agent.NewSession(agent.SessionConfig{Headless: true})
defer session.Close()
session.Navigate("https://example.com")
obs, _ := session.Observe()
session.Click("#submit")
_, diff, _ := session.ObserveDiff()
session.FillFormSemantic(map[string]string{
"Email": "user-example", "Password": "secret",
})
result, _ := session.AnnotatedScreenshot()
session.ClickLabel(7)
session.OpenTab("pricing", "https://example.com/pricing")
session.SwitchTab("default")
frameworks, _ := session.DetectedFrameworks()
state, _ := session.ComponentState("#app")
session.EnableNetworkCapture("/api/")
captured := session.CapturedRequests("/api/users")
session.StartRecordingPlaybook("login-flow")
pb, _ := session.StopRecordingPlaybook()
agent.SavePlaybook(pb, "login.json")
session.SaveProfile("session.json")
session.LoadProfile("session.json")
session.Markdown()
session.ReadableText()
session.AccessibilityTree()
session.ObserveWithBudget(500)
Core Library (Go)
Gin-like Engine/Context/Group/HandlerFunc with middleware composition. The lowest-level Go API β use it when you want full control of task lifecycle, named groups, and middleware chains:
engine := browse.Default(browse.WithHeadless(true))
engine.MustLaunch()
defer engine.Close()
engine.Use(middleware.Stealth())
engine.Use(middleware.Retry(middleware.RetryConfig{MaxAttempts: 3}))
engine.Use(middleware.Timeout(30 * time.Second))
admin := engine.Group("admin", middleware.BasicAuth("admin", "secret"))
admin.Task("export", func(c *browse.Context) {
c.MustNavigate("https://app.example.com/admin")
table, _ := c.ExtractTable("#users")
c.Set("data", table)
})
engine.RunGroup("admin")
Middleware
| Category | Middleware |
|---|
| Resilience | Retry, Timeout, CircuitBreaker, RateLimit, Bulkhead |
| Auth | BearerAuth, BasicAuth, CookieAuth, HeaderAuth |
| Anti-detection | Stealth (10 patches: webdriver, plugins, WebGL, etc.) |
| Network | BlockResources, WaitNetworkIdle |
| Utilities | ScreenshotOnError, SlowMotion, Viewport |
CLI
The CLI is a deliberate one-shot tier: each command launches a browser,
navigates, runs one read/diagnostic action, prints the result, and exits.
Stateful interactive flows (clickβtypeβsubmit, multi-tab, live cookie/network
manipulation) live in the MCP server (scout mcp serve) and chat UI
(scout ui serve), which keep a session alive across actions.
CLI defaults to visible browser (--headless to hide):
scout navigate <url>
scout observe <url>
scout markdown <url>
scout readable <url>
scout accessibility <url>
scout screenshot <url> [--output f]
scout pdf <url> [--output f]
scout extract <url> <selector>
scout table <url> <selector>
scout auto-extract <url>
scout eval <url> <expression>
scout form discover <url>
scout frameworks <url>
scout app-state <url>
scout aria <url>
scout vitals <url>
scout console <url>
scout dialog <url>
scout auth-wall <url>
scout cookies <url>
scout watch <url> [--interval=5s]
scout pipe <command> [selector]
scout record <url> [--output f]
scout mcp serve
scout ui serve [flags]
scout version
Architecture
scout/
βββ browse.go, engine.go, context.go # Gin-like API
βββ page.go, selection.go # CDP page & element interaction
βββ recorder.go # Action playbook recording (Navigate/Click/Type β JSON)
βββ middleware/ # stealth, resilience, auth, network
βββ agent/ # AI agent API (50+ methods)
β βββ session.go # Session lifecycle, Navigate, Click, Type
β βββ observe.go, diff.go # Observe, ObserveDiff, cost estimation
β βββ content.go # Markdown, ReadableText, AccessibilityTree
β βββ form.go # DiscoverForm, FillFormSemantic, MatchFormField
β βββ annotate.go # AnnotatedScreenshot, ClickLabel
β βββ network.go # EnableNetworkCapture, CapturedRequests
β βββ spa.go # DetectedFrameworks, ComponentState, GetAppState
β βββ tabs.go # OpenTab, SwitchTab, CloseTab, ListTabs
β βββ playbook.go # StartRecording, ReplayPlaybook, SavePlaybook
β βββ interact.go # Hover, DragDrop, SelectOption, ScrollTo
β βββ profile.go # CaptureProfile, ApplyProfile, SaveProfile
β βββ selector.go # Playwright :text() selector translation
β βββ budget.go # ObserveWithBudget, EstimateTokens
β βββ nlselect.go # SelectByPrompt, fuzzy NL element matching
β βββ batch.go # ExecuteBatch, sequential multi-action
β βββ vision.go # HybridObserve, FindByCoordinates
β βββ trace.go # StartTrace, StopTrace, action tracing
β βββ screencast.go # StartScreenRecording / StopScreenRecording β video via captureScreenshot polling + ffmpeg encode
β βββ iframe.go # SwitchToFrame, SwitchToMainFrame
β βββ vitals.go # WebVitals (LCP/CLS/INP)
βββ internal/cdp/ # WebSocket CDP client (context-aware)
βββ internal/launcher/ # Chrome process management
βββ cmd/scout/ # CLI + MCP server (87 tools)
βββ docs/ # Landing page (GitHub Pages)
Security
Vulnerability scanning runs on every push and PR via nox. Findings are uploaded to GitHub code scanning, annotated inline on PRs, and gated against .nox/baseline.json so regressions block merges. The status badge in the header reflects the latest main-branch scan.
nox also drives dependency remediation in place of Dependabot β the Nox Remediate workflow runs weekly (Monday 06:00 UTC) and on demand, executing nox fix against fresh OSV.dev findings and opening a single PR with the verified upgrades.
nox scan -severity-threshold high .
nox fix -input findings.json
License
MIT