Drive a real Playwright browser: annotated screenshot + structured page state after every action.
io.github.AlexKay28/clickcast MCP Server
The io.github.AlexKay28/clickcast MCP server drives a real Playwright browser. After every action, it can provide an annotated screenshot and structured page state, supporting workflows that require end-to-end or visual inspection of page behavior.
π οΈ Key Features
Real Playwright browser control
Annotated screenshot output after each action
Structured page state output after each action
π Use Cases
Browser automation tasks in Chromium-based environments
E2E testing and visual testing workflows
Page state verification alongside captured screenshots
β‘ Developer Benefits
Developer tooling for browser automation and inspection
Supports MCP and Model Context Protocol integrations
Works with Python-based development workflows
β οΈ Limitations
Described capabilities focus on screenshot annotation and structured page state per action; additional tool coverage is not provided in the available data.
ποΈ clickcast β Playwright browser automation, screenshots, and GIF/MP4/WebP recording for AI agents
Give AI agents visual + structured feedback about live web UIs β and give humans deterministic demo reels while you're at it.
What's new (v0.3.0) β live agent control via a new MCP server, pixel-level clickcast diff, accessibility semantics fused with the pixel-grid overlay, an official CI GitHub Action, day-one Homebrew/apt packaging, and a self-healing first run. See CHANGELOG.md for the full notes.
clickcast Β· scripted tour of tailwindcss.com β sticky AβB arrow, actions panel bottom-right, symmetric open/close pairs
See docs/ONE_PAGE_NAVIGATION_ORDER_TIPS.md for the nine principles behind why this reel reads as legibly as it does β and the scenario template you can copy for your own reels.
clickcast drives a real browser through a website and hands back two things:
A watchable reel β GIF / MP4 / WebP / raw frames.
A machine-readable JSON sidecar β every step's selector, timings, per-step frame paths, discovered elements, and post-action page state (title, URL, console errors, failed requests). Versioned. See docs/feedback-schema.md.
Point it at a URL and it will auto-discover the interactive elements and build a tour for you, or hand it a small YAML scenario for a scripted, repeatable walkthrough.
Broken down β pip install clickcast (requires Python β₯ 3.10) gets you the
CLI; clickcast install downloads Chromium (~one-time, ~180MB β kept out of
the pip package itself since it's versioned independently and every project
doesn't need every engine); --with-deps also pulls the system libraries
Chromium needs (Linux only, may prompt for sudo); clickcast doctor
confirms everything above actually worked.
Forgot the second step? Any command that needs a browser (auto, run,
shot, elements, mcp) detects a missing engine itself and offers to
install it right then instead of failing β say yes once and it retries
automatically. That prompt only fires in an interactive terminal; CI/scripted
runs fail fast with the exact fix command instead of hanging on stdin.
Homebrew (macOS/Linux)
bash
brew install --build-from-source ./Formula/clickcast.rb # works today, clone this repo first
brew install AlexKay28/clickcast/clickcast # tap not bootstrapped yet -- see docs/packaging/homebrew.md
apt (Debian/Ubuntu, .deb)
bash
bash scripts/build_deb.sh 0.2.9 && sudo dpkg -i dist-deb/clickcast_0.2.9_amd64.deb # works today, clone this repo firstsudo apt install clickcast # hosted repo not bootstrapped yet -- see docs/packaging/apt.md
Both native packages skip the Chromium download (~180MB, versioned
independently of clickcast) and never install a second ffmpeg (clickcast
already bundles one via imageio[ffmpeg]) -- run clickcast install --with-deps chromium once after either install path. Full design rationale,
what's live today vs. what needs one-time bootstrapping, and the exact
bootstrap steps: docs/packaging/homebrew.md,
docs/packaging/apt.md.
npm (clickcast + clickcast-mcp)
Two npm packages -- thin Node wrappers whose postinstall provisions an
isolated Python venv and pip-installs the pinned PyPI clickcast, since
there's no way to ship clickcast's Playwright/Pillow/ffmpeg runtime as pure
JS. Built specifically because the MCP ecosystem's install pattern is
npx <package>, not pip install:
bash
npx -y clickcast-mcp # MCP entry point -- execs `clickcast mcp`
npx clickcast --version # general CLI wrapper
Both are published on the npm registry. To run against your working copy
instead of the published version:
Full design rationale
(including why the shared provisioning code is a vendored copy rather than
a file: dependency), what's live today vs. what needs bootstrapping, and
the exact bootstrap steps: docs/packaging/npm.md.
For AI agents β copy this prompt into your chat
Paste the block below into your coding agent (Claude Code, Cursor, Copilot Chat, Codex, etc.) to teach it clickcast in one message. The agent will install the tool, verify the environment, generate a visual + machine-readable report for your project, and know how to gate CI on the results.
text
You are going to use `clickcast` (https://pypi.org/project/clickcast/) to visually verify
this project's frontend. It produces a watchable GIF reel + a JSON sidecar with
per-step selectors, timings, page state, error codes, and a graph of the app's
pages. Follow this exact sequence:
1. INSTALL
pip install clickcast
clickcast install --with-deps chromium # first time only, ~180 MB
clickcast doctor # bail out if this fails
2. LOAD FULL CAPABILITIES
Run `clickcast skill --json` and read the output. It lists every subcommand,
flag, docs URL, and sidecar schema pointer. Prefer this over guessing.
3. RUN A QUICK TOUR OF THE APP
If we have a running URL (dev server / staging / production):
clickcast auto <URL> --for-humans --emit-events --out tour.gif
If we have a specific flow to verify, write a YAML scenario and run:
clickcast run tour.yml --emit-events --out tour.gif
`--for-humans` gives a legible reel for the user to watch;
`--emit-events` prints a machine-readable JSONL line you can parse.
4. READ THE SIDECAR (JSON at `<gif>.json`, schema v3)
For each step, gate on the structured fields β NOT regex over prose:
- `status` : "ok" | "failed" | "skipped"
- `error_code` : "timeout" | "locator_missing" | "cross_origin" |
"navigation_error" | "selector_ambiguous" | "other"
- `skip_reason` : "optional_no_reaction" | "pre_action_failed" |
"element_vanished" | "cross_origin_bounce"
- `page_state` : title, url_after, console_errors, page_errors,
network_failed
The top-level `graph` block gives page nodes + navigation edges you can
use to reason about the app's shape, not just the sequence you ran.
5. WATCH STDERR FOR ADVISORIES
clickcast prints `β <message> β see <docs-url>` lines for known anti-
patterns (nav-heavy tour, click without DOM reaction, very short reel,
cross-origin bounce, incoherent cursor styling). Each has a stable
kebab-case id you can dedupe or gate on.
6. FOR CI REGRESSION GATES
`clickcast assertions <sidecar>.json --baseline golden.json` diffs the
run against a committed baseline; nonzero exit on drift. Byte-identical
across runs (timestamps, frame paths, and URL query strings excluded).
Reference docs (all in-repo, load lazily as needed):
- Sidecar shape: https://github.com/AlexKay28/clickcast/blob/main/docs/feedback-schema.md
- Human-legible reel authoring: https://github.com/AlexKay28/clickcast/blob/main/docs/ONE_PAGE_NAVIGATION_ORDER_TIPS.md
- Agent integration: https://github.com/AlexKay28/clickcast/blob/main/docs/ai-integration.md
If something is unclear, run `clickcast <subcommand> --help` before asking me.
Once the agent has this, ask it something concrete like "run clickcast auto against http://localhost:3000 and tell me which clicks had DOM reactions" β it now has everything it needs.
Live agent control (MCP)
Everything above is batch mode: record a whole tour, then read back a GIF + sidecar. clickcast mcp is the live counterpart β an MCP server that drives one action at a time (goto/click/type/scroll/...) and hands back clickcast's richer per-call payload (annotated frame, page_state, grid coordinates, an enumerated error_code) instead of a bare screenshot, so an agent can react before deciding the next step.
Reach for mcp when an agent needs to explore and react live; reach for auto/run for a repeatable, one-shot artifact (CI, docs, release notes). Full tool reference, client config for Claude Code / Claude Desktop, and the schema doc: docs/mcp.md Β· docs/mcp-tool-schema.md.
First run β 30 seconds
bash
clickcast auto https://example.com --out tour.gif
Produces two files:
tour.gif β the reel
tour.gif.json β the AI-consumable sidecar (schema_version: 1, spec at docs/feedback-schema.md)
Flags: --out, --format, --headful, --slowmo MS, --url URL (retarget the first goto step β see below), --var key=value (repeatable β substitute {{ key }} inside the scenario), --no-sidecar.
CLI flags override the scenario's meta: block.
Point an existing scenario at a different environment with --url β no YAML edits, no {{ URL }} templating:
bash
clickcast run tour.yml --url https://staging.example.com/app
--url rewrites the first goto step's URL and wins over --var URL=.... Only the first goto is touched β later goto steps are usually intra-app navigation from the entry point, so they stay put.
Dump the discovered interactive elements β useful for authoring selectors.
Each entry additionally carries an accessibility block (Playwright's own
role / accessible name / interactive state, fused with the pixel-grid
overlay's grid_cell when --grid is on) β see
docs/feedback-schema.md.
bash
clickcast elements https://example.com --json > elements.json
clickcast elements https://example.com --grid --grid-pitch 50 --json
Flag
Default
Notes
--limit N
20
Cap on returned elements.
--json
off
Emit machine-readable JSON on stdout.
--viewport WxH
1280x800
--device NAME
β
Playwright preset (e.g. "iPhone 15", "Pixel 8").
--engine E
chromium
chromium / firefox / webkit.
--headful
off
Show a real browser window.
--lang LOCALE
β
e.g. en-US.
--dark
off
Emulate prefers-color-scheme: dark.
--slowmo MS
0
Delay each Playwright op by N ms.
-v / --verbose
β
Repeatable.
--grid
off
Populate each element's accessibility.grid_cell (#196).
--grid-pitch N
100
Major-line spacing in px, used for grid_cell.
--grid-color HEX
#FFFFFF33
Unused by elements beyond validation β kept symmetric with auto/run/shot.
--grid-style full|ruler
full
Unused by elements beyond validation β kept symmetric with auto/run/shot.
clickcast doctor # human-readable
clickcast doctor --json # machine-readable, non-zero exit on failure
config
Read / write persistent defaults.
bash
clickcast config path # print the user config file path
clickcast config list # every effective value + source
clickcast config get engine
clickcast config set engine firefox
Set values land in the user TOML at clickcast config path. See Configuration for precedence.
install [enginesβ¦]
Wrapper over playwright install. Default engine: chromium.
bash
clickcast install # chromium only
clickcast install firefox webkit # add more
clickcast install --with-deps chromium # Linux: pull system libs (needs sudo)
mcp
Start a stdio MCP server for live agent control β see Live agent control (MCP) above and docs/mcp.md. Requires pip install 'clickcast[mcp]'.
bash
clickcast mcp --grid # every action tool's returned frame carries the coordinate overlay
Flags: --engine, --viewport, --device, --headful, --lang, --dark, --grid/--grid-pitch/--grid-color/--grid-style β all defaults for start_session when the connecting agent doesn't override them.
Scenario format
A scenario is plain YAML: a meta: block and a list of steps:. Full worked examples: docs/scenarios/.
yaml
meta:name:WorldSightbroadtourengine:chromium# chromium | firefox | webkitviewport:1280x800device:null# or "iPhone 15", "Pixel 8", "iPad Pro"fps:12dwell:1.0# default seconds after each stepformat:gif# gif | mp4 | webp | framesout:worldsight.gifsteps:-goto:https://worldsight-weld.vercel.appwait:networkidlelabel:OpenWorldSight-click:"text=3D"label:Switchto3Dglobedwell:2.0-hover:"[aria-label='Rankings']"-click:"[aria-label='Rankings']"label:OpenRankings-type:into:"#search"text:"Japan"label:SearchJapan-select:in:"#metric"value:"GDP"-scroll:to:footer
Supported actions
Action
Example
Notes
goto
goto: https://β¦
Navigate. Pair with wait.
click
click: "text=Compare"
CSS, text=β¦, or role=β¦ selectors β Playwright syntax. Also accepts click: { selector: ..., wait: networkidle } β wait (same shape as goto's) blocks after the click; pair it with a target that triggers client-side/SPA navigation.
dblclick
dblclick: ".cell"
Also accepts wait, same as click.
hover
hover: ".menu"
Reveals CSS :hover state.
type
type: { into: "#q", text: "Japan", delay: 40 }
delay is per-char ms.
press
press: "Enter"
Or press: { key: "Ctrl+A", selector: "#in" }.
select
select: { in: "#m", value: "GDP" }
in: in YAML β into internally.
scroll
scroll: { to: "footer" } or scroll: { by: 600 }
Element or pixel scroll.
wait
wait: 1.5 or wait: networkidle or wait: ".map-loaded"
Number = seconds, string = load-state or selector.
screenshot
screenshot: { full_page: true }
Force a frame capture.
Every step also accepts label, dwell, optional: true (don't fail the run if the selector is missing β sidecar records status: "skipped"), and repeat: N.
Variable substitution: {{ key }} inside any string, injected via --var key=value.
Python API
Fluent, chainable β every builder returns self:
python
from clickcast import Reel
reel_path = (
Reel("https://worldsight-weld.vercel.app", viewport=(1280, 800), fps=12)
.goto(wait="networkidle")
.click("text=3D", label="Switch to 3D globe", dwell=2.0)
.click("[aria-label='Rankings']", label="Open Rankings")
.scroll(to="footer")
.save("worldsight.gif") # or .save("tour.mp4", quality=8)
)
Async variant for callers already inside a running event loop:
from clickcast import discover
elements = discover("https://example.com", limit=10)
Skip the sidecar with save(..., no_sidecar=True).
Pre-push iteration on a local static build
Building a static site and reeling it locally before pushing? Use Reel.serve_dir (or the standalone serve_directory helper) as a context manager β it starts a threaded HTTP server on a free port, yields the base URL, and tears the server down on exit. No more python3 -m http.server 8091 & invocations that leak past your shell session and collide with the next iteration.
python
from clickcast import Reel
with Reel.serve_dir("./public") as url:
Reel(url).goto().click(".chip").save("out.gif")
# Server is gone; the port is free.
Defaults are safe for dev iteration: loopback-only bind (127.0.0.1), OS-picked free port, ThreadingHTTPServer so parallel browser requests don't queue. Override any of them explicitly:
python
from clickcast.serving import serve_directory # importable without going through Reelwith serve_directory("./dist", port=8091, bind="0.0.0.0", threading=False) as url:
...
bind="0.0.0.0" exposes the server to your LAN β opt-in, not the default.
Reading the sidecar
Every recording run writes <out>.json alongside the media file.
python
from clickcast.feedback import load
report = load("tour.gif.json")
for step in report.steps:
if step.status == "failed":
print(f"step {step.index} ({step.action}) failed: {step.error}")
print(" frames:", step.frames)
if step.page_state:
print(" console errors:", step.page_state.console_errors)
On GitHub Actions? .github/actions/clickcast
packages this recipe (plus the visual diff gate below, browser caching,
and a PR comment with the reel + a summary table) as a reusable Action β
see docs/ci/README.md. This section stays the
canonical recipe for every other CI platform, and for anyone who'd rather
not depend on a third-party Action for two lines of shell.
Every reel writes a JSON sidecar, but raw sidecars carry timestamps, frame
filenames, and query-string tokens β none of which are stable across runs.
For a proper CI regression gate, use clickcast assertions (or
Reel.assertions()) to distill the sidecar down to the shape that
actually matters: step count, per-step action / label / status, and the
per-step error counters.
The distilled shape is byte-identical across runs of the same scenario
against the same URL (schema: docs/assertions-schema/v1.json).
Diff it against a committed baseline; non-zero exit on drift.
Bootstrap the baseline once:
bash
clickcast run tour.yml --out reel.gif
clickcast assertions reel.gif.json > tests/golden-tour.json # commit this
Then in CI (2 lines):
bash
clickcast run tour.yml --out reel.gif
clickcast assertions reel.gif.json --baseline tests/golden-tour.json
Exit 0 means the target UI produced the same step ordering, statuses, and
error-signal counts as when the baseline was captured; anything else is
real drift and the command prints per-line descriptions like
step 2: status changed 'ok' -> 'failed'.
Excluded from the distilled shape on purpose: wall-clock timestamps,
per-step duration_ms, frames filenames, resolved URLs (including
query-string tokens), cursor_xy. If you need those in your gate too,
diff the raw sidecar with your own tooling β the assertion set is the
narrow "did the UI still behave" contract, not the wire-level snapshot.
Its visual companion: clickcast diff
assertions is structural β it never looks at a pixel. clickcast diff
(and Reel.visual_diff()) is the pixel-level companion: it pairs up two
sidecars' steps and pixel-diffs their frames, reporting a percent-changed
and a list of changed bounding regions per step, plus region-highlighted
diff images. Reach for assertions to catch "the flow broke" (wrong step
count, a step that started failing); reach for diff to catch "the flow
ran fine but the button moved / the color changed / the layout shifted."
The two compose β run both in CI for full coverage, or diff alone if
pixel drift is your only concern.
Exits non-zero when any step's changed pixel percentage exceeds
--fail-above, or when a step couldn't be paired with its baseline
counterpart at all (mismatched step counts fall back to label matching;
anything still unmatched is flagged rather than silently skipped).
--out collects the region-highlighted diff images plus a summary.json
for CI artifacts; --threshold tunes the per-pixel noise floor. Clickcast's
own overlays (progress bar, action label, actions panel, cursor + ripple)
are excluded from the diff by default β otherwise every run would flag its
own chrome as a regression β pass --no-exclude-overlays for a strict
raw-pixel diff. Same signal from Python via reel.visual_diff(run_path, baseline_path).
Configuration
Precedence (highest β lowest):
CLI flags
Scenario meta: block
CLICKCAST_* environment variables
Project ./clickcast.toml
User TOML (path via clickcast config path)
Built-in defaults
Every Config field can be set at any of these layers: engine, viewport, device, headful, slowmo, lang, dark, proxy, fps, dwell, format, quality, loop.
Project TOML β flat or [defaults]-wrapped both work:
Session launches a Playwright browser at the requested viewport/device.
Actions run one step at a time (click, type, scroll, β¦) with normalised timings and cursor tracking.
Recorder captures a pre-frame + N padding frames per step (deterministic filenames, byte-identical copies for padding).
PageStateCollector subscribes to page events for the sidecar.
Encoder produces the final artifact; ReportBuilder finalises the JSON.
The annotator (clickcast.annotate.Annotator β click ripples, cursor trail, caption bar, progress bar) ships as a library API in v0.1. Automatic wiring into auto / run outputs is planned for v0.2 (see Roadmap).
Troubleshooting
Blank frames β the site is a SPA; increase --initial-wait, or add wait: networkidle (or a specific selector) to the first step.
ffmpeg not found β imageio-ffmpeg bundles a static binary; falls back if missing. Choose gif / webp if you'd rather skip MP4 entirely.
Selector not found β clickcast elements <url> shows what's actually clickable. Or mark the step optional: true.
Can't reach an internal site β set CLICKCAST_PROXY, or proxy in the scenario meta: block.
Chromium missing β clickcast install. On Linux CI add --with-deps.
Sidecar shape changed β the current schema is versioned at src/clickcast/feedback/schema/v1.json; a future v2 (see #29) will add a graph block without breaking v1 consumers.
On-frame HUD β fixed header/footer with step index, action verb, target role, URL. OCR-legible so LLMs can read the reel as a strip of images.
BFS UI exploration β clickcast explore <url> treats the app as a state graph: discover β click β discover the new state β recurse. Bounded, deterministic, with visited-state dedup.
Sidecar schema v2 β adds a top-level graph block (nodes = distinct page states, edges = (from, to, action, transition_kind)).
Automatic annotation of auto / run outputs.
Related projects
vercel-labs/webreel β a TypeScript tool for authoring polished demo videos. Easy to confuse with this project given the similar name and overlapping output, so: clickcast is a Python tool aimed primarily at AI agents that need a visual modality onto a live web UI, and secondarily at humans who want reproducible demo reels. If you want hand-authored marketing videos, webreel is the better fit.