Local videos or YouTube/Bilibili URLs -> timestamped transcript, keyframes, contact sheets. Offline.
io.github.vsh5dvsch7-png/yueying — Model Context Protocol (MCP) Server
This MCP server converts local video files or YouTube/Bilibili URLs into timestamped transcripts, keyframes, and contact sheets. It operates offline, producing derived media artifacts from the provided video inputs. The server description indicates an emphasis on transforming video sources into searchable or reviewable outputs.
🛠️ Key Features
Accepts local video files and YouTube/Bilibili URLs
Generates timestamped transcripts
Extracts keyframes
Produces contact sheets
Runs offline
🚀 Use Cases
Create transcripts aligned to video timestamps
Extract keyframes for review or referencing
Generate contact sheets for quick visual indexing
Use without network dependencies (offline processing)
⚡ Developer Benefits
Predictable offline workflow for video-to-artifacts conversion
Output artifacts suitable for building video review or indexing pipelines
⚠️ Limitations
Described inputs are limited to local videos and YouTube/Bilibili URLs
Only the offline behavior and listed outputs are specified
Point Claude, Cursor or any MCP client at a video and get back a timestamped transcript plus keyframe contact sheets — offline, no API key. Local files first; URLs (YouTube, Bilibili, Douyin, Xiaohongshu, TikTok, Vimeo, …) are videos you are entitled to process, fetched via yt-dlp at ≤720p and deleted after processing by default.
Yueying (阅影) means "read video". One package gives you an MCP server, a CLI and an agent skill.
What you get
Claude Desktop: a YouTube link is pasted, the watch_video tool runs for about 40 seconds, and Claude answers with timestamped key points
Claude Desktop with yueying connected: paste a link, wait about forty seconds, get the video back as
timestamped notes. This video ships captions, so speech recognition never ran, and the model asked for
the transcript only. Keyframes and contact sheets come back through get_frames when it needs to see
the screen. Demo video: GitInGifs: Git Branches by GitLab, CC BY.
Contact sheet from a 24-second demo clip (four app screenshots with Chinese narration). The yellow label on every tile is the keyframe number and timestamp; the model cites them back to you.
The transcript of the same clip — local speech recognition, language auto-detected as Chinese:
(The app is called 月读; ASR heard the homophone 阅读. Speech recognition does that to names — the model corrects it from the on-screen text in the frames.)
Every video becomes one folder:
code
report.md index for the model: metadata, chapters, contact sheets, keyframes, transcript
transcript.txt paragraphs with [mm:ss] timestamps
transcript.srt subtitles for any player
grid_01.jpg … 3x3 contact sheets, 9 keyframes each, in time order
frames/ full-size keyframes, e.g. f003_00m15s.jpg
manifest.json machine-readable result (paths, segments, chapters, options)
Why yueying
Captions first, Whisper only when needed. Platform subtitles are used when they exist. Otherwise local faster-whisper: large-v3-turbo on an NVIDIA GPU, small on CPU, automatic CPU fallback — nothing is uploaded, no key.
ffmpeg bundled. Works on Windows 11 out of the box (imageio-ffmpeg); no PATH fiddling.
Token-efficient. Keyframes are taken at scene changes, near-duplicates dropped, then packed into 3x3 contact sheets with burned-in timestamps. One sheet ≈ 1–2K tokens for nine moments; one transcript with [mm:ss] paragraphs.
Chinese platforms and the rest. Bilibili (multi-part, collections, member videos with your browser login), Douyin, Xiaohongshu — and YouTube, TikTok, Vimeo, X and every other yt-dlp site.
Zero API keys, zero telemetry. The only network traffic is the video site you name and one Whisper model download. See the privacy policy.
Benchmark: a 6-minute Bilibili video → report in ~90 s on an RTX 5060 laptop; on CPU with model=small expect ~1–2 min per 10 min of speech.
winget install astral-sh.uv # Windows
brew install uv # macOS
curl -LsSf https://astral.sh/uv/install.sh | sh # Linux / macOS
Warm up and check everything once (installs the package, probes the GPU, downloads the speech model, runs a 2-second smoke test, prints config to paste):
bash
uvx yueying mcp --setup
Add the server to your client (below), then ask: "Watch C:\videos\lecture3.mp4 and turn the steps into notes" or "What does this video say about docker compose: https://www.bilibili.com/video/BV…".
Claude Desktop
%APPDATA%\Claude\claude_desktop_config.json (Windows) · ~/Library/Application Support/Claude/claude_desktop_config.json (macOS). Fully quit and reopen Claude afterwards.
Windows note: Claude Desktop does not always see your PATH — if the server fails to start ("spawn uvx ENOENT"), use the absolute path, e.g. "command": "C:\\Users\\<you>\\.local\\bin\\uvx.exe" (where uvx prints it). Logs: %APPDATA%\Claude\logs\mcp-server-yueying.log (~/Library/Logs/Claude/ on macOS). Keep wait_seconds at its default there; see the RUNNING rule.
Claude Code
bash
claude mcp add --transport stdio --scope user yueying --env PYTHONUTF8=1 -- uvx yueying mcp
Or drop this repo's .mcp.json into a project (it ships with "timeout": 1800000 so one watch_video call can wait for a long video). To raise Claude Code's tool timeout globally, set MCP_TOOL_TIMEOUT=1800000 (ms) in your environment. The repo is also a Claude Code plugin (.claude-plugin/plugin.json: server + skill).
Cursor
Click the Add to Cursor badge above, or put the same JSON in ~/.cursor/mcp.json (global) or .cursor/mcp.json (project):
MCP Servers → Configure (cline_mcp_settings.json). timeout is in seconds; the five read-only tools are safe to auto-approve. Step-by-step agent instructions: llms-install.md.
Direct links for hosts that accept custom URL schemes: cursor://anysphere.cursor-deeplink/mcp/install?name=yueying&config=eyJjb21tYW5kIjoidXZ4IiwiYXJncyI6WyJ5dWV5aW5nIiwibWNwIl0sImVudiI6eyJQWVRIT05VVEY4IjoiMSJ9fQ== and vscode:mcp/install?%7B%22name%22%3A%22yueying%22%2C%22command%22%3A%22uvx%22%2C%22args%22%3A%5B%22yueying%22%2C%22mcp%22%5D%2C%22env%22%3A%7B%22PYTHONUTF8%22%3A%221%22%7D%7D.
Without uv (pip / pipx) and Windows one-click
bash
pip install yueying # or: pipx install yueying
yueying mcp --setup # prints a config with the absolute path of the yueying-mcp executable
Use that absolute path as "command" with no args (Windows: ...\Scripts\yueying-mcp.exe; also works as python -m yueying mcp). Windows users without Python tooling can double-click install.cmd from a checkout: it creates %LOCALAPPDATA%\yueying\venv, installs the Claude Code skill, runs yueying mcp --setup and prints the JSON block with the right path.
GPU
bash
uvx --from "yueying[cuda]" yueying mcp # NVIDIA: adds the CUDA runtime wheels (cuBLAS, cuDNN)
pip install "yueying[cuda]"
Device and model are chosen automatically (model=auto: large-v3-turbo on CUDA, small on CPU); if the GPU trial fails, recognition falls back to CPU by itself.
The image is CPU-only (containers get no GPU by default), so it defaults to the small model.
Mount your videos read-only and give the tools container paths (/videos/lesson.mp4); results and
the downloaded Whisper weights live in the /data volume. In a client config the command is
docker and args are ["run", "--rm", "-i", "-v", "yueying-data:/data", "-v", "/your/videos:/videos:ro", "yueying"].
First call for any video: an absolute local path or a URL. mode: full (transcript + keyframes), transcript, frames.
DONE overview: title, source, duration, text source, folder, files, chapters, contact-sheet ranges, transcript in [mm:ss] paragraphs — or RUNNING with stage/percent/ETA, or ERROR with a plain-English hint.
Blocks up to wait_seconds (0–1500). Transcript truncated at max_chars with a get_transcript start time. Cached per video; refresh=true reprocesses.
video for the read tools accepts the video_id from watch_video/list_videos, the results folder, or the same path/URL you gave watch_video.
The RUNNING rule
Processing can take minutes, and most hosts cap a tool call at about a minute. So watch_video waits at most wait_seconds, then answers RUNNING video_id=… · stage 2/4 speech recognition 40% · elapsed 46 s · est. ~1 min remaining. The agent simply calls watch_video again with the same video — it re-attaches to the same job (options are ignored while it runs; refresh=true restarts). Recommended wait_seconds:
Host
wait_seconds
Why
Claude Desktop
45 (default)
hard ~60 s client timeout
Cursor
45 (default)
60–120 s
Claude Code
up to 1500
with .mcp.jsontimeout / MCP_TOOL_TIMEOUT = 1800000 ms
Cline
up to 1500
with "timeout": 1800 (s)
Progress notifications are sent every 1.5 s for hosts that display them. One video is processed at a time per server; extra requests queue.
First run: the first speech recognition downloads a Whisper model once (~480 MB small on CPU, ~1.6 GB large-v3-turbo on GPU). uvx yueying mcp --setup does this ahead of time; otherwise the RUNNING line says "first run downloads ~… this can take several minutes".
Where files go
Root: $YUEYING_OUT_DIR if set, else ~/yueying_out. One entry per video, named <slug>-<video_id> (yt-<id>, bili-<BV>, or the file name) — never renamed; the title lives in manifest.json.
code
~/yueying_out/
└── bili-BV1xx-3f9a2c1e/
├── report.md transcript.txt transcript.srt manifest.json
├── grid_01.jpg … grid_07.jpg
├── frames/ f001_00m02s.jpg … (+ frames/extra/ for get_frame_at)
├── .job only while a job runs
└── _download/ only with YUEYING_KEEP_SOURCE=1
Environment variable
Meaning
Default
YUEYING_OUT_DIR
root folder for results (absolute, ~ ok)
~/yueying_out
YUEYING_MODEL
default for the model parameter
auto
YUEYING_DEVICE
auto / cuda / cpu
auto
YUEYING_LANG
language of report.md written by the server (en/zh)
en
YUEYING_COOKIES_FROM_BROWSER
default browser for cookies (chrome, edge, firefox, …)
unset
YUEYING_KEEP_SOURCE
1 keeps the downloaded ≤720p source in _download/ (enables exact-moment frames for URLs)
unset
YUEYING_MAX_JOBS
pipelines running at once per server
1
YUEYING_JOB_TIMEOUT
hard limit per video, seconds
7200
HF_HOME
Hugging Face cache (Whisper weights live here)
HF default
HF_ENDPOINT
mirror, e.g. https://hf-mirror.com
huggingface.co
PYTHONUTF8
set to 1 on Windows to avoid mojibake
—
Disk budget: ≈ 25 MB per hour of video; 300–600 MB/h more with YUEYING_KEEP_SOURCE=1. Nothing is deleted automatically — list_videos shows sizes; delete a folder to free space; watch_video(refresh=true) reprocesses one video. Editing a local file changes its size/mtime and therefore gets a new entry.
Supported sources
Local files: anything ffmpeg reads — mp4, mkv, mov, webm, avi, flv, ts, and audio (mp3, m4a, wav, …). Audio-only input gives a transcript without frames.
URLs: every site yt-dlp supports. Fetched at ≤720p and deleted after processing unless YUEYING_KEEP_SOURCE=1.
Bilibili: without login Bilibili serves 480p — enough for slides and code. For HD or member-only videos pass cookies_from_browser="chrome" (or edge, firefox, brave, chromium, safari); on Windows close Chrome first, it locks its cookie database. Multi-part videos and collections: p= links are separate entries; the CLI's --all processes them all.
Not for live streams or images. Short links (b23.tv, v.douyin.com) are processed but not de-duplicated against their long form (the server never resolves URLs itself).
Also a CLI and an agent skill
bash
yueying video.mp4
yueying "https://www.bilibili.com/video/BVxxxx" --ui-lang en
yueying "https://www.youtube.com/watch?v=xxxx" --out ./notes/xxx
yueying lesson1.mp4 lesson2.mp4 "https://www.bilibili.com/video/BVyyyy"# several at once, one folder each + index.md
yueying "https://www.bilibili.com/video/BVxxxx" --all # every part of a multi-part video / collection
yueying --install-skill # Claude Code skill -> ~/.claude/skills/yueying
Default output: ./yueying_out/<name>/ (parent folder when several inputs). CLI log lines and report.md are Chinese by default (--ui-lang en for English); 0.3 will flip the default to English.
Flag
Meaning
--out DIR
output folder (default ./yueying_out/<name>; the parent folder when several inputs)
--all
when the URL is a Bilibili multi-part video / collection / playlist, process every entry (default: only the first)
--lang zh
spoken language code (zh, en, ja, …); default auto-detect
--model auto
Whisper model: auto / tiny / base / small / medium / large-v3 / large-v3-turbo (CLI default). auto = large-v3-turbo on an NVIDIA GPU, small on CPU
--device cpu
force CPU (auto / cuda / cpu)
--interval 3
roughly one keyframe every N seconds. Default by duration: 2 s under 1 min, 3 s under 3 min, 6 s under 10 min, 12 s under 30 min, 20 s beyond
--frames 30
maximum number of keyframes (default by duration, cap 150; 300 with --interval)
--scene 0.2
scene-change sensitivity 0–1, lower = more sensitive (default 0.3)
--no-dedupe
keep frames that are almost identical to the previous one (default drops them: < 2 % of thumbnail pixels changed)
--no-asr
no speech recognition even without subtitles (pictures only)
--no-frames
no keyframes (text only)
--force-asr
run speech recognition even when subtitles exist
--cookies-from-browser chrome
download with your browser login (Bilibili HD / member videos, sign-in-gated YouTube)
--keep
keep the downloaded source video
--ui-lang en
language of report.md headings and labels: zh (default) or en
--json
print one line of manifest JSON at the end (for scripts)
--install-skill
install the agent skill into ~/.claude/skills/yueying
mcp
run the MCP server (mcp --setup, mcp --check, mcp --version)
The skill (src/yueying/skill/SKILL.md) tells a coding agent to prefer the MCP tools when present and otherwise run the CLI and read report.md plus the contact sheets. Tools that support the Agent Skills standard can copy ~/.claude/skills/yueying/SKILL.md into their own skills folder.
Compared with similar projects (September 2026)
yueying
claude-video
claude-real-video
mcp-video-analyzer
MCP server
yes
no (skill only)
no (skill)
yes (Node)
Offline speech recognition
yes — subtitles first, local faster-whisper otherwise
cloud Whisper fallback
ASR-first
whisper installed separately
GPU auto-detect + CPU fallback
yes
–
–
–
ffmpeg bundled
yes
–
manual ffmpeg
–
Windows tested
yes (Windows 11)
–
–
–
Contact sheets (3x3)
yes
–
–
–
Burned-in timestamps on frames
yes
–
–
–
Bilibili / Douyin / Xiaohongshu
yes
–
–
–
"–" means the project did not advertise the feature when we looked; check their READMEs, they may have moved on.
Privacy policy
yueying collects nothing and has no telemetry, analytics, crash reporting or update checks. All processing is local. The only network connections are (1) to the video site of the URL you pass, through yt-dlp, and (2) to Hugging Face (or HF_ENDPOINT) to download a Whisper model once. Outputs are stored in your folder until you delete them. The transcript and any frames you request are sent only to the model your MCP client is configured to use — that transfer is governed by your client's and provider's terms, not by yueying. Questions: GitHub issues. Full text: docs/privacy.md.
Troubleshooting / FAQ
"No result received" in Claude Desktop. Keep wait_seconds at 45 (the agent then re-calls watch_video), and run uvx yueying mcp --setup once so the first call is not also the model download.
spawn uvx ENOENT / server fails to start. The host cannot see your PATH: use the absolute path to uvx (where uvx / which uvx) or to yueying-mcp as "command".
Console windows flash on Windows. Update to 0.2.0+: child processes are started without a window. If you still see them, you are running an old install (uv cache clean yueying).
Mojibake / ????? in titles. Add "env": { "PYTHONUTF8": "1" } to the server config (all snippets above include it).
Bilibili error 412 / "-352". The site wants a login: cookies_from_browser="edge" or "chrome" (close Chrome first on Windows).
YouTube "Sign in to confirm you're not a bot". Same fix: cookies_from_browser. Also try updating yt-dlp: uv cache clean yueying or pip install -U yt-dlp.
Slow on CPU.model="small" is already the automatic choice without an NVIDIA GPU; use mode="frames" when only the pictures matter, or mode="transcript" to skip keyframes.
Names, numbers and code are wrong in the transcript. Expected with any ASR — the agent is told to trust on-screen text; ask it to get_frame_at the moment.
Model download is slow or blocked (mainland China). Set HF_ENDPOINT=https://hf-mirror.com in the server env (the CLI switches to the mirror automatically when huggingface.co is unreachable).
GPU error (CUDA / cuDNN / out of memory).model="small" or YUEYING_DEVICE=cpu; install the CUDA wheels with yueying[cuda].
Want to reprocess with different settings.watch_video(video=…, refresh=true, …) — it kills a running job for that video, deletes the entry and starts again.
The server process never loads yt-dlp, Whisper or CUDA itself; all heavy work runs in a child process that is killed with the server. Everything lives in src/yueying/:
File
Responsibility
cli.py
command-line entry; runs the pipeline; mcp subcommand dispatch; --install-skill
mcp_server.py
the MCP server: six tools, job runner, progress parsing, DONE/RUNNING/ERROR rendering, --setup