Turn URLs, files and text into media, transcripts, summaries, translations and PDF.
io.github.LatentNoise/content MCP Server
The io.github.LatentNoise/content server turns URLs, files, and text into media, transcripts, summaries, translations, and PDF documents. The project describes a single self-hosted engine that can be accessed via multiple interfaces, including an MCP server alongside other clients and APIs.
π οΈ Key Features
Processes URLs, files, and text
Produces media, transcripts, summaries, translations, and PDF
Uses one self-hosted engine
Exposes an MCP server as one integration point
π Use Cases
Converting web links or provided content into media
Generating transcripts and summaries from text/inputs
Producing translated outputs from URLs/files/text
Creating PDF documents from supported inputs
β‘ Developer Benefits
Works via an MCP server interface
Built on an underlying REST API (listed as βunderneath them allβ)
Available alongside a CLI, typed SDK, and REST API
β οΈ Limitations
Capabilities listed are limited to turning URLs/files/text into the documented output types (media, transcripts, summaries, translations, PDFs).
Turn URLs, files, and text into media, knowledge, and documents.
One self-hosted engine. HomeTube client for YouTube β Content Studio for everything else β HomeTube browser extension β MCP server β CLI β Typed SDK β REST API underneath them all
Content is a self-hosted engine that turns supported sources into media,
transcripts, summaries, translations, images, metadata, Markdown, PDF, and
more. The backend analyzes each source, determines what it can actually
produce, plans and runs the job, records its history, and delivers the results.
Everything else is a client of one public contract: HomeTube for the
focused YouTube experience, Content Studio for general workflows, a
Chromium extension for the tab you are already on, the content-mcp
server for agents, a CLI for terminals and cron, a typed Python SDK for
applications, the REST API for every other language, and Content
Console to watch the engine work. Pick one or several β none of them is a
layer the others have to pass through, and none of them holds business logic of
its own.
HomeTube is the quickest way to see what that means: paste a YouTube URL,
pick what you want out of it, and watch the files land in your library.
HomeTube demo β paste a YouTube URL, choose the outputs, watch the job deliver into the library HomeTube in Content β paste a YouTube URL, choose what you want,
follow the job into your library. More about HomeTube β
NOTE
A proven workflow, now growing into a platform. The standalone
HomeTube container has passed
300,000 package downloads on GitHub Container Registry. It keeps running
and stays maintained β nothing breaks, and there is no deadline. Content is
where that work continues, and where its users are invited to move at their
own pace.
The two are separate projects: standalone HomeTube has not been retrofitted to
run on Content. The HomeTube app in this repository is a new UI on the
Content engine, carrying the same workflow forward.
Coming from standalone HomeTube? β
Quick start
Docker Compose is the only runtime prerequisite. Install from the published
images β nothing to clone or build:
bash
mkdir content && cd content
curl -fsSLO https://raw.githubusercontent.com/LatentNoise/content/main/deploy/docker-compose.yml
curl -fsSL -o .env https://raw.githubusercontent.com/LatentNoise/content/main/.env.example
docker compose up -d
Open Content Studio for the general workflow, or HomeTube if your
source is YouTube. You can also skip both web apps entirely and drive the same
engine from the MCP server, CLI, SDK, extension, or REST API.
Everything stays in the installation folder: ./data holds the database, jobs,
and artifacts; finished files are delivered to ./playground/output. Set
CONTENT_DELIVERY_DIR_HOST in .env to point at your own NAS or media library
instead β its sub-folders then become the destination choices offered in the
clients.
Update later with:
bash
docker compose pull && docker compose up -d
Build from source instead β for development or to include the
optional speech-to-text runner
bash
git clone https://github.com/LatentNoise/content.git
cd content
cp .env.example .env
docker compose up -d --build
The repository's docker-compose.yml adds local build: definitions beside
the same images. deploy/docker-compose.yml is
the build-free deployment file used above; a test keeps the two aligned.
One source, many artifacts
The same engine serves media and document workflows:
text
ββ My Conference.mp4
ββ My Conference - audio.opus
YouTube URL βββ Content ββββββΌβ My Conference - subtitles - en.srt
ββ My Conference - subtitles - fr.srt
ββ My Conference - transcript.txt
ββ My Conference - summary.md
ββ My Conference - summary.pdf
Web page / text / .md file βββ Content βββ¬β Article.txt
ββ Article - summary.md
ββ Article - translation.md
ββ Article.pdf
Each branch starts with a declarative request. You describe the results;
Content resolves what is possible, plans the work, runs the available tools,
and records where every artifact came from. No yt-dlp flags, ffmpeg pipelines,
transcription glue, or LLM orchestration in the client.
The clients
One engine, one contract, several independent front doors. Every client below
speaks the same GenerationRequest, sees the same resolved capabilities, and
produces the same artifacts β they differ in ergonomics, not in what they can
ask for. HomeTube is entirely optional; so is every other row.
Paste a URL, choose the media and related artifacts, and follow the job into
your library β the workflow demonstrated at the top of this page.
With HomeTube in Content, you can:
download a video or playlist as video or audio, with quality, codec,
container, audio-language, and subtitle choices;
remove or mark sponsored segments with SponsorBlock, cut clips, use
server-side cookie credentials, and embed subtitles, chapters, thumbnails,
and metadata;
give every artifact a readable name and deliver it into a filesystem library
watched by Plex, Jellyfin, or Emby;
ask the same source for a transcript, summary, thumbnail, or metadata β not
just the downloaded media β when the required runners are available.
HomeTube is deliberately focused: it does not accept file, upload, or
inline-text sources, and it does not expose the broader document workflow or
PDF output. For those, use Content Studio, the API, CLI, or SDK. HomeTube is
a Content client, not a layer every Content user must pass through.
It has no settings of its own. It reads the engine's configuration, so what
changes HomeTube lives in the .env beside your docker-compose.yml: the
delivery library whose sub-folders become the destination choices
(CONTENT_DELIVERY_DIR_HOST), the languages pre-selected for audio and
subtitles (CONTENT_LANGUAGE_PRIMARY, CONTENT_LANGUAGES_SECONDARIES,
CONTENT_VO_FIRST, CONTENT_LANGUAGE_PRIMARY_INCLUDED_IN_SUBTITLES), and the
server-side cookie file that unlocks age-restricted or members-only videos
(CONTENT_CREDENTIALS β HomeTube only ever shows its id; the file never leaves
the server). Nothing is ever pre-selected that the source does not offer, and
no default is final.
HomeTube's README documents each
variable and works the language rules through a concrete example β a Japanese
talk, a French speaker, and exactly what ends up pre-checked. The
full variable table
lists everything the engine accepts.
Coming from standalone HomeTube?
HomeTube β the standalone
Streamlit app β keeps running and stays maintained. Nothing breaks, and there
is no deadline. Content is where the work continues, so this is a move you make
when you are ready, not one you are forced into.
What comes across. The workflow you know: paste a URL, choose quality, codec,
container, audio languages and subtitles; remove or mark sponsored segments with
SponsorBlock; cut clips; unlock restricted videos with a server-side cookie file;
embed subtitles, chapters, thumbnails and metadata; and deliver readable file
names into a library watched by Plex, Jellyfin, or Emby. Playlists come across
too β each entry is analyzed and planned on its own, named per item, and
delivered as its own artifact.
What is new. The same source can also produce a transcript, a summary, a
translation, generated thumbnails, metadata, Markdown, or PDF. And the engine is
no longer reachable only from a web form: Studio, the Chromium extension, an MCP
agent, the CLI, the SDK, and plain HTTP all ask for the same things.
What has not made the trip yet. HomeTube's playlist synchronization β the
plan/apply diff that keeps a local folder in step with a playlist over time
(rename detection, archiving or deleting removed videos, relocation after a
naming or location change) β does not exist here. Content downloads a playlist;
it does not yet re-sync one. It is roadmapped as part of
M2. If playlist sync is your main use, keep standalone
HomeTube for that job.
One practical note. Both HomeTube apps default to port 8501. Change one of
them in your .env if you want to run the two side by side.
The general-purpose web app: several sources in one request, every output type
the engine can resolve, with their options, preferences, and constraints. The
form is capability-driven β it renders what /capabilities answers for
your source, so a capability the engine gains appears without a UI release, and
what cannot be produced is shown with the server's reason instead of silently
disappearing.
Content Studio β the general-purpose request builder Content Studio β build general URL, file, upload and text requests
Studio takes a file two ways, and the distinction matters when the engine runs
on another machine:
From this device β the bytes are sent to the engine and become an
upload source (ADR 0020).
Several files at once become several sources. Streamlit buffers the upload in
the UI container, so the ceiling here is deliberately lower than the API's
(200 MB by default, and the picker states the current value rather than
letting you discover it by failing).
On the server β a path the engine can already read, under a configured
allowed input root. With this Compose setup that root is ./playground/input.
For genuinely large files, skip the browser: the SDK streams from disk with
client.upload_file(path).
Send the page you are watching to your engine without opening a UI. Manifest
V3, Chromium only (Chrome, Brave, Edge, Vivaldi, Opera, Arcβ¦), no build step β
what the browser loads is exactly the files in the zip:
download content-browser-extension-chromium-v<version>.zip from the latest
release and unzip it (Chromium loads a folder, not a zip);
open chrome://extensions, turn on Developer mode, Load unpacked,
select the folder;
open a video and click the extension.
It normalizes the tab's URL (youtu.be, Shorts, and embeds become the
canonical watch URL; a list= on a watch URL means that video, not the
playlist), asks the engine what the source can produce and offers only that,
prefills the name with the naming engine's own proposal, offers the library's
existing folders plus new folderβ¦, then follows the job and shows where each
file landed. Every network call happens in the service worker β the engine
sends no CORS headers by default, and that is what lets the extension work
against a stock instance with nothing to configure
(ADR 0016).
HomeTube for Content β the Chromium extension popup on a video page Browser extension β send the current tab to Content, from the tab itself
The popup's footer always names the engine it is talking to, so "where is this
sending my video?" is answered on screen. Its README also carries an honest
path-by-path verification table β what has actually been driven in a
browser, and what has not.
Content gives an MCP-compatible agent a controlled way to turn URLs into real
artifacts, not just talk about them. The official content-mcp server
exposes intention-level tools β not one per endpoint β to:
analyze a URL and discover what it can produce;
request one or several outputs and choose a library destination;
monitor the job, cancel it, and report actionable failures;
find the artifacts and their delivered paths;
read small text artifacts inline while keeping large media out of the model
context.
The agent never needs shell access, yt-dlp syntax, or backend internals.
Nothing to install.uvx fetches the server on first use and caches it, so
the client owns the whole lifecycle (updating stays explicit β uvx --refresh,
or pin content-mcp@x.y.z):
bash
# Claude Code
claude mcp add content \
--env CONTENT_API_URL=http://localhost:8010 \
-- uvx content-mcp
For Claude Desktop, Cursor, and other MCP clients using the standard JSON
shape:
Prefer a pinned executable on your PATH β for an offline machine, or to control
when the version changes? Install it and name it directly
("command": "content-mcp", or -- content-mcp for Claude Code):
Analyze this lecture, save its audio into Talks, create a transcript and a
structured summary if the available runners allow it, and tell me where
every artifact landed.
The normal tool journey is
get_config β analyze_source β generate β get_job β get_artifact. Downloads to
the agent's own machine are bounded by a single variable
(CONTENT_MCP_DOWNLOAD_DIR, ~/Downloads/Content by default): a destination
pointing outside it is refused, not clamped β widening that is the operator's
decision, not something a prompt can talk the server into.
What it can and cannot do, briefly β the
MCP guide has the full three
tables, and each row there was driven over stdio against a running engine
rather than inferred:
Works
URLs and playlists Β· local files (uploaded to the engine, so a laptop can drive a NAS) Β· PDFs, read for their text layer Β· video, audio, subtitles, transcripts, summaries, translations, chapters, thumbnails, metadata, Markdown, PDF Β· delivery into your library Β· cookie-authenticated sources
Does not
no live progress (get_job is a poll) Β· no .docx/.epub/.odt/.rtf, no OCR for scans Β· no transcoding Β· no playlist sync Β· nothing deletes anything Β· the API has no authentication
Needs a runner
Summaries, translations and derived chapters want a local Ollama or a cloud key; transcripts want subtitles, or the optional Whisper runner. The engine reports these as unavailable up front rather than failing halfway
By design
stdio, not an HTTP endpoint β the server runs where you do, which is what lets an agent hand it ~/Documents/report.pdf and have the bytes uploaded to an engine on your NAS. It also means no open port on a system that has no authentication. For a client that cannot spawn a process β Open WebUI, hosted UIs β mcpo bridges it today
uv tool install content-cli
export CONTENT_API_URL=http://nas.local:8010
content analyze "https://www.youtube.com/watch?v=β¦"
content video "https://β¦" --height 1080 --subs en,fr --watch
content audio "https://β¦" --format opus --playlist --watch
content submit request.json --watch # a raw GenerationRequest, or -
content jobs ; content job <id> ; content artifacts <id>
content download <artifact_id> -o out.mkv
--watch follows the job's event stream rather than re-asking for its
status, and the exit code carries the outcome so a script can chain on it
(ADR 0021):
0 succeeded, 2 partially succeeded, 1 failed or cancelled. 2 is
deliberately distinct β a playlist that yielded five videos of six is not a
failure, but a script that treats it as a success will quietly move on with
missing files.
The one official API client: the CLI, the MCP server, and the web apps all
speak through it, so the engine's rules are never duplicated. Python 3.11+,
httpx and pydantic only.
python
from content_sdk import ContentClient, outputs
with ContentClient("http://localhost:8010") as client:
analysis = client.analyze(outputs.url_source("https://www.youtube.com/watch?v=β¦"))
job = client.generate(analysis.id, [outputs.audio_output()])
job.wait()
for artifact in job.artifacts:
print(artifact.display_filename, artifact.delivered_path)
AsyncContentClient mirrors the whole surface. Analyses are addressable, so
every call accepts either an analysis_id or inline sources
(ADR 0014);
client.upload_file(path) streams a local file to the engine and hands back a
ready-to-use source; and every non-2xx becomes a typed exception carrying the
stable error codes (NotFound, Gone, ValidationError).
/api/v1 is the contract every client above is built on, and it is stable
enough to build a ninth client against: requests describe desired outputs
rather than technical operations, error codes are machine identifiers that do
not change (only their human messages do), and "valid but not implemented" is
a different, explicit answer from "invalid". Reserved fields are refused
rather than silently ignored. The V1 API has no built-in authentication β
see deployment and security.
The operations console: strictly observability and control, and deliberately
not a way to create downloads (that is Studio and HomeTube).
Content Console β jobs, runners, storage and configuration, live Content Console β observe and control the engine
Overview β version, cache, concurrency, analysis TTL, a live jobs pulse,
the installed runners with their availability, storage paths.
Environment β every CONTENT_* variable with its effective value,
whether it was set or fell back to a default, and a one-line description.
Secrets are never shown β only their presence and length.
Jobs β a filterable list plus full detail: steps, the submitted
GenerationRequest, artifacts with their provenance, ordered events,
per-step logs, cancel and retry, with an opt-in 5-second auto-refresh that
polls only while jobs are in flight.
Storage & cache, and a contract & API tab with the schemas, Swagger,
ReDoc, and a raw API tester.
What Content can produce
Content resolves availability for each analyzed resource and installed runner.
Clients show what can actually be produced instead of presenting a static list
and failing halfway through a job.
Output
Typical inputs and behavior
Video
Media URLs and local media files; stream selection, remux, fast or frame-accurate cutting, and SponsorBlock handling
Audio
Media URLs as source audio, Opus, MP3, or M4A; local files keep their native audio stream
Subtitles
Manual or automatic tracks selected by language
Transcript
Existing subtitles, or audio with the optional Whisper runner
Summary
Transcript or readable text through a local or opt-in cloud LLM
Translation
Subtitles or transcripts through an LLM; subtitle timings stay aligned
Chapters
Chapters declared by the source, or chapters derived from a transcript through an LLM
Thumbnail and keyframes
Published artwork or frames extracted from video
Metadata
Normalized, provider-independent resource information
Markdown and plain text
Readable web pages, text and Markdown files, PDFs (their text layer), and inline text
PDF
Readable sources or outputs such as summaries, transcripts, and translations
Media acquisition works with YouTube and the
sites supported by yt-dlp.
Playlists can produce artifacts for each member in a single traceable job.
Finished files can be delivered directly into any mounted filesystem library,
including folders watched by Plex, Jellyfin, or Emby.
Local and optional AI runners
Content has no mandatory cloud dependency. AI-backed outputs activate when a
compatible runner is available:
Summaries, translations, and derived chapters: connect a local
Ollama instance, or explicitly configure an Anthropic
or OpenAI API key. Cloud runners can be excluded per request.
Transcription from audio: run a speech service next to the engine β
docker compose --profile speech up -d and CONTENT_SPEECH_URL=http://speech:8000,
or speech.enabled: true in the Helm chart. It is
speaches, behind the OpenAI audio
API, and the same container also does text-to-speech. Audio is only ever sent
to a service on a private network. Transcripts from existing subtitles do not
need it.
Unavailable optional runners do not make Content unhealthy. The affected
capability is reported as unavailable, and an impossible request is refused
before the job starts. See
deployment and configuration for the full
inventory.
Why build on Content?
Declare intent, not tooling. Requests describe outputs; yt-dlp, ffmpeg,
Whisper, LLMs, and PDF renderers remain replaceable implementation details.
Self-hosted and local-first. Run it on a workstation, NAS, or homelab.
There is no account, telemetry, or mandatory cloud service.
One public contract. Studio, HomeTube, MCP, the CLI, SDK, extension, and
REST API converge on the same domain instead of drifting into parallel
feature sets.
Human files, not pipeline debris. The engine names every artifact and can
place it directly in the library you already use.
Observable work. Jobs expose states, ordered events, progress, logs,
artifacts, provenance, cancellation, and retry.
Honest capabilities. βValid but unsupported,β βunavailable on this
source,β and βbrokenβ are different answers, reported before or during the
correct stage.
How it works
The public request describes desired outputs. It does not prescribe technical
operations:
Content Backend (REST /api/v1) β domain, planning, execution, persistence
β
Content Python SDK β official Python API client
β β β
CLI MCP Web apps β thin clients of the public contract
Browser extension β JavaScript client of /api/v1
The engine analyzes source facts, resolves feasible capabilities against the
installed implementations, builds a deterministic execution plan, and runs it
as an observable job. The database is the source of truth; yt-dlp, ffmpeg,
transcription, LLM, and PDF implementations stay behind dedicated boundaries.
Content targets single-host Linux/macOS deployments on amd64 or arm64. Docker
Compose runs the API and embedded worker together, with the web apps as
separate API clients. SQLite and artifact storage persist locally; no Redis,
Celery, Kubernetes, or cloud infrastructure is required.
The commented .env.example covers UI selection, image
versions, ports, storage and delivery, language preferences, server-side
cookie credentials, optional LLM/STT runners, CORS, and notifications. No
release check makes an outbound request unless you configure one.
The V1 API has no built-in authentication: anyone who can reach it can submit
jobs, upload files and read every artifact. Keep it on a trusted network, or put
an authenticating reverse proxy in front of any externally reachable instance β
and do not publish the port. What an attacker who reaches the API can and cannot
do is written out in docs/operations/threat-model.md.
See deployment for configuration, health
checks, data layout, authenticated sources, and production guidance, and
ADR 0024
for why that is the V1 answer.
Development
The repository is a monorepo containing the backend, web applications, CLI,
MCP server, Chromium extension, and Python SDK. The root Makefile is the main
entry point:
bash
make install # create the development environment and install packages
make validate # formatting, linting, and hermetic tests: the official gate
make test-ui # hermetic Streamlit application tests
make validate-all # official gate plus UI and opt-in external-tool tests
In plain words: you may read, audit, modify and run Content β at home, for
yourself, or inside your organisation. The one thing you may not do is offer it,
or something substantially similar built from it, as a competing commercial
product or service.
Every version becomes Apache 2.0 two years after it is published. That
promise is in the licence itself and cannot be taken back.
Copyright 2026 Yann Orieult. "Content" and "Latent" are names of the project and
of its author's work; a fork must not use them for its own product.
Versions up to and including 0.8.4 were released under
AGPL-3.0-or-later and remain
available under it, permanently. See COMMERCIAL.md for the
full summary; no separate commercial offering is currently available.
Governance
Content is developed and maintained by Yann Orieult under a single-maintainer
model. Bug reports, ideas, and design feedback are welcome through issues. Code
contributions are not acceptedβunsolicited pull requests are closed without
review; forking is the intended path for independent changes. See
CONTRIBUTING.md and GOVERNANCE.md.
Report security issues privately as described in SECURITY.md.