LLM-operable Midjourney pipeline: 21 MCP tools to compose, generate, curate with vision, and log.
The Model Context Protocol (MCP) serverio.github.laffeyp/cascade-img provides an LLM-operable Midjourney pipeline. It exposes 21 MCP tools to compose, generate, curate with vision, and log results. The project is described as a system for midjourney-automation controlled through conversational prompt cards.
π οΈ Key Features
LLM-operable Midjourney pipeline
21 MCP tools for compose, generate, curate (vision), and log
Vision-assisted curation
π Use Cases
Image generation workflows driven by LLM prompts
Automation and orchestration for Midjourney-based tasks
Building AI agents for generative image production
β‘ Developer Benefits
MCP server and tool-based integration
Uses topics covering CLI, Python, and discord automation (discord-bot)
β οΈ Limitations
Scope is tied to a Midjourney pipeline (βdirect Midjourney by conversationβ)
cascade-img β direct Midjourney by conversation: a director writes a prompt card, a helper robot carries it to a colossal artist robot forging framed paintings, and hands the cut-out winner back
Generate Midjourney images by conversation instead of by hand. You describe what you want; your AI assistant composes the prompt, fires it, inspects the grid with vision, crops the best quadrant, cleans it up, and logs what worked.
code
You: "I need a flat-design mountain icon, centered, simple shapes, transparent background"
Agent: reads prompt log β composes prompt from parts β fires imagine β
waits β inspects 2x2 grid with vision β picks best quadrant β
crops it β removes background β saves β logs what worked
cascade-img is an MCP server with 23 tools that plugs into Claude, Cursor, Codex, or anything that speaks MCP. Midjourney is the first backend; Flux, DALL-E, and Imagen are on the roadmap. There's also a CLI.
Not a programmer? Open an AI assistant that can run commands (Claude Code, Cursor, or Cline), point it at this repo, and say: "Read RUNBOOK.md and set up cascade-img on this machine, then let me make images by describing them to you." It does the technical parts. You just need a Midjourney subscription and to copy a few values from Discord.
Quick Start
You need: a paid Midjourney subscription, a Discord account with the MJ bot in a channel, and Python 3.12 or newer.
bash
pip install cascade-img
Or from source:
bash
git clone https://github.com/laffeyp/cascade-img
cd cascade-img/packages/python
pip install -e .
This puts three commands for operating cascade-img on your PATH: cascade-mj-bridge (the daemon), cascade-mcp (the MCP server), and cascade-mj (the CLI). Installing also adds a fourth command, cascade-trace-check β a diagnostics validator (not part of the generation loop) that replays a recorded event log and checks it against the vocabulary's declared event ordering and timing rules.
Configure β you need four values from the Discord desktop app (channel ID, server ID, imagine version, and your user token). Takes about five minutes. RUNBOOK.md walks through each one step by step.
bash
cp"$(python -c 'import cascade_img, pathlib; print(pathlib.Path(cascade_img.__path__[0]) / ".env.example")')" .env# Fill in the four values per RUNBOOK.md, then validate:
cascade-mj-bridge --check-env --pretty
Start the daemon in one terminal, then connect from another:
bash
cascade-mj-bridge # leave running β holds the Discord connection
Connect your AI assistant β add to your MCP config (Claude Desktop, Cursor, Cline):
Or point your assistant at this repo and ask it to read AGENTS.md β it'll wire everything up.
Or use the CLI:
bash
echo'{
"mountain-icon": {
"subject": "a flat-design icon of a mountain, centered, simple shapes",
"aspect_ratio": "1:1"
}
}' > assets.json
cascade-mj mountain-icon --registry assets.json --upscale all --pretty
The 23 Tools
Category
Tools
What they do
Onboarding
cascade_guide
Returns the full operating manual in one call β the loop, every tool, the failureβaction table. Call it first; the generation and curation tools are gated until it's read.
Compose and fire prompts, poll for results, check daemon health, trigger Midjourney actions (upscale, vary, pan)
Catch-up
channel_recent, adopt_message
See what the human did by hand in Discord and claim those results into the pipeline β adopted messages become normal jobs that curation and mj_action work on
Composition
compose_prompt, compose_video
Build prompts from structured parts β subject, moodboard, style refs, aspect ratio, negatives β not freeform text
Extract quadrants from grids, remove backgrounds, trim whitespace, build sprite sheets, score results with vision, promote winners to final output
Working memory
log_append, read_prompt_log
Append-only prompt log the agent reads before every run β what was tried, what worked, what didn't. Persists across sessions.
Every call returns {ok, result} or {ok: false, error: {code, remediation}}. Branch on the stable code, not the message. Full tool reference in AGENTS.md.
How This Differs
Other open-source Midjourney tools focus on the generation step β fire the prompt, hand back the image. cascade-img does the work around that:
Vision-based self-curation β the agent inspects its own output and picks the best quadrant
Structured prompt composition β prompts built from parts (subject, style, identity, constraints), not raw strings
Working memory β append-only log persists across sessions; each run reads what came before
MCP-native β 23 tools that plug into Claude, Cursor, Codex, or anything that speaks MCP
Pluggable backends β Midjourney now, Flux/DALL-E/Imagen on the roadmap
How It Works
One daemon, two stateless clients, all over local HTTP:
cascade-mj-bridge β the daemon. Only process that talks to Discord. Holds the live connection and tracks in-flight jobs. Must stay running.
cascade-mcp β the MCP server. Stdio by default (Claude Desktop / Cursor / Cline); --http <port> for HTTP. Stateless β start and stop freely.
cascade-mj β the CLI. Takes an asset ID and a registry, composes the prompt, fires, waits, writes to the log.
Prompts are composed from structured parts, not written as raw strings:
python
from cascade_img.prompt.composer import PromptComposer, Subject, StyleStack, IdentityStack
prompt = PromptComposer().compose(
Subject(
text="a flat-design icon of a mountain",
constraints=["centered", "simple shapes", "transparent background"],
),
# Both optional. moodboard is a Midjourney personalization code;# sref/oref are reference-image URLs you'd set up in MJ first.
style=StyleStack(moodboard="abc123def", sref="https://cdn.example.com/style.png"),
identity=IdentityStack(oref="https://cdn.example.com/ref.png", ow=1000),
aspect_ratio="1:1",
version="7",
)
All three entry points emit structured JSON and follow the same {ok, result | error: {code, remediation}} envelope. Every failure carries a stable error code (e.g. DISCORD_401, MJ_UUID_MISSING, UPSCALE_BUTTON_FAILED) with a machine-readable remediation β so a caller branches on the code, not the message. The full catalog of log events and error codes is in vocabulary/0.1.json, and a trace checker (cascade-trace-check) enforces event ordering over recorded runs.
Midjourney terminology
prompt β the text + flags you send Midjourney. grid β the 2x2 set of four candidates returned per prompt. quadrant / U1-U4 β the four cells; "U2" means upscale the second. upscale β render one cell at full resolution. aspect ratio (--ar) β output shape. sref β an image whose style to borrow. oref β an image whose subject identity to keep across poses. moodboard (--p) β a saved personalization profile. stylize (--s) β how strongly MJ applies its own aesthetic.
Channel catch-up + message adoption β landed on main: channel_recent and adopt_message let the agent see what the human did by hand in Discord and act on it (design); still to come: more MJ commands (/describe, /blend, Vary Region inpaint, /tune), retro-U-press on adopted grids, internal refactoring
Every backend implements one interface, so a later release can chain them β generate on one provider, refine on a second (e.g. Flux Kontext), upscale on a third.
Repository Layout
code
cascade-img/
βββ packages/python/ # the Python package (cascade_img)
β βββ src/cascade_img/ # prompt/, interfaces/, backends/, curation/, vocabulary/
β βββ tests/ # behavior tests
β βββ tools/ # live smoke walk
βββ examples/ # three walkthroughs of the operating loop
βββ vocabulary/0.1.json # event log-line catalog
βββ *.md # README, ARCHITECTURE, RUNBOOK, AGENTS, CAPABILITIES, ...
Disclaimer
This tool automates Midjourney through a Discord user account. A paid Midjourney subscription is required. Both Discord and Midjourney's Terms of Service prohibit user-account automation. This is the same mechanism used by every open-source MJ tool (midjourney-proxy, midjourney-api, etc.) β there is no public Midjourney API. Use at your own risk.
The backend interface is pluggable β Flux, DALL-E, and Imagen are on the roadmap.