AI image and video generation, editing, and region repair via Gemini, OpenAI, and Grok.
Pixel Surgeon MCP server enables AI image and video generation, editing, and region repair via Gemini, OpenAI, and Grok. It orchestrates multiple models for transplant-grade region fixes and creative edits, exposed as an MCP server. This description uses repository-provided name, description, readme excerpt, and related tooling to guide integration and discovery.
๐ ๏ธ Key Features
AI image and video generation, editing, and region repair
MCP server for AI image & video generation, editing, and transplant-grade region repair
Powered by Gemini 3.1 Flash Image, OpenAI GPT Image 2, Grok Imagine, and Veo 3
An MCP server that gives Claude (or any MCP client) the ability to generate images, edit them, fix garbled text, and create videos โ all through natural language.
How it works
pixel-surgeon-mcp is a multi-provider image generation server. You can use any combination of providers and switch between them per-request:
Gemini (Google) โ balanced
Google's image generation pipeline uses a two-stage approach: Gemini 3.1 Pro reasons about your prompt, then Gemini 3.1 Flash Image renders the pixels. Supports 9 aspect ratios at 512/1K/2K/4K resolution. Best price/performance ratio, with a free tier available.
OpenAI GPT Image 2 โ highest quality
OpenAI's latest image model with dramatically improved text rendering and visual fidelity. Supports flexible resolutions โ pixel-surgeon maps your chosen size and aspect ratio to the optimal pixel dimensions automatically. Quality levels: medium (fast) and high (print-ready). Excellent for infographics, diagrams, and text-heavy images where other models struggle. Slower and more expensive.
Grok Imagine (xAI) โ fastest
xAI's Aurora-powered image model. Fastest generation speed and lowest cost. Supports 7 aspect ratios at fixed resolutions (~1K). Good for rapid prototyping and iteration.
Veo 3 (Video)
For video, the server calls Veo 3 with async polling โ generating both video and ambient audio. Supports 16:9 and 9:16 at 5s or 8s duration.
Region repair
AI image models struggle with text-heavy images. The fix tools solve this by sending smaller regions to the provider, then stitching the results back with histogram-matched compositing for seamless blending.
Tools
Tool
Description
generate_image
Text-to-image generation (single image)
generate_images
Parallel batch generation (1-8 images)
generate_video
Text-to-video via Veo 3 with audio (5s or 8s)
edit_image
Edit an existing image with natural language instructions
fix_image
Grid-based tile repair for garbled text (2x2, 3x3, etc.)
fix_region
Targeted region repair with automatic aspect ratio snapping
Force a specific model per-call via the model tool parameter, or set DEFAULT_IMAGE_MODEL env var.
Gemini automatic fallback
If a Gemini generation call fails with a billing / prepay error, the server automatically retries on the free-tier gemini-2.5-flash-image model. The viewer shows a yellow banner when this happens. Free-tier limits: 1K max resolution, 10 RPM, 500 RPD.
Style presets
All generation and edit tools support an optional style parameter:
Duval Software's signature retro-futurist infographic style. 1960s Space Age meets 1980s arcade. Cathode blue, amber, and salmon palette. Great for diagrams and system overviews.
Prepayment required. Gemini 3.1 Flash Image and Veo 3 require billing and prepaid credits. The free-tier fallback (2.5 Flash) has limited resolution and rate limits. See Google AI pricing.
git clone https://github.com/j-east/pixel-surgeon-mcp.git
cd pixel-surgeon-mcp
npm install
npm run build
Image output
Generated images are saved to ~/Pictures/pixel-surgeon/. A local browser viewer auto-launches on first use for full-resolution previews with model selection, respin controls, and search.
Development
bash
npm run dev # tsx watch mode
npm run build # compile TypeScript
npm run start # run compiled server
Key implementation details
Aspect ratio snapping โ crops are adjusted to the nearest Gemini-supported ratio while preserving center point
Human-in-the-loop โ interactive_fix opens a browser crop UI, blocks via Promise until the user submits, fires parallel Gemini calls, and lets the user pick the best result
MCP size limits โ full-resolution images are saved to disk; downsampled versions (< 950KB) are returned in MCP responses
Contributing
PRs are welcome! We're especially looking for:
New style presets
Add entries to the STYLE_PRESETS object in src/index.ts. Your PR should include:
The preset definition (name, prompt prefix, default aspect ratio)
2-3 example images generated with the preset (drop them in your PR description)
A short description of the visual style for the README table
Model adapters
The server currently supports Gemini, OpenAI, Grok Imagine, and Veo 3. We'd love adapters for other image/video generation APIs โ Stable Diffusion, Flux, etc. If you're interested in adding one, open an issue first so we can align on the interface.
Built by Duval Software
pixel-surgeon-mcp is maintained by John Evans, part of the engineering team at Duval Software โ a software engineering firm in Jacksonville Beach, FL building AI-powered tools and custom integrations. If you need MCP servers, AI pipelines, or production tooling built, get in touch.
License
MIT
Install
Configuration
Environment variables
GOOGLE_API_KEYsecret
Google AI API key (enables Gemini image/video generation)