The prediction MCP — score your prompt before you generate, so you never waste a credit.
Dali by Lulu (io.github.Lulu-The-Narwhal/dali) MCP Server
Dali by Lulu is a “prediction MCP” that scores a prompt before generation to avoid wasting a credit. It is identified as io.github.Lulu-The-Narwhal/dali and provides 16 tools. Topics include ai, claude, generative-ai, mcp, mcp-server, model-context-protocol, and prompt-engineering.
🛠️ Key Features
Prediction step that scores the prompt before you generate
“Never waste a credit” positioning described in the server summary
🚀 Use Cases
Reducing unnecessary generation by evaluating prompts up front
Supporting prompt-engineering workflows for generative-ai and claude
⚡ Developer Benefits
Prompt pre-scoring to inform whether to proceed with generation
Fits into Model Context Protocol (model-context-protocol) server tooling
⚠️ Limitations
Only high-level behavior and tool count (16) are described in the provided excerpt; specific tools and inputs/outputs are not included
Score a prompt for a generation target (0–100) and, if it's weak, return the rewrite brief — in ONE call.
Reads intent with a fast heuristic keyword analyzer, scores the prompt, then:
• score ≥ 70 (A/B) → returns the scorecard and tells you to proceed.
• score < 70 (C/D/F) → returns the scorecard PLUS a rewrite brief so you can
fix it without a second call. Write the enhanced prompt from the brief, then
call track_enhancement(original, enhanced, generator).
Returns a ScoreCard (overall, grade A–F, per-dimension breakdown, what's missing,
anti-patterns, verdict) plus needs_enhancement, and enhancement_brief when weak.
category (optional): the ad vertical (e.g. "wellness", "beauty") — when set and
conversion priors exist, the brief upgrades to a conversion-justified rewrite.
Supported generators: veo3, higgsfield, midjourney, flux, kling, sora, imagen…
Get a rewrite brief for this prompt + generator. YOU write the enhanced prompt from the brief.
Returns a structured brief with score_before, rewrite_brief, and llm_instructions.
category (optional): the ad vertical (e.g. "wellness", "beauty") if known.
When set and conversion priors exist for it, the brief upgrades from craft
advice to a conversion-justified one, backed by real ad-performance data.
IMPORTANT: After you write the enhanced prompt, you MUST call
track_enhancement(original_prompt, your_enhanced_prompt, generator) immediately.
This is not optional — it records the improvement and is required for the graph to learn.
Record an enhancement pair in the Dali graph brain.
Call this AFTER you write an enhanced prompt from score_prompt's brief or enhance_prompt.
This records the before→after improvement so the graph learns which rewrites consistently
push scores up — enriching creative_patterns and community_benchmark over time.
Returns before/after scores so you can confirm the delta.
Recommend the best generator for your creative concept and per-generation budget.
Analyzes the concept's creative signals (motion, style, subject type, use case)
and matches them to generators within your budget. Returns a ranked list so you
can make an informed choice before scoring the actual prompt.
Parameters2
concept
string
required
What you want to make — subject, style, mood, format, use case
budget_usd_max
number
optional
Max USD per generation attempt (default $1.00)
Raw schema
{
"type": "object",
"properties": {
"concept": {
"type": "string",
"description": "What you want to make — subject, style, mood, format, use case"
},
"budget_usd_max": {
"default": 1,
"type": "number",
"description": "Max USD per generation attempt (default $1.00)"
}
},
"required": [
"concept"
],
"additionalProperties": false
}
score_variations
Score 2–8 prompt variations for the same generator and rank them best-to-worst.
Use this when you've drafted multiple versions of a prompt and want to pick the winner
without burning generation credits. Returns a ranked list with per-dimension comparison
so you can see exactly why one variant beats another.
Parameters2
prompts
array
required
List of 2–8 prompt variants (same creative intent, different wording)
Community graph: which patterns consistently produce high-grade prompts for this generator?
Powered by the V3 graph brain (Supabase PostgreSQL). Every scored prompt contributes.
Returns top patterns by type, enhancement unlocks, and cross-model universal patterns.
Find community A/B-grade prompts structurally similar to yours.
Uses graph traversal (Memgraph) to locate prompts that share the most
creative patterns with your input and scored A or B on the same generator.
Returns what those prompts did right — so you can adopt the same moves.
Use this when:
- Your prompt scored C or below and you want inspiration
- You want to see how the community solved the same creative problem
- You need concrete A-grade examples, not abstract advice
Parameters2
prompt
string
required
The prompt to find neighbors for.
generator
string
required
The generation model (veo3, midjourney, flux, etc.)
Raw schema
{
"type": "object",
"properties": {
"prompt": {
"type": "string",
"description": "The prompt to find neighbors for."
},
"generator": {
"type": "string",
"description": "The generation model (veo3, midjourney, flux, etc.)"
}
},
"required": [
"prompt",
"generator"
],
"additionalProperties": false
}
enhancement_path
Show the most reliable path from a bad grade to an A on this generator.
Mines the Dali graph for all F/D → A/B enhancement pairs and surfaces
the patterns that appear most consistently in the 'after' side.
These are the highest-ROI moves for this specific generator.
Use this when:
- A prompt just scored D or F and you're not sure what to fix
- You want to know which improvements matter most for a specific generator
- You want to understand generator-specific enhancement strategy
Parameters2
generator
string
required
The generation model (veo3, seedance, kling, etc.)
starting_grade
string
optional
The grade you're starting from — 'F', 'D', or 'C' (default 'F')
Raw schema
{
"type": "object",
"properties": {
"generator": {
"type": "string",
"description": "The generation model (veo3, seedance, kling, etc.)"
},
"starting_grade": {
"default": "F",
"type": "string",
"description": "The grade you're starting from — 'F', 'D', or 'C' (default 'F')"
}
},
"required": [
"generator"
],
"additionalProperties": false
}
dali_version
Current Dali MCP version and changelog.
Check this whenever you want to know what tools are available,
what changed in the latest release, or which version is running.
Score an actual ad IMAGE (not the text prompt) for conversion — before you spend.
Conversion lives in the pixels, so this scores the real creative and gives you
ONE answer combining two views, in a single call:
• HEADLINE score = how much it visually resembles PROVEN WINNERS (Vertex
embedding vs the live winner corpus). The sharpest predictor — it reads the
whole look and self-solves archetype (a premium ad resembles premium winners,
not scammy direct-response ones).
• WHAT TO CHANGE = the specific winning attributes it's missing (Gemini vision
vs category priors) — the actionable detail.
• DEFECT GATE = generation defects (extra fingers, garbled text, warped anatomy).
Use it on a generated image, a mockup, or any ad you're about to run.
Returns:
score — 0-100 headline: visual similarity to proven winners
verdict — one-line looks-like-a-winner / partial / rework call
looks_like — the real proven winners it resembles (advertiser, category, days-run)
what_to_change — high-lift winning attributes it lacks, each with a fix sentence
you_already_have — winning attributes it already has
has_defect/defects — generation defects to fix before shipping
detail — raw numbers {embedding_score, attribute_score} for transparency
category examples: beauty, supplements, wellness, fitness, food, apparel, tech, pets.
Leave category empty for a cross-vertical look-alike match + defect QA.
Score an ad creative YOU are looking at (e.g. a pasted/attached image) against
the winning corpus — no URL needed. Use this when the user shares an image in the
conversation: read the creative yourself and fill in what you see, and Dali scores
it against what wins in the category (3,800+ proven winners), returning the
conversion verdict and exactly which winning attributes it's missing.
You (the model) provide the visual read; Dali provides the winning-data scoring.
(For a fetchable image URL, prefer score_creative — it adds the embedding
similarity headline, which needs the real pixels.)
Fill these from looking at the image:
category — vertical: beauty, wellness, supplements, fitness, food, apparel, tech, pets
lighting — warm lighting | natural light | studio light | dramatic lighting | clinical bright | dark moody | neon
subject — single person | group | product only | no person | before after
subject_age — young adult | middle age | senior | child | none
format — ugc selfie | testimonial | product hero | lifestyle | chart infographic | text meme | comparison
text_density — none | light | heavy
dominant_emotion — calm | excited | trust | fear | aspiration | neutral
eye_contact — true if a person looks at camera
offer_visible — true if a price/discount/offer is shown
defects — list any generation defects (extra fingers, garbled text, warped anatomy); [] if clean
Returns: conversion_score (0-100), verdict, matched (winning attributes it has),
missing (high-lift attributes to add, each with a fix sentence), has_defect/defects.
Find YOUR winning ad formula from your own numbers — paste your ads export.
The category prior is a cold-start fallback; the real signal is what wins in
YOUR account. Paste an ads CSV (a creative image-URL column + a performance
column — CPA / CTR / ROAS / purchases) and Dali runs vision on your winners vs
losers and returns the attributes that separate them, plus how your account
compares to the industry median.
If an email is supplied, the formula is saved and emailed with a ready-to-paste
Claude prompt wired to Dali — so scoring the next creative is one step.
Returns:
formula — attributes over-represented in your winners (value, winner%/loser%, lift)
benchmark — your median vs the vertical's industry median (when category given)
analyzed — how many winners/losers were read, and the metric direction
saved — whether the lead+formula were captured (only when email supplied)
Score your creative against what's actually winning in the ad market — before you spend the credit.
Most AI generation failures are predictable. A weak prompt, an off-formula creative — you can't tell until after you've burned the token. Dali scores it first, and it doesn't grade against opinions or generic "prompt tips." It grades against a real, living corpus of proven-winning ads — creatives still running in the market months after launch, scraped, embedded, and ranked. Two jobs:
score_prompt — judge the prompt before you generate (craft: camera, motion, lighting, model-native language).
score_creative — judge the actual image against proven winners (does it look like what converts, and what's missing).
Every wasted generation has a real cost — a Seedance retry is ~$6. The live dashboard tracks what the community has saved by catching bad creatives before they burned a credit.
code
You: "make a video ad for our glass serum bottle"
dali::score_prompt(prompt, "veo3")
→ 8/100 Grade: F
→ no camera move · no motion · no lighting · 8 words
→ Verdict: Generic stock footage guaranteed. Enhance first.
→ enhancement_brief included (score < 70):
① lead with camera — Veo 3's #1 lever: "Slow dolly", "Orbital push"
② describe physics: "a drop falls", "liquid ripples", "glass refracts"
③ lighting type + quality: "warm backlight", "rim-lit edges"
↳ [Camera]. [Subject + motion]. [Lighting]. [Mood]. [No text.]
✦ YOUR LLM rewrites using the brief:
"Slow orbital push around a glass serum bottle on white marble. A single
amber drop falls in extreme slow motion, catching warm backlight. Macro:
liquid gold ripples outward from impact. Rim-lit edges, soft studio
diffusion. Premium, clinical. No text."
dali::score_prompt(enhanced, "veo3")
→ 91/100 Grade: A ✓ Safe to generate.
The real winning data layer
This is what makes Dali more than a prompt linter. The scores are grounded in real ads that are actually winning, not hand-written rules.
How the corpus is built — longevity is the outcome signal. We scrape the public Meta Ad Library. An ad still running months after it launched is one the advertiser keeps paying for — a proven winner. That "still-running-after-N-days" longevity is a market-validated label you can't fake, and it's the spine of the whole dataset.
The pipeline (offline → serving). The tools never scrape or embed on the fly — they read pre-built stores:
code
scrape Meta Ad Library → proven winners (longevity label)
→ Gemini vision → creative attributes (lighting, format, before/after, offer…)
→ prevalence SQL → winning-pattern lift per vertical (winners vs baseline)
→ Vertex embeddings → BigQuery VECTOR_SEARCH (nearest proven winners, cosine)
→ graph edges (Memgraph) → (:Pattern)-[:WINS_IN {lift, n}]->(:Category)
So when score_creative runs, it embeds your image and finds the actual winning ads it most resembles by full visual signature — then tells you which winning attributes you're missing. When enhance_prompt runs with a category, the rewrite brief is backed by real market lift ("before/after shows up in 78% of winning wellness ads, 4× baseline"), not craft opinion.
Honest scope. The winner label is longevity (a strong market-validated proxy), not per-ad conversion rate — measured CVR validation is in progress. The corpus grows on a schedule, so coverage per vertical keeps deepening. What you get today: your creative scored against what's demonstrably surviving in the live market.
code
dali::score_creative(image_url, "beauty")
→ score 62/100 — partial resemblance to proven winners
→ looks_like: Frøya Organics (ran 411d), tashportcosmetics (884d), Face Reality (346d)
→ what_to_change: winners use "before/after" 4× more · offer-visible 1.8× more
→ defects: none
→ Verdict: Partial — strong resemblance, but add the high-lift attributes before spending.
pip install dali-mcp
claude mcp add dali -- python -m dali.server
The self-hosted package exposes the prompt-scoring tools locally. The creative-scoring tools (score_creative, analyze_winning_formula) and the winning-ad corpus run on the hosted server — connect via the hosted MCP to use them.
Tools
Score the creative — against real winners
Tool
What it does
score_creative(image_url, category)
Score an actual ad image. Embedding similarity to proven winners is the headline score; also returns the winners it resembles, which winning attributes it's missing, and generation defects — in one call
score_creative_from_view(category, …)
Score an image you're looking at (pasted/attached in the chat) — no URL. The model reads the creative's attributes and Dali scores them against the winning corpus (verdict + what to change). Use for images shared in-conversation; score_creative (URL) adds the embedding headline
analyze_winning_formula(csv, category, email)
Paste your own ads export (creative URL + CPA/CTR/ROAS) → your winning formula vs your losers, plus how you compare to the industry median
Score the prompt — before you generate
Tool
What it does
score_prompt(prompt, model, category?)
Grade 0–100 with a per-dimension breakdown and verdict. When the score is weak, the rewrite brief is returned in the same call. Reads intent with the conversation LLM (understands negation, any language)
enhance_prompt(prompt, model, category?)
Returns a structured rewrite brief — YOUR LLM writes the enhanced prompt. With a category, the brief is backed by real winning-ad lift
track_enhancement(original, enhanced, generator)
Record a before/after pair in the graph brain — trains community patterns
score_variations(prompts, generator)
Rank a list of prompt variants in one call — highest to lowest
suggest_generator(concept, budget_usd_max)
Pick the best model for your concept + budget
The graph brain & meta
Tool
What it does
creative_patterns(model)
Community top patterns for this model from the graph
community_benchmark(prompt, model)
Compare your prompt against community top scorers
prompt_neighbors(prompt, model)
Find A/B-grade prompts that share your patterns (score the prompt first, so its patterns are in the graph)
Natural language + contentClass and style.presets API params
Imagen 4 (Google): deprecated — use gemini-3.5-flash with image output. Dali still scores legacy Imagen prompts via the imagen model key but don't build new things on it.
Platform supersets
Higgsfield and Runway are aggregator platforms — they proxy multiple underlying models under one API. The model you pick matters more than the platform name:
Platform
Model selector
Underlying model
Higgsfield
veo3
Google Veo 3.1
Higgsfield
seedance
ByteDance Seedance 2.0
Higgsfield
kling3
Kling 3
Higgsfield
wan2-7
Wan 2.7
Higgsfield
image2video
Higgsfield native
Runway
veo3
Google Veo 3.1
Runway
gen4_turbo
Runway Gen 4.5
Runway
seedance
ByteDance Seedance 2.0
Dali scores for the underlying model's native prompt language, not the platform wrapper. Pass the model name (veo3, kling, seedance…), not the platform name.
Why model-specific?
Generic prompt optimizers don't know that:
Veo 3.1 needs camera movement specified above everything else
Kling 3 supports multi-shot scene labels natively in the prompt
Flux responds to camera body and lens names like a photographer ("Sony A7 IV, 85mm f/1.4")
Midjourney V8.1 reads prose + parameters, not keyword lists
Higgsfield simulates physics — you describe materials in motion, not motion abstractly
Minimax uses [Pan left] bracket syntax for camera moves — plain text camera commands are ignored
Ideogram V4 needs text quoted exactly in the prompt for typography accuracy
Wan 2.7 generates native audio — include sound descriptions alongside visuals
Dali has a separate scoring rubric and rewrite brief for each model. Your LLM does the creative rewriting — Dali provides the intelligence.
Model guides live in dali/data/guides/{model}.json on the hosted server. Found practitioner patterns that consistently produce high-grade results? Open an issue with the model, the pattern, and a sample prompt + result. The best contributions come from Reddit, Discord, and YouTube — real practitioners, not official docs.