Multi-LLM dev harness, MCP-operable: bugs, cycles, gates. Verdicts are exit codes, never opinions.
Model Context Protocol (MCP) Server: io.github.jrullan/ducklab
io.github.jrullan/ducklab is a multi-LLM, self-hosted development harness that is MCP-operable. It runs a project's development cycle in fixed roles, using test gates where verdicts are exit codes rather than opinions, and supports local models alongside OpenAI-compatible or Anthropic endpoints.
๐ ๏ธ Key Features
Multi-LLM dev harness with fixed roles
Development cycle: brief โ requirements โ spec โ plan โ build โ review โ release
Test gates with verdicts as exit codes (not model opinions)
Local-first model support (llama.cpp, vLLM) plus OpenAI-compatible/Anthropic endpoints
Human or MCP/agent operation with recorded, attributed decisions
๐ Use Cases
Self-hosted, Linux-first development workflow orchestration
Multi-agent code generation and code review with gated outcomes
Reproducible runs using recorded data in .ducklab/
โก Developer Benefits
MCP operability for other agents and tooling integration
Clear phase-based pipeline for full development-cycle automation
Exit-code verdicts for consistent gating
โ ๏ธ Limitations
Linux first (per the excerpt); no other OS support details provided
Tooling is described as having a harness/engine, but no specific tool list or toolCount data provided
A self-hosted harness that runs a project's full development cycle with
several LLMs in fixed roles, under test gates that only you sign.
In one block: self-hosted development harness (Go engine + CLI + desktop,
Linux first) ยท brief โ requirements โ spec โ plan โ build โ review โ release ยท
verdicts are exit codes, never model opinions ยท local models first (llama.cpp,
vLLM) beside any OpenAI-compatible or Anthropic endpoint ยท operable by humans
or by other agents over MCP with recorded, attributed decisions ยท
Apache-2.0 ยท develops itself (the run records in .ducklab/ are the
receipts). Agents: start at AGENTS.md and llms.txt.
You give it a brief. It writes requirements, a spec and a plan; builds tasks
with one model or several arguing; runs your project's real test gate; and
stops for you before anything is committed. Every model call is logged. No
model ever decides a verdict.
A live council intake: the architect streams a requirements draft, the reviewer approves, and the run stops at a human gate A real council intake, recorded live and sped up: the architect streams the draft, a different model reviews it, the budget ticks in cents โ and the run stops at your gate. Total cost of what you just watched: $0.07.
It was built for local models first. Two of the seats that built most of it
are a vLLM box on the LAN and a llama.cpp server on localhost, both priced
at zero; hosted models sit beside them in the same roster, measured by the
same evidence.
Why this exists
Most agentic coding tools assume one strong model and trust it. Ducklab
assumes several cheap models and trusts none of them:
The gate decides, never a model. A verdict is a command's exit code.
A test-first run measures a green baseline before any test is written,
the red over the new test after, and every accept reproduces the gate
from a clean checkout of the committed sha โ nothing lands that did not
reproduce, and an accept whose reproduction fails takes its own commit back.
Decorrelation everywhere. A different model reviews; a reviewer never
learns who wrote the code (absent from the payload, not hidden in the UI);
tournament judges choose blind; council critics read the draft, not each
other.
Work is a contract. A task's deliverables are the implementer's
numbered checklist; it reports on each by number, the reviewer checks each
against the diff, and an undelivered item summons the rubber duck โ an
advisor seat that wakes only on measured distress (brake refusals, failure
streaks, red gates) and answers none, a note that sends the implementer
straight back to work, or stop.
Seats are chosen on evidence. Every duckling carries a scorecard โ
in-seat pass rate from your own runs, cost per run, coding index โ and the
roster board suggests seats from it, with the ranking criteria yours to
reorder. Suggestions are rare and justified: pass rates rank by their
Wilson lower bound, three runs minimum, locals never win on a $0 price.
Nothing is unbounded. Turns, tokens, cost, wallclock, tool output,
shell commands โ every ceiling visible and liftable mid-run, on the record.
Your documentation is not bounded by the model's window. Attach a wiki
to a stage and a big seat reads it whole; a small seat gets each document
digested to fit, the full text one ref_read call away, and the gate
names any document nobody opened. A 32k local model can be briefed by a
quarter-million characters of reference material โ the harness carries the
working memory.
The record does not round up: every run with its verdict, its cost, and whether its accept reproduced green from a clean checkout.
Ducklab is developed inside ducklab. The plan, the bugs, the releases and
the accepted tasks went through its own loop, driven by the same local and
hosted models it measures; recent features (per-run worktrees, the
merge-proof accept, the acceptance receipts, the governance write guard)
were built by the duck and gated by a person. To check the claim yourself:
bash
git clone https://github.com/jrullan/ducklab && cd ducklab
go build -o ducklab-cli ./cmd/ducklab
for r in .ducklab/runs/*/receipt.json; do ./ducklab-cli proof verify "$r"; done
Receipts ship with every accept since v0.7.0: the committed sha, the gate
command, its exit code, and the clean-checkout reproduction verdict โ
facts a third party re-derives, never assessments.
Status
v0.7.0 plus the phase-3 work now on main: every build and test run
executes in its own git worktree (your checkout is never touched),
acceptance rebases the run branch, re-runs the gate on the rebased commit
and merges fast-forward only, and an operator can re-close a finished run
as landed when its work reached main outside the engine. Before that:
seven stages, five modes, the roster board with measured scorecards,
reference documents with automatic digestion, skills managed from the
desktop, a seated consultant chat (vision verified before images are
sent), bug reports with screenshot evidence, adopt surveys with a
deterministic coverage check, provider-aware queueing that states why a
run waits, escalation suggestions when a seat measurably hits its
ceiling, acceptance receipts (ducklab proof verify), releases,
autopilot, a CLI, a desktop app, and an MCP server โ in the
official MCP registry as
io.github.jrullan/ducklab โ so another model can operate the loop with
recorded, attributed decisions.
docs/status.md tracks all acceptance criteria and does
not round up. Where code and spec differ, the difference is recorded in
docs/decisions/.
Install
Needs Go 1.25+, Node 22+ for the desktop, and git.
Linux
The CLI and engine are pure Go. The desktop is a Wails v3 app and needs the
GTK/WebKit development packages:
bash
sudo apt install libgtk-3-dev libwebkit2gtk-4.1-dev # Debian/Ubuntu names
make desktop && make install
On Ubuntu 24.04+ the desktop also needs an AppArmor profile โ see
decision 0003 and
packaging/apparmor/.
macOS
bash
xcode-select --install # the desktop build links against WebKit
brew install go node
make desktop && make install
Honesty note: ducklab is developed and exercised daily on Linux. The CLI and
engine compile-check for darwin/arm64 on every make cross, but no desktop
build has been verified on a Mac yet โ the first person to try it is the
test, and make install gives you the CLI and engine either way. Please
report whatever breaks.
Both
make install installs to ~/.local/bin โ make sure it is on your PATH.
It warns when the desktop binary predates frontend/src, because it will
happily install a stale one.
Frontend development without the desktop
To exercise the frontend in a browser, run the engine and Vite in separate
terminals, then open the browser with its connection details. The fake engine is
the quickest option; the same flow can use a real engine with its opt-in CORS
flag:
bash
# Fast, scripted data (recommended for UI work)
go run ./cmd/fake-engine --port 8787 --token fake-token
npm run dev --prefix frontend
# open http://localhost:5173/?engine=http://127.0.0.1:8787&token=fake-token# Or use real engine data (development only; keep the origin explicit)
go run ./cmd/ducklab-engine --allow-origin http://localhost:5173
npm run dev --prefix frontend
# open http://localhost:5173/?engine=http://127.0.0.1:<engine-port>&token=<engine-token>
The real engine remains same-origin restricted by default. --allow-origin
enables exactly one browser origin and is intended for local frontend development
and visual audits; it does not change authentication or the loopback bind. Without
this flag, a browser's cross-origin failure can look like a dead session.
The engine and token query parameters are available only in Vite dev
builds. They can also be supplied as VITE_DUCKLAB_ENGINE and
VITE_DUCKLAB_TOKEN environment variables. The desktop shell continues to use
its injected window.ducklab connection.
Three binaries
What it is
ducklab-engine
The daemon. Owns every run. Binds 127.0.0.1 only, bearer token rotated each start.
ducklab
The CLI client. Holds no state; it asks the engine.
ducklab-desktop
The desktop app. Also a client, also holds no state. Starts (or adopts) the engine itself.
Provider keys come from the engine's environment at call time โ export them
before it starts, or launch the desktop through a wrapper that loads them
from your keyring. The app tells you when the engine it adopted is missing a
key this app has, with the restart button beside the words.
A cycle, end to end
From the desktop: Projects โ New project, then Cycle โ Draft it. From
a terminal:
bash
cd ~/dev/myproject
git init # ducklab needs a git repo
ducklab project init --name MyProject # auto-starts the engine if none is running
ducklab intake --from brief.txt # brief โ requirements
ducklab intent # your briefs, verbatim, and what each one changed
ducklab spec # requirements โ spec
ducklab plan # spec โ milestones and tasks
ducklab run T-001 # build it
ducklab run accept r-20260729-... # commit it
ducklab review T-001 # read the commit
ducklab release plan --bump minor # what shipped
Your words are part of the record: every brief is kept verbatim as an
INT-nnn entry before any model reads it, and the requirements it added or
changed point back to it โ so a requirement can always answer "who asked for
this, and in what words".
Each stage writes a .proposed file first and waits for you. accept
promotes it; reject restores exactly what the run wrote and nothing else;
"request changes" sends any draft โ spec, plan, release notes โ back with
your note. Nothing is committed without you (or without the autonomy level
you explicitly granted).
Reference documents ride any stage: --ref ~/wiki/product/ (or the
attach door in the desktop) loads files or whole directories as background
for the architect โ grounded by two rules the prompt states outright: the
approved requirements own the scope, and where a reference and the code
disagree, the code is the truth. When the corpus outgrows the seat's
context, each document is digested once (cached by content hash), the full
text stays reachable through the ref_read tool, and the proposal card
lists any document no seat ever opened.
Adopting an existing codebase works the same way: intake reads the code
and writes as-built requirements, the spec marks its sections as-built, and
the plan stays deliberately empty โ new work then enters through bug reports
and plan amendments, which is how ducklab itself is developed.
Your project declares its own truth in .ducklab/project.toml: the gate
([verify] โ with link_deps and setup for what a clean checkout needs),
how the app launches ([run] with a preflight), and how the project's own
binaries are rebuilt ([install]) so the whole loop runs without leaving
ducklab. Stack capabilities propose run.command choices from build metadata
(for example Meson executables, Cargo bins, Go main packages, or Node scripts),
but a person must adopt one. Once configured, every otherwise-green build gate
executes run.smoke when declared, otherwise it falls back to run.command.
This lets an interactive application declare a headless smoke without opening
its real UI during every gate. run.smoke_expect = "exit" | "live" defines
success: exit requires status 0 inside run.smoke_timeout_s; live requires
the process to survive the whole observation window. When omitted, an explicit
run.smoke defaults to exit, while a run.command fallback or an app with a
URL/health endpoint defaults to live. The gate and UI always state the
effective expectation, so an interactive program that exits cleanly in 20 ms
cannot masquerade as a booted product.
Gate and shell process trees always receive DUCKLAB_RUN_ID and DUCKLAB_PROJECT_ID. For example, excercise-tracker can use DATABASE_URL=test_db_${DUCKLAB_RUN_ID} in [verify].tests, and a compose preflight can use ${DUCKLAB_PROJECT_ID} as its per-run project name. Ducklab guarantees identity only; provisioning and teardown remain the project's.
--key-env is the name of an environment variable, never a key. No key
is written to config, sent over the API, or kept in shell history.
Seat suggestions come with their evidence: pass rates from your own runs, cost per run, coding index. You decide.
The desktop's Roster view is where seats are assigned: drag from the
Flock onto a mode's seat, globally or per project, with each duckling's
evidence on the card and the engine's suggestions beside the seats. Coding /
intelligence / agentic indices come from OpenRouter's benchmarks endpoint
when a duckling lives there; your own runs supply the rest.
The fleet that built this repo
This is not a recommendation list. It is this repository's own run record
(454 recorded runs, ~2,300 seat assignments as of 2026-08-24), so you can
see what actually held which seat. Any OpenAI-compatible endpoint slots in
the same way.
Duckling
Model
Served by
Seats held
What the record says
beelink-local
Qwen3.6-35B-A3B (Q4 GGUF)
llama.cpp (Vulkan) on a Ryzen AI Max 395, on-desk
465 (the most-seated duckling in this repo)
judge, scribe, reviewer. Free.
luna
gpt-5.6-luna
OpenRouter
455
implementer workhorse: 77% measured pass rate at ~$0.02/run.
atom-local
Qwen3.8-27B
vLLM on a DGX Spark on the LAN
352
architect and scribe; it wrote the release notes. Free.
k3
Kimi K3
OpenRouter
348
triage, architecture drafts, question advisor.
terra
gpt-5.6-terra
OpenRouter
303
the heavier implementer, ~$0.28/run.
glm52
GLM-5.2
OpenRouter
160
the reviewer seat: 81% measured over 261 reviews.
qwen38-max
Qwen3.8-Max
OpenRouter
140
the advisor (the rubber duck). 88% measured.
pato-sonnet
Claude Sonnet 4.5
OpenRouter
7
the expensive seat, used when cheaper ones measurably hit a ceiling.
Two notes for accuracy. First, "built with local models" here means the
local seats held judgment and documentation roles (judge, reviewer,
scribe, architect) while cheap hosted models did most of the typing;
about a third of all seat assignments ran on hardware in this room.
Second, the pass rates above are measured on my runs (ducklab duckling scorecard, Wilson lower bound). Yours will differ, and that is the
point: the roster works from your record, not from a leaderboard.
The same machinery on real work: a council revising ducklab's own spec, 4.5M tokens in, paused once on a budget it asked to lift.
The five modes
ducklab run T-001 --mode <mode>
Mode
What it does
solo
One duckling. The yardstick everything else is measured against.
pair
Implementer and reviewer, decorrelated. Between them the advisor โ the rubber duck.
tournament
Contestants build the same task in isolated worktrees; a judge picks, blind.
split
An architect decomposes; subtasks run in parallel; integration is file copies, no model involved.
council
Several models on one document, for intake, spec, plan and review. One drafts, the others critique blind, the first revises.
What it will not do
Invariants, enforced in code:
A model never decides a verdict. A gate is a command's exit code.
A green candidate is applied byte-for-byte; nothing is re-generated
after it passed.
A reviewer never learns who wrote the code.
Nothing lands that did not reproduce from a clean checkout of the
commit being merged.
A reject undoes what the run wrote, and nobody else's work.
Every budget (tokens, cost, turns, wallclock) has a ceiling you can see.
Secrets never touch project state.
The engine is loopback-only. There is no remote mode.
Skills
A skill is a directory with a SKILL.md โ under .ducklab/skills/ for one
project, or in the machine-wide skills directory to serve every project
(project shadows global on a name collision). The documentation-only form
has no script and is the default: a recipe a model reads and follows. The
architect reads survey guides before an adopt (skill_list is in its
prompt), the consultant reads them in chat, and only the implementer can
skill_run an executable one.
Skills are administered from the desktop (gear โ Skills): list with
scope badges and validation problems, read, edit the whole SKILL.md,
run with arguments, delete. A skill a duckling writes during a run shows
there greyed pending acceptance until its run is accepted โ proposing a
skill goes through the same gate as proposing code.
bash
ducklab skill new house-style
ducklab skill run changelog-entry --arg summary="..."
The consultant
Every project seats a consultant (a Common seat on the roster board):
the model behind the "chat about this" doors and the free-form chat in the
guide rail. It reads the code, the runs, the boards and the skills โ never
writes โ and takes images: paste a screenshot of a broken view and ask.
Vision is verified, not assumed: a declared-vision seat is probed with a
real image request once, and a text-only seat refuses the paste with words
instead of hallucinating an answer.
The seated consultant answering a question about the repo it reads.
Operating ducklab from another model
ducklab mcp serve exposes the whole loop over stdio as an MCP server: an
external model reads each result, decides gates (with a required, recorded
reason โ decisions land as approved_by: mcp:<client>, never as "human"),
answers questions, files bugs, amends plans and starts work. The engine's
next lists are the law: an operator cannot take an action a person could
not.
Contributing
See CONTRIBUTING.md โ how to build, how the tests guard
the architecture, how work flows through ducklab's own loop, and where to
start. The short version:
bash
make # vet, test, build the frontend
go test ./... # 39 packagescd frontend && npx vitest run
License:Apache-2.0. Contributions are accepted under the
same terms (ยง5 of the license โ no CLA). The Ducklab name and the duck are
the maintainer's (ยง6).
Specification
The code implements a written specification, in this repo:
docs/spec/ (00-VISION through 08-DESKTOP-UI) is the
normative layer โ vision, invariants, protocol contracts, acceptance
criteria. What the system IS today lives in .ducklab/docs/ โ the as-built
requirements, spec and plan the loop itself maintains, each version signed
at a human gate. Where the two differ deliberately, the difference is
recorded in docs/decisions/; the diff between them is
the roadmap, and the alignment stage computes it.