assurance

Software that reports success does not thereby prove it
Every tool here answers one question about work that has already happened, and refuses to answer it
when it cannot. Nothing consults a model. Every result is arithmetic you can recompute yourself.
Start with the sentence that shows what that means in practice:
$ pip install assurance-budget
$ assurance-budget runs.jsonl
0 of 3 runs hit a limit β 1 was going nowhere first
r-002
Stopped: 3 rounds repeating fetch(url=api/invoices) and failing the same way
(timeout) with nothing new read and no part of the goal closer. Continuing would
spend the rest of this run's budget on the same result.
Not tested by this log: iterations, retries, seconds. The log carries no events of
that kind, so this is silence rather than a pass.
Read the last line again. The run passed three of the four limits β and the tool says so is not
the same as says nothing. Most software reports the absence of a failure as a success, which is how
a check that never ran becomes a green tick. This one names what it could not test, in the same
breath as what it could.
That is the whole idea, and it is why these are separate from any product: a claim you can check
is worth more than a claim you have to trust.
Three questions, three commands
Each installs on its own. None needs the others, an account, a service, or a network.
Did the work cover what it was supposed to cover?
pip install assurance-cli
assurance check ~/reports
No config and no corpus file β it reads the cadence, the span and what is absent from the filenames.
A folder with no regular cadence is told so rather than handed a ratio.
Where did the run's budget go, and where did it go nowhere?
pip install assurance-budget
assurance-budget runs.jsonl --fail-on-exhausted
The expensive runs are rarely the ones that crash. They are the ones that retried the same failing
call fourteen times and finished with a plausible answer and a bill. Ceilings are enforced by code
the caller cannot talk out of them.
May this task proceed, for the person who asked?
pip install assurance-authority
assurance-authority team.json
1 of 3 tasks may proceed for the person who asked β 1 moved owner β 1 refused
team roster intern-42 proceed
Q3 margin memo intern-42 escalate_ownership -> CFO
Priya (intern) may not receive finance-confidential, and CFO may. The task moves to
CFO rather than the answer moving to Priya (intern).
payroll extract agent-a refuse
Drafting agent may not receive payroll, and nobody offered can. The task stops here.
The middle row is the product. The intern may not have the margin memo; the CFO may. So the task
moves to the CFO β she is told it moved, and never told the figure. An agent fetching it as a service
account and handing her the answer is a permission-laundering machine with your company's name on it.
The rule all three follow
A denominator we cannot establish is refused, never invented. A tool that answers "0 of 36" for a
folder it did not understand is worse than one that says it does not know, because you cannot argue
with a number that was made up.
All six packages
The three commands above are the way in. These are the parts they are made of, each installable on
its own and versioned on its own β a release tag names its package (cli-v0.5.10), because a bare
version number is ambiguous between six.
| package | what it is |
|---|
assurance-core | the decision layer as a pure library β no I/O, no model, no framework. Coverage, corpus census, staleness, drift, tool pinning, the rule of two |
assurance-cli | six commands, each a CI gate: check, diff, pin, drift, deps, init |
assurance-mcp | four MCP tools, read-only by construction, for Cursor / Claude Desktop / any MCP client |
assurance-deps | what a pip install or npm install is about to execute, read without executing it β and what could not be read |
assurance-budget | where a run spent, and where it went nowhere. Ceilings a caller cannot raise |
assurance-authority | whether a task may proceed for the person who asked, and what happens when it may not |
budget and authority had their own repositories until 2026-09-09. One package per repository
meant a reader had to find four front doors and work out how they related before anything happened,
which is the opposite of the point. Those repositories are private as of 2026-09-11, so their old
URLs no longer resolve β the history and the releases are here and on PyPI, and every PyPI name is
unchanged.
Two more worth knowing about once you are past the first command:
assurance pin --check
assurance drift runs.jsonl
drift reports no labels, no judge and no benchmark β it says whether a change is distinguishable
from noise, and refuses when there is not enough history to say. Its
README leads with the false-alarm rates of the textbook methods it
rejected, because that is the part worth checking.
Layout
packages/core/ assurance-core β generated; see below
packages/cli/ assurance-cli
packages/mcp/ assurance-mcp
packages/budget/ assurance-budget
packages/authority/ assurance-authority
skills/ agent skills that use the tools above
packages/core/ is generated and must not be hand-edited. It is scrubbed out of a private
upstream by a publisher that rewrites the whole tree, so an edit made here is destroyed on the next
run and never reaches anyone. Everything else in this repo is ordinary hand-written code, and pull
requests are welcome against it.
Honest limits
check opens .csv, .tsv and .xlsx only. Anything else in the folder is counted and
named, not silently skipped.
- The span is inferred from the earliest and latest filenames unless you pass
--from / --to,
which means a report missing from either end of the range cannot be detected. Pass the range
when you know it.
expected is never inferred in the MCP tools. A denominator nobody can argue with is not an
answer.
- No cross-document inference. It produced 21 false positives on a real corpus, so it is refused.
Contributing
CONTRIBUTING.md has the setup β it is the sequence that was actually run, and
the note about upgrading pip first is load-bearing on Python 3.10.
Issues labelled good first issue
are scoped so the hard part is already decided in the issue text.
"I ran this on my own folder and the answer looked wrong" is a first-class issue and needs no
fix attached. That is how most of what is fixed here was found β including a folder of twenty-eight
files that reported thirty-three absent months which had never existed.
Licence
Apache-2.0.