Deterministic security scan of MCP servers, agent skills and npm/PyPI packages. Runs locally.
ai.skilltotal/skilltotal (MCP)
This MCP server performs a deterministic security scan of MCP servers, agent skills, and npm/PyPI packages. It is designed to run locally, producing static-analysis results relevant to LLM security and supply-chain security workflows, with support for reporting artifacts such as SARIF.
π οΈ Key Features
Deterministic security scanning
Covers MCP servers, agent skills, and npm/PyPI packages
Focuses on static-analysis and llm/prompt-injection-related security
Supports SARIF as an output format
π Use Cases
DevSecOps checks for model-context-protocol (MCP) server security
Auditing agent skills for security issues
Assessing npm/PyPI packages for supply-chain security concerns
β‘ Developer Benefits
Local execution for reproducible scanning
Static-analysis oriented findings for security review
Fits into security-tools pipelines using SARIF
β οΈ Limitations
No tool count or specific MCP tools/transport details are provided in the available data.
AI Component Security Platform β open-source CLI engine.
SkillTotal statically analyzes AI-related components β agent skills/plugins, MCP servers, npm /
Python packages, repositories, and AI-generated projects you upload as an archive or file β to
surface supply-chain risks, dangerous capabilities, prompt-injection surfaces, and data-exfiltration
paths before the component is installed or trusted. Point it at a path, a git URL, an
npm: / pypi: package, or a project archive (.zip / .tar.gz) / single file.
Try it online (no install, no account):www.skilltotal.ai β
the website runs this same engine. Prefer the CLI? pipx install skilltotal (below).
It analyzes only the component itself β never your user, company, environment,
deployment, or runtime context. Every score and finding is derived exclusively from the
files inside the component.
Core principle: every confirmed finding carries evidence (file, line range, code
snippet). Anything that cannot be evidenced is placed in needs_review, never in
findings, and never affects the score.
Why SkillTotal
100% local & offline β the component's code never leaves your machine. No account,
no API token, no cloud upload (unlike cloud scanners that send your components to a backend).
Safe to point at untrusted components β the engine analyzes without ever running them on
your machine. (Optional dynamic analysis is a separate paid service that runs only in our
isolated sandbox, with your consent.)
Zero runtime dependencies, pure Python stdlib β auditable and easy to vendor/air-gap.
Deterministic β regex + AST, no LLM in the static engine; the same input always yields
the same report.
Evidence-anchored & low false-positive β every finding points at an exact file:line.
Standards-aligned β every component gets a behavioral trait fingerprint mapped to the
Cloud Security Alliance (CSA) agentic threat model, MAESTRO threat-model layers, and
MITRE ATLAS tactics β including a three-way execution-context read (embedded static
credential β delegated OAuth/OIDC β least-privilege scoped identity) that shows the blast radius
of a compromise, not just that a secret exists.
Free and open source (Apache-2.0) β the full static report is free, forever.
Measured, not asserted
Detection claims are cheap, so the numbers behind them are published with the data and the code
that produced them.
The whole MCP registry, scanned β every distinct component in
the official registry, 17,535 of them, in one deterministic run.
88.2% expose tools to an agent, 64.8% can reach the network, 21.1% can execute shell commands β
and the risk distribution underneath is far flatter, because a capability scores zero here.
Raw JSON Β· the harness.
Detection efficacy β recall and precision on a labelled corpus,
regenerated every release and enforced by CI as a floor.
Every one of these reproduces: same input, same engine, same output. Nothing is executed and no
LLM is involved.
Install
Requires Python 3.10+. Zero runtime dependencies. git is required only for scanning
remote URLs.
Recommended for the CLI β pipx (isolated install; also works on
Debian/Ubuntu where bare pip install is blocked by PEP 668):
bash
pipx install skilltotal
Or into a virtual environment / as a library:
bash
pip install skilltotal
From source (development):
bash
pip install -e ".[dev]"
Usage
bash
# Human-readable report
skilltotal scan ./path/to/component
# Scan a remote repository (shallow git clone)
skilltotal scan https://github.com/owner/repo
# Scan a project archive or a single file (e.g. an AI-generated project downloaded as a ZIP)
skilltotal scan ./my-project.zip
skilltotal scan ./app.tar.gz
skilltotal scan ./suspicious.py
# Scan a package from a registry (latest, or a pinned version)
skilltotal scan npm:left-pad
skilltotal scan npm:left-pad@1.3.0
skilltotal scan pypi:requests
skilltotal scan pypi:requests==2.31.0
# JSON to stdout
skilltotal scan ./component --json
# SARIF 2.1.0 (GitHub Code Scanning / IDE)
skilltotal scan ./component --sarif --output report.sarif
# Write the report to a file (SARIF if --sarif, else JSON)
skilltotal scan ./component --output report.json
# CI gate: exit code 2 by severity level or by risk score
skilltotal scan ./component --fail-on-high # alias for --fail-on high
skilltotal scan ./component --fail-on medium
skilltotal scan ./component --fail-on-score 50
# Skip paths (repeatable; combined with the config file's `exclude`)
skilltotal scan ./component --exclude "vendor/*" --exclude "*.min.js"# Opt-in provenance for npm:/pypi: sources (registry metadata -> needs_review, never scored)
skilltotal scan npm:some-lib --provenance
# Baseline: snapshot current findings, then suppress them on later scans
skilltotal scan ./component --write-baseline .skilltotal-baseline.json
skilltotal scan ./component --baseline .skilltotal-baseline.json --fail-on-high
# Diff two versions of a component: what changed between them?# Each side is any scannable source (path/archive/git/npm:/pypi:) or a saved --json report.
skilltotal diff npm:some-lib@1.2.3 npm:some-lib@1.2.4
skilltotal diff ./old-checkout ./new-checkout --json
skilltotal diff old-report.json new-report.json
# CI gate: fail (exit 2) if the new version INTRODUCES a high/critical finding
skilltotal diff npm:some-lib@1.2.3 npm:some-lib@1.2.4 --fail-on-new high
# Pre-install guard: allow/block decision (exit 2 on block) you can chain before installing
skilltotal guard npm:some-mcp-server && claude mcp add some-mcp-server -- npx some-mcp-server
skilltotal guard --installed # check every AI component already on this machine
skilltotal guard npm:x --block-on malicious # block only on malicious indicators# Inventory: discover AI components already installed on this machine and scan them# (reads agent configs for Claude Desktop/Code, Cursor, Windsurf, VS Code, Gemini, and# local skills; derives an npm:/pypi:/local source per MCP server and runs the engine)
skilltotal inventory
skilltotal inventory --json
skilltotal inventory --no-scan # list only, do not scan
skilltotal inventory --project . # also include this project's agent configs
skilltotal inventory --sbom # AI-BOM: CycloneDX 1.6 JSON of your agent stack,# scan verdicts attached as component properties# List every detection rule
skilltotal rules list
skilltotal rules list --json
Baseline suppresses findings by a stable fingerprint of
(rule id, file, code snippet) β independent of line numbers, so it survives edits.
Suppressed findings are removed before scoring and do not affect the risk score.
Diff reports new / resolved / changed findings, evidence-level additions and removals
(matched by the same line-independent fingerprint as the baseline, so pure line shifts are
not noise), capability changes, and the risk-score delta. --fail-on-new LEVEL gates only
on risk the new version introduces β existing accepted findings never trip it, so it fits
upgrade reviews ("is 1.2.4 riskier than the 1.2.3 we already vetted?") without a baseline
file.
Guard is the install-time answer to "should I trust this component right now?".
Malicious indicators always block; scored risk at/above --block-on blocks;
capabilities alone never block β a legitimate MCP server with shell/network access
passes, so the guard stays quiet enough to leave enabled everywhere (unlike a raw
--fail-on high gate, which would trip on most of the ecosystem's honest capability
findings).
Provenance (--provenance, opt-in) adds registry-metadata signals for npm: /
pypi: sources: recently published, deprecated / yanked, no recent releases, no
repository link. Metadata is context about a component, not component content β so these
signals go to needs_review and never affect the score or verdict, and the default
scan stays 100% component-only and offline.
Project config (optional) β commit a .skilltotal.toml instead of repeating flags
(CLI flags override it):
toml
fail_on = "high"# low | medium | high | criticalfail_on_score = 50# or gate on the 0-100 risk scoreexclude = ["vendor/*", "*.min.js"]
ignore = ["ST-NET-PY"] # rule ids to dropbaseline = ".skilltotal-baseline.json"# Per-rule policy: reviewable gate decisions that live in the repo, not in a dashboard.[policy]"ST-SHELL-PIPE-EXEC" = "block"# gate trips (exit 2) whenever this rule fires,# even with no fail_on configured"ST-DYN-PY" = "warn"# explicit accept-but-show: reported, still counts toward# the risk score, but exempt from the fail_on severity gate"ST-SENS-WORD" = "ignore"# suppressed entirely (same effect as `ignore`)
Suppress a single finding inline with a # skilltotal:ignore (or # skilltotal:ignore[ST-ID])
comment on its line.
python -m skilltotal ... works identically to the skilltotal console script.
A configured gate tripped (--fail-on/--fail-on-high severity, --fail-on-score, or diff --fail-on-new)
Gate semantics:--fail-on/--fail-on-high trip on the severity of any single finding,
not the aggregate risk_score. A component can report risk_level: low (score 0) and still fail
the gate if it has a high-severity finding β including a powerful capability (e.g. shell or
network access), which is reported but never scored as malicious. To gate on the score instead,
use --fail-on-score; to accept known findings, use a baseline, an inline
# skilltotal:ignore[ST-ID], or a per-rule [policy] action (block / warn / ignore).
CI / GitHub Action
Run SkillTotal in CI and surface findings in your repository's Security β Code scanning tab.
yaml
# .github/workflows/skilltotal.ymlname:SkillTotalon: [push, pull_request]
permissions:contents:readsecurity-events:write# required to upload SARIF to Code Scanningpull-requests:write# required only for comment-on-pr (optional)jobs:scan:runs-on:ubuntu-lateststeps:-uses:actions/checkout@v4-uses:pezhik/skilltotal@v0.53.0with:source:.# a path, a git URL, or an npm:/pypi:<name> specfail-on:high# fail the build on a high/critical finding (or 'none')comment-on-pr:'true'# post a sticky summary comment on pull requests (optional)
The action installs the CLI, scans source, uploads SARIF (so findings appear inline on pull
requests and in Code Scanning), and fails the job on a high/critical finding unless
fail-on: none. On pull requests, comment-on-pr: 'true' posts a single summary comment (risk
level, score, findings, capabilities) and updates it in place on later runs β it needs
pull-requests: write and is off by default. Pin the action to a released tag (see
Releases) and, optionally, pin the engine version
with the version: input. Prefer plain CLI? It is the same thing:
skilltotal scan . --sarif --output skilltotal.sarif --fail-on-high.
# .pre-commit-config.yamlrepos:-repo:https://github.com/pezhik/skilltotalrev:v0.53.0hooks:-id:skilltotalargs: [".", "--fail-on-high"] # scan the repo; block the commit on a high/critical finding
Then pre-commit install. The hook installs the CLI in its own environment and scans the repo
on commit; tune the scan with the same flags as the CLI (e.g. --exclude, --fail-on).
Use as an MCP server
Let your agent check a component before installing it. skilltotal mcp runs the engine
as a stdio MCP server (stdlib-only, still zero dependencies) β register it in Claude
Code/Desktop, Cursor, Windsurf, or any MCP client:
Tools exposed: scan_component (full report for a path / git URL / npm: / pypi:
source), diff_components (upgrade review: what changed between two versions), and
list_rules. Scans run locally with the same never-execute static engine β the component's
code is not uploaded anywhere.
Add a status badge
Scan a component on skilltotal.ai and each report offers an
"Add this badge" snippet β a small SVG that always reflects the component's latest scan and
links back to the full report. Drop it in your README so visitors see the risk at a glance:
Copy the exact, ready-to-paste markdown from the report page β it fills in the badge URL for you.
Methodology
SkillTotal performs static security analysis of AI components β MCP servers, agent
skills/plugins, npm and PyPI packages, and AI-generated projects/repositories. The engine
combines capability analysis, dangerous-pattern detection, privilege analysis, supply-chain
(install-time) analysis, prompt-surface analysis, and data-flow correlation (e.g. secret
access combined with network egress). Findings are mapped to risk categories and contribute to
a 0β100 risk score; capabilities are reported but never inflate the score β capability β risk.
Nothing is executed and no LLM is called, so results are deterministic and reproducible.
manifests, dangerous tools (shell/fs/network/credential), server commands
Prompt surface
"ignore previous instructions", "reveal system prompt", exfiltration phrasing
Coverage by component type
Legend: β analyzed by default for this component type Β· β οΈ the engine detects this, but
that surface is uncommon for this type β so it is flagged only when the component actually contains
it (e.g. prompt-injection text inside an npm/PyPI package) Β· β not applicable to this type Β·
π§ planned (SkillTotal Cloud).
Columns are the component types SkillTotal scans. AI project = a scanned repository or folder
β an agent skill/plugin, an AI-generated codebase, or a set of prompts/configs β that is not a
published npm/PyPI package.
Category
MCP
npm
PyPI
AI project
Prompt injection / instruction override
β
β οΈ
β οΈ
β
Tool poisoning (MCP tool metadata)
β
β
β
β οΈ
Dangerous capabilities (shell / fs / network)
β
β
β
β οΈ
Data exfiltration (secret access + egress)
β
β
β
β οΈ
Secret theft / sensitive-path access
β
β
β
β οΈ
Dynamic code execution
β
β
β
β οΈ
Obfuscation (decode-and-execute)
β
β
β
β
Hidden-Unicode smuggling
β
β
β
β
Embedded secrets (hardcoded keys/tokens)
β
β
β
β
Install-time / supply-chain hooks
β οΈ
β
β
β
Overprivileged / auto-approved tools
β
β
β
β οΈ
Runtime behavior analysis
π§
π§
π§
π§
Sandbox analysis
π§
π§
π§
π§
Typical findings
An MCP tool can execute arbitrary shell commands
A package downloads and runs code from an external URL
Access to credential locations (~/.aws, ~/.ssh, .env) detected
Dynamic code execution (eval / exec) detected
Prompt-injection / instruction-override phrasing in a tool description or skill
Sensitive-data access combined with outbound network egress
Hardcoded API keys or tokens
An MCP server with auto-approved or overprivileged tools
Untrusted input (environment, sys.argv, a request/response body) flowing into exec or a
shell β a proven injection path, not just a dangerous API in isolation
An agent skill does more than its declared allowed-tools allow (undeclared capability /
least-privilege violation)
Out of scope
SkillTotal statically analyzes a single component's own files. It does not execute code,
observe runtime behavior, or assess your environment, deployment, or infrastructure. It is not
a substitute for:
a penetration test
an application-security (app-sec) review
an architecture / design review
a cloud-security or infrastructure assessment
a Kubernetes / container runtime audit
a business-logic review
a manual code review
Runtime behavior and sandbox analysis are planned for SkillTotal Cloud (paid).
Output
A normalized report containing the component identity, a risk score (0β100) and
risk level (low / medium / high / critical), detected capabilities (each
evidence-backed), a behavioral trait fingerprint (with a CSA / MAESTRO / MITRE ATLAS
crosswalk), findings, needs_review, and metadata. See
docs/report-schema.md and docs/scoring.md.
Every finding also carries its OWASP Agentic Skills Top 10 category ids (owasp), emitted in
both the JSON report and SARIF (native taxonomies/relationships);
docs/owasp-agentic-skills-mapping.md explains the coverage
(AST01βAST05) and the honest gaps. For MCP servers,
docs/mcp-owasp-mapping.md maps SkillTotal's checks to the OWASP MCP
Security Cheat Sheet (and names the runtime controls a static engine can't cover).
The report's traits array is a behavioral fingerprint β a higher-level projection over the
findings (e.g. execution_authority, embedded_credential, untrusted_perception, and the
emergentexfil_correlation combination) β each mapped to the Cloud Security Alliance
trait-based model, a MAESTRO threat-model layer, and a MITRE ATLAS tactic where there is an honest
fit. It is descriptive and never affects the score; see
docs/trait-crosswalk.md.
Architecture
The package under skilltotal/ (except cli.py) is a pure, side-effect-free library so the
same engine can power the future web app and enterprise SaaS. See
docs/architecture.md.
Development
bash
pip install -e ".[dev]"
pytest
Accuracy notes
Python is analyzed via an AST (resolves import aliases, tells open(p,'w') from a
read, ignores API names that only appear in strings/comments). Node.js/config use regex.
Test code (__tests__/, *.test.*, tests/, conftest.py, β¦) is demoted to
needs_review β it is not executed by consumers, so it does not affect the score.
Ambiguous signals (bare secrets/credentials words, lone base64 blobs, "before
answering" phrasing, minified files) go to needs_review, never to findings.
Hidden Unicode (ASCII-smuggling tag characters, Trojan-Source bidi overrides,
zero-width chars) is detected and decoded β a real evasion used to smuggle instructions
past human review. See tests/manual_eval/ for calibration against real-world attacks.
Shell execution covers subprocess/os.system, asyncio.create_subprocess_*, Node
child_process, and common process-spawning libraries (Python sh/plumbum/pexpect/
invoke/fabric; Node zx/execa/cross-spawn/shelljs/tinyexec/node-pty).
MCP dangerous tools are classified by name/description both in JSON manifests and when
defined in code (server.tool("run_command", β¦), @mcp.tool over def read_file).
Limitations: detection is at the call/import level. Capability via an unrecognized
higher-level library (e.g. a git library that writes files internally, a browser library)
may not be flagged as a raw filesystem/shell call. Capabilities indicate presence, not
proven misuse.
Open source vs SkillTotal Cloud
SkillTotal is open core. This engine (analysis + all detection rules + CLI) is open source
and complete on its own β run it locally or in CI, free, offline, with zero runtime
dependencies. It tells you what a component does, with evidence.
Paid features are delivered only via SkillTotal Cloud (the website) and explain why it
matters: LLM interpretation and prioritization of findings, dynamic sandbox execution,
hosting, scan history, and monitoring. They are server-side services on top of this engine β
their code is not part of this repository. See docs/open-core.md.