Open, sourced, confidence-scored intelligence on AI startups - a free PitchBook alternative.
OpenPitch MCP Server (io.github.Avierovich/openpitch)
This MCP server provides open, sourced, confidence-scored intelligence on AI startups, positioned as a free PitchBook alternative. Its dataset is intended to be MCP-native, fully-sourced, and updated daily, and it focuses on the AI companies venture capitalists care about.
Status: v0.1.3 β functional. The pipeline, reconciliation engine, MCP server, and
dashboard all work end-to-end. Coverage and source breadth keep growing via the daily run.
PitchBook and CB Insights cost $20k+/year β and for fast-moving AI startups, their data is often months stale, because human verification is slow. For a company growing 3Γ a year, a figure verified six months ago can be off by multiples.
Meanwhile, the real numbers are already public: founders state ARR on podcasts weeks before any database, funding hits SEC filings, hiring velocity reveals growth. They're just scattered, unstructured, and contradictory β exactly the problem an AI agent is built to solve.
OpenPitch's bet is latency, not coverage. For the AI companies that matter, a fresh, fully-sourced, confidence-scored number beats a verified-but-stale one. We don't claim certainty β we show you the receipts.
What you get
Ask your coding agent, get an answer with receipts:
code
> what's Sierra's valuation, with sources?
Sierra β AI agents for customer service (sierra.ai)
Valuation $15.4B [consensus Β· confidence 0.96] Β· as of 2026-05
β³ 10 public sources Β· Reuters Β· CNBC Β· The Information Β· qz.com
β³ $950M round closed May 2026 β led by Tiger Global and GV
(A real answer from the committed data β check it against the live dashboard.)
Every number carries its source, a confidence score, and a tracked history of how it changed.
Features
ποΈ Mines podcasts β founders leak metrics on podcasts before any database catches them. We transcribe and extract them.
π§Ύ Always sourced β every figure links to its origin (podcast timestamp, filing, article). No black-box numbers.
π Confidence-scored β built from source reliability, speaker authority, corroboration, and freshness (confidence decays as data ages).
π Reconciles conflicts β when sources disagree, you get a consensus range + a contradiction flag, not a silent guess.
π§ Learns which sources to trust β sources that prove right over time earn more weight.
π Version-tracked β the git history is the audit log. See exactly how a company's reported ARR evolved.
π‘ Composable β emits typed events other agents subscribe to (newsletters, press alerts, investor outbound).
π€ A2A-discoverable β ships an A2A agent card so agent ecosystems can find and describe it.
π§― Grounding β give your AI a sourced, confidence-scored fact base so it stops making up AI-company numbers.
β‘ 60-second install β no key, no signup; works in your agent in under a minute.
πΈ Genuinely free β runs entirely on free tiers. No cost to run, no cost to use.
Quickstart β use it in Claude Code / Codex
No API key. No signup. No cost. The data is already built and committed; the MCP server just reads it, and your agent does the reasoning.
Fastest β zero install (reads the committed data from the public repo, no clone):
bash
uvx openpitch-mcp
Or install the package:
bash
pip install openpitch # the MCP server (mcp is a core dependency)
openpitch-mcp # start the read-only server
Or run from a clone (for the pipeline / to rebuild data):
bash
git clone https://github.com/Avierovich/openpitch && cd openpitch
python -m venv .venv && source .venv/bin/activate
pip install -e ".[pipeline]"# core + pipeline LLM deps
openpitch seed # build the data/ database from the committed seed (offline, no key)
Then point your agent at the local server:
jsonc
// MCP config (Claude Code / Codex) β zero-install via uvx:{"mcpServers":{"openpitch":{"command":"uvx","args":["openpitch-mcp"]}}}// (or "command": "openpitch-mcp" if you pip-installed the package)
Ask your agent: "What's Cognition's ARR, with sources and confidence?" β it calls get_metric/get_provenance and answers from committed data (and will flag the public-source discrepancy).
Or just browse the data
π Live dashboard β avierovich.github.io/openpitch (sourced company cards, refreshed daily) β or build locally: openpitch build-dashboard
π Raw data β data/companies/ β plain JSON, diffable, yours to use
π€ A2A Agent Card β generated at dashboard/dist/.well-known/agent.json
Data status: live, refreshed daily by CI. Figures are probabilistic, public-source intelligence β every number carries its source, confidence score, and date, and open quality items are tracked in public. See the methodology and the correction workflow.
The git repo is the database. There's no server to run. See the FRD for the full design.
Build on it (composability)
OpenPitch emits typed, confidence-scored events when something material changes β so other agents can react:
You're buildingβ¦
Subscribe to
OpenPitch becomesβ¦
A newsletter agent
all material events
your content pipeline's data source
A press/PR workflow
funding/valuation events, confidence β₯ 0.8
your "time to call the company" trigger
Investor outbound
universe entries, growth thresholds
your targeting signal
Events ship on MCP and a raw events/feed.jsonl. Schemas are versioned. See the events spec.
How we compare
OpenPitch is complementary to the incumbents, not a rip-and-replace. We win a narrow wedge; we lose on breadth and verification β and we're honest about both.
PitchBook / CB Insights
Crunchbase
Harmonic
MAGNiTT / Wamda
OpenPitch
Price
$20kβ100k/yr
Freemium
Custom
$/regional
Free & open
Freshness
Weeksβmonths
Variable
Days
Weeks
Daily
In your AI agent (MCP)
β
β
β
β
β
Every figure sourced + confidence-scored
β
β
β
β
β
Contradiction detection
β
β
β
β
β
Coverage breadth
βββ
βββ
ββ
β (MENA)
narrow (by design)
Verified, diligence-grade
β
β
β
β
β (probabilistic)
The honest pitch:the free, fresh, AI-native first look β every number sourced β before you pull the expensive verified report. For an investment decision, you still need the incumbents. Full mapping, feature matrix & pricing: docs/COMPETITIVE-ANALYSIS.md Β· spreadsheet.
Coverage
Global AI startups β 140+ profiled across 12 sectors (including Chinese AI labs and European names Western trackers miss), with a top 50 dynamically ranked by VC attention (valuation + funding activity β not ARR, to avoid circularity). The list moves as attention shifts; companies entering/leaving the top 50 is itself a tracked signal, and auto-discovery grows the universe daily.
MENA AI/tech segment β a dedicated regional set (an open, AI-native alternative to MAGNiTT/Wamda). Honest caveat: MENA disclosure is lighter than the US, so this segment launches with lower confidence/coverage, clearly labeled.
OpenPitch is transparently probabilistic. Many figures are estimates derived from public, self-reported, sometimes-contradictory sources. We surface confidence and provenance precisely so you can judge for yourself. This is not investment advice, and figures are not guaranteed accurate. Always verify before acting.
Roadmap
Seed universe (global AI + MENA segment) + auto-discovery (news, funding digests, 21-sector backfill, China feed)
Core data model + reconciliation engine (confidence, consensus, contradiction) β tested
Contributions welcome β especially new source adapters (one file each) and watchlist curation. See the FRD for architecture.
Who built this
OpenPitch is built and run by Mohamed Abdulhadi,
a product manager β working with AI agents (Claude Code) that wrote much of the code and now
operate the daily pipeline and its public data corrections. That's not a footnote; it's the
product demonstrating itself: an agent-native database, built and maintained agent-natively,
with every commit and correction in the open. Questions, feedback, or collaboration β
connect on LinkedIn or open an issue.