vibe-hnindex
Local MCP server — index your repo once, search it in every AI session
Keyword (SQLite FTS5) · Semantic (Qdrant + configurable embeddings) · Hybrid — local or cloud embedding providers

MCP server (vibe-hnindex) version: v0.14.0 · hnindex-cli v0.14.0 — Docs · Changelog · GitHub Releases
What this does
vibe-hnindex is a Model Context Protocol server. After you index a folder once, assistants (Claude, Cursor, Windsurf, Antigravity, …) can search that codebase with paths and line ranges — data is stored locally (SQLite + optional Qdrant). Embeddings use Ollama (default), OpenAI, Voyage, Gemini, or an OpenAI-compatible API; see Embedding providers. Cloud providers receive the content being embedded; vectors use Qdrant (Docker, local, or Qdrant Cloud with QDRANT_API_KEY).
Documentation
📚 Full docs site: docs.hnindex.cloud — 16 pages covering Getting Started, Configuration, Tools Reference, Guides, and Code Agent.
| Page | What you'll learn |
|---|
| Introduction | What vibe-hnindex does, key features, how it works |
| Installation | Node, Ollama, Qdrant setup + MCP config |
| Quick Start | 5-minute walkthrough with CLI + agent skill |
| Configuration | All 25+ env vars with embedding model comparison |
| Search | 6 search modes, regex, fuzzy, streaming, cache |
| Code Agent 🆕 | code_session + code_apply with safety scopes |
| Setup MCP | Per-platform config (Claude, Cursor, Antigravity, VS Code...) |
Also available in-repo: docs/getting-started.md, docs/configuration.md, docs/tools-reference.md.
CLI installer (hnindex)
Optional — writes the MCP JSON for you (merge-safe, same npx -y vibe-hnindex block as in the docs):
npm install -g hnindex-cli
hnindex init --mcp antigravity
hnindex init --list
hnindex init-skill --target claude
hnindex init-skill --list
hnindex update
See docs.hnindex.cloud for full documentation.
Install in 5 steps
- Node.js — v22+ (nodejs.org). CI verifies Node 22 and 24. SQLite uses N-API binaries bundled for supported platforms; see Troubleshooting → Windows if
npm i vibe-hnindex fails.
- Embedding provider — choose OpenAI, Voyage, Gemini or a compatible API, or use local Ollama: install from ollama.com, then:
ollama pull bge-m3:567m and keep ollama serve running (or set OLLAMA_URL to a remote server).
- Qdrant — for semantic/hybrid search:
docker run -d --name qdrant -p 6333:6333 qdrant/qdrant (or use Qdrant Cloud). Keyword-only search works without Qdrant.
- MCP config — add the server to your assistant’s MCP settings. Minimal example (self-hosted Qdrant):
{
"mcpServers": {
"vibe-hnindex": {
"command": "npx",
"args": ["-y", "vibe-hnindex"],
"env": {
"OLLAMA_URL": "http://localhost:11434",
"OLLAMA_MODEL": "bge-m3:567m",
"QDRANT_URL": "http://localhost:6333",
"SEARCH_STREAM_ENABLED": "true",
"CODE_AGENT_ENABLED": "true",
"CODE_AGENT_SCOPE": "moderate",
"CHAT_MEMORY_ENABLED": "true"
}
}
}
}
- Restart the IDE or assistant, then in chat ask to index a path and search — see First steps.
For Qdrant Cloud, add QDRANT_API_KEY and set QDRANT_URL to your HTTPS cluster URL — details in Getting started.
Code Graph (v0.14.0)
Build a TypeScript/JavaScript symbol graph with index_code_graph(path, project_name) using SQLite alone. index_codebase also builds it; file indexing and watching update relationships automatically. Use find_references, callers, and graph_context to inspect source evidence and bounded relationship context. No graph database service is required.
AST chunking preserves function/class boundaries where possible. After upgrading, run index_codebase once to rebuild chunks and vectors. See Code Graph guide for resolution limits, offline usage, token budgets and quality evaluation.
Optional rerank
Hybrid search combines keyword results and embedding vectors using RRF. Optional reranking supports Voyage (RERANK_PROVIDER=voyage, VOYAGE_API_KEY or RERANK_API_KEY) and custom HTTP (RERANK_PROVIDER=http, RERANK_URL). Without a reranker, or if it fails, the original retrieval ranking is preserved. Exact symbol/regex modes skip reranking.
| Env | Role |
|---|
SEARCH_RERANK | false disables reranking. |
SEARCH_RERANK_POOL | Candidate pool before final trim (default 50; at least the requested limit). |
RERANK_PROVIDER | none, http, voyage; defaults to http when RERANK_URL is set, otherwise none. |
RERANK_MODEL | Voyage model, default rerank-3-lite. |
RERANK_URL | Custom endpoint; HTTP contract is {query, documents} -> {scores}. |
RERANK_API_KEY | Optional Bearer credential; Voyage can use VOYAGE_API_KEY. |
RERANK_TIMEOUT_MS | Full request/body timeout, default 15000. |
See Configuration.
Timeouts
To prevent hanging when Ollama or Qdrant are unresponsive, vibe-hnindex applies timeouts on all external calls. You can tune these via environment variables:
| Env | Default | Controls |
|---|
OLLAMA_TIMEOUT_MS | 30000 (30s) | Max wait for Ollama /api/embed and /api/tags calls |
QDRANT_TIMEOUT_MS | 15000 (15s) | Max wait for Qdrant API calls (search, upsert, etc.) |
SEARCH_TIMEOUT_MS | 60000 (60s) | Overall timeout for the entire search operation |
Set any of these to a higher value if you have a slow machine or large dataset. Set to 0 to disable the timeout for that layer (not recommended).
Google Antigravity
Use the same mcpServers block as above, but save it in Antigravity’s MCP file:
| |
|---|
| File | mcp_config.json under .gemini/antigravity/ in your user folder |
| Windows | C:\Users\<your-username>\.gemini\antigravity\mcp_config.json |
| macOS / Linux | ~/.gemini/antigravity/mcp_config.json |
| UI | ⋮ menu → MCP → Manage MCP Servers → View raw config |
Step-by-step: Integrations → Google Antigravity.
Features (short)
| |
|---|
| Search | 6 modes: keyword (FTS5+BM25), semantic (Qdrant vectors), hybrid (RRF fusion), regex, symbol, auto |
| Code Agent | code_session — 1 call replaces 5-15 searches. code_apply — safe code changes with auto test/lint/typecheck |
| Chat Memory 🆕 | Auto-track tool calls, semantic search via Qdrant, persistent AI context across sessions |
| Streaming | Parallel keyword+semantic search (~1.5-2× faster), 4-phase progress notifications |
| Fuzzy Search | Levenshtein distance auto-corrects typos ("fucntion" → "function") |
| Smart Context | Task-aware context: impact analysis, test file detection, similar code patterns |
| Storage | SQLite on disk + Qdrant for vectors; 100% local, no cloud required |
| Indexing | Incremental (SHA-1 hash), parallel workers (~3-4× faster), watch mode (auto re-index on save), 40+ languages, .hnindexignore |
| Resilience | Keyword search works without Qdrant or Ollama; graceful degradation |
| Benchmark | Built-in benchmark_search tool — compare streaming vs non-streaming, all search modes |
| Multiple Embedding Providers (v0.13.0) | Ollama, OpenAI, Voyage, Gemini and OpenAI-compatible APIs; provider/model-aware vector collections |
| Multiple Embedding Models | bge-m3 (default), nomic-embed-text, qwen3-embedding, mxbai-embed-large, and more |
Architecture
graph TB
subgraph Input["📂 Input"]
A["💻 Your Codebase<br/>.ts .py .go .rs ..."]
end
subgraph Server["⚙️ vibe-hnindex MCP Server"]
B["🔍 Search Router<br/>keyword | semantic | hybrid"]
C["🔀 RRF Fusion"]
end
subgraph Storage["💾 Storage"]
D[("SQLite<br/>FTS5 + Keyword")]
E[("Qdrant<br/>Vector Embeddings")]
end
subgraph Memory["🧠 Chat Memory (v0.12)"]
F[("SQLite<br/>Chat Context")]
G[("Qdrant<br/>Chat Vectors")]
end
subgraph Infra["🏗️ Infrastructure"]
H["Ollama<br/>Embeddings"]
I["Qdrant<br/>localhost:6333"]
end
subgraph Output["🤖 AI Clients"]
J["Claude · Cursor · Windsurf<br/>Antigravity · VS Code"]
end
A -->|"index_codebase"| Storage
A -->|scan| H
B -->|"keyword"| D
B -->|"semantic"| E
B -->|"hybrid"| C
C --> D
C --> E
B -.->|"auto-track"| F
F --> H
H --> G
D --> J
E --> J
H -.-> I
style F fill:#6366f1,color:#fff
style G fill:#6366f1,color:#fff
style B fill:#f59e0b,color:#000
style J fill:#22c55e,color:#fff
How indexing & search work →
License
MIT — see LICENSE.
Contributing
Issues and PRs: github.com/AndyAnh174/vibe-hnindex.
Ho Viet Anh (AndyAnh174) · hovietanh147@gmail.com · GitHub