Agent♥︎Age
Catalog

io.github.HaseebKhalid1507/velocirag

Official

by HaseebKhalid1507 · Python

Lightning-fast RAG for AI agents. 4-layer fusion, ONNX Runtime, sub-200ms search.

MCP Server: io.github.HaseebKhalid1507/velocirag

This MCP server exposes VelociRAG, a “four-layer retrieval fusion” RAG solution for AI agents. It performs vector similarity, BM25 keyword matching, knowledge graph traversal, and metadata filtering, fused via reciprocal rank fusion and cross-encoder reranking. Retrieval runs on ONNX Runtime (no PyTorch), with sub-200ms warm search and incremental graph updates.

🛠️ Key Features

  • Four retrieval layers: vector similarity, BM25, knowledge graph traversal, metadata filtering
  • Fusion via reciprocal rank fusion and cross-encoder reranking
  • ONNX Runtime execution; no PyTorch, no GPU, no API keys
  • Incremental graph updates
  • MCP-ready with an MCP server and a Unix socket

🚀 Use Cases

  • Agentic RAG for fast retrieval in local AI setups
  • Semantic search combining keyword, vector, graph, and metadata constraints

⚡ Developer Benefits

  • MCP server for agent integration
  • Local inference via CPU-first ONNX Runtime execution
  • Multiple retrieval methods unified in a single pipeline

⚠️ Limitations

  • Source indicates no GPU usage (CPU inference implied)
  • Does not claim support for PyTorch-based or API-key workflows

Topics

ai-agentsbm25faissknowledge-graphllmmcpmcp-serveronnxpythonragretrieval-augmented-generationsemantic-searchvector-databasevector-searchagentic-ragcpu-inferenceembeddingslocal-aimodel-context-protocolretrieval