The Problem
Every time you send an image or long prompt to GPT-4o / Claude / Gemini, you burn 1,000+ tokens on processing that could happen locally for free.
Traditional: Image -> Cloud LLM (1,200 tokens) -> Answer
LatentGate: Image -> Local Ollama (FREE) -> Cloud LLM (200 tokens) -> Answer
Features
| Feature | Description |
|---|
| Local-First | Vision and text compression runs on Ollama (free, no API key needed) |
| Token Optimizer | Deterministic, fact-preserving compression: 66% fewer tokens on a realistic dev corpus in ~1ms (benchmark) |
| MCP Server | Works with Claude Desktop, Cursor, Cline, Continue, Zed |
| Selective Decoding | For video, only call API when scene changes (~2.85x fewer calls) with cosine similarity |
| Text Compression | Long prompts, conversations, RAG docs compressed locally |
| Speed Optimized | Connection pooling, model preloading, parallel processing |
| Multi-Provider | OpenAI, Anthropic, Google, Groq, DeepSeek, Together, Azure, AWS Bedrock, Ollama, or any OpenAI-compatible endpoint |
| REST API | FastAPI server for web application integration |
| Video Processing | Direct video file input with automatic frame extraction |
| Cost Tracking | Persistent cost tracking with SQLite analytics and exportable reports |
| Async Support | Non-blocking async methods for FastAPI, aiohttp, etc. |
| Streaming Responses | Stream responses from remote LLMs |
| Config Persistence | YAML/TOML config files with environment variable overrides |
| Structured Logging | JSON-formatted logging with rotation and correlation IDs |
| Docker Support | Dockerfile and docker-compose for easy deployment |
| Plugin System | Custom processors for domain-specific compression |
| Multi-Language | Support for 30+ languages with automatic detection |
Use it in Claude
LatentGate plugs into Claude as an MCP server. Its optimizer tools work offline, with no
Ollama and no API key β Claude reads big logs, JSON dumps and docs through it and spends a
fraction of the context. A 400-line error log goes from 14,400 to ~100 tokens.
Requires uv (uvx fetches LatentGate from PyPI on first run).
The first run downloads dependencies, which can exceed Claude's 30-second MCP startup limit
on a slow connection; run this once beforehand (later starts take ~2s):
uvx --from "latent-gate[mcp,tokens]" latent-gate --optimizer-benchmark
Claude Code β plugin (MCP server + a skill that tells Claude when to use it):
/plugin marketplace add KathanModh259/latent-gate
/plugin install latent-gate@latent-gate
Claude Code β MCP server only:
claude mcp add latent-gate -- uvx --from "latent-gate[mcp,tokens]" latent-gate-mcp
Claude Desktop β add to claude_desktop_config.json and restart:
{
"mcpServers": {
"latent-gate": {
"command": "uvx",
"args": ["--from", "latent-gate[mcp,tokens]", "latent-gate-mcp"]
}
}
}
Then just ask: "Read logs/app.log with latent-gate and tell me why orders fail."
| Tool | Needs Ollama | What it does |
|---|
read_file_optimized | no | Read a text file and return its optimized form (use for files you read, not files you edit) |
optimize_text | no | Optimize text you already have, optionally toward a question / max_tokens budget |
count_tokens | no | Count tokens (tiktoken o200k_base) |
compress_image | yes | Describe an image locally as a ~150-token scene payload |
compress_text / compress_conversation / compress_documents | yes | Optimizer + fact-checked local-LLM compression |
get_stats | yes | Session statistics |
Quick Start
Install
pip install latent-gate
pip install latent-gate[mcp]
pip install latent-gate[api]
pip install latent-gate[video]
pip install latent-gate[embeddings]
pip install latent-gate[langchain]
pip install latent-gate[bedrock]
pip install latent-gate[tokens]
pip install latent-gate[all]
Pull Ollama Models
ollama pull llava:7b
ollama pull llama3:8b
One-Command Quickstart
chmod +x scripts/quickstart.sh
./scripts/quickstart.sh
This starts everything: Ollama β pulls models β API server β website. See scripts/quickstart.sh for options like --no-pull, --no-website, --port 9000.
CLI Usage
latent-gate photo.jpg "What is in this image?" --provider ollama -v
latent-gate --text "Your long prompt here..." --provider ollama -v
latent-gate --text-file prompt.txt --provider openai -v
latent-gate photo.jpg "Analyze" --text "Extra context..." -v
latent-gate photo.jpg "Describe" --json -v
cat prompt.txt | latent-gate --text-file - --compress-only --deterministic --level balanced
latent-gate --optimizer-benchmark
latent-gate --benchmark --benchmark-output reports/benchmark.json
latent-gate-api
Production Hardening
For API deployments, restrict direct image-path reads to trusted directories:
set LATENTGATE_ALLOWED_IMAGE_ROOTS=C:\safe-images;D:\datasets
latent-gate-api
Benchmark before releases so speed and savings are measured, not guessed:
latent-gate --benchmark --json
Python API
Image Query
from latent_gate import LatentGatePipeline, PipelineConfig
config = PipelineConfig(
vision_model="llava:7b",
predictor_model="llama3:8b",
remote_provider="openai",
remote_model="gpt-4o-mini",
)
with LatentGatePipeline(config) as pipeline:
result = pipeline.query("photo.jpg", "What is in this image?")
print(result["answer"])
print(f"Tokens sent: ~{result['tokens_estimated']}")
print(f"Timing: {result['timing']}")
Text Compression
result = pipeline.query_text("Your 500-word prompt here...", mode="auto")
messages = [
{"role": "user", "content": "Help me with Kubernetes setup"},
{"role": "assistant", "content": "Sure! What's your target configuration?"},
{"role": "user", "content": "3 nodes, t3.large, us-east-1 with autoscaling"},
]
result = pipeline.query_conversation(messages, "Now give me the setup commands")
documents = ["doc1 text...", "doc2 text...", "doc3 text..."]
result = pipeline.query_documents(documents, "How do I implement JWT refresh?")
result = pipeline.query_universal(text="Explain this code...", image="screenshot.png")
Batch Processing
results = pipeline.query_batch(image_paths, "Describe each scene")
results = pipeline.query_batch(image_paths, "Describe each scene", parallel=True, max_workers=4)
results = pipeline.query_batch_texts(text_list, question="Summarize each")
Streaming
for token in pipeline.query_stream("photo.jpg", "Describe this"):
print(token, end="", flush=True)
for token in pipeline.query_text_stream("Long prompt...", mode="compress"):
print(token, end="", flush=True)
REST API
Start Server
latent-gate-api
LATENTGATE_HOST=127.0.0.1 LATENTGATE_PORT=9000 latent-gate-api
$env:LATENTGATE_HOST="127.0.0.1"; $env:LATENTGATE_PORT="9000"; latent-gate-api
set LATENTGATE_HOST=127.0.0.1 && set LATENTGATE_PORT=9000 && latent-gate-api
Endpoints
| Method | Endpoint | Description |
|---|
GET | /health | Health check (Ollama connection status) |
GET | /stats | Session usage statistics |
POST | /query/image | Image query |
POST | /query/text | Text compression |
POST | /query/conversation | Conversation compression |
POST | /query/documents | RAG document compression |
POST | /query/universal | Auto-detect input type |
POST | /query/image/upload | Upload image for query |
Example Requests
import requests
response = requests.post("http://localhost:8000/query/image", json={
"image_path": "photo.jpg",
"question": "What is in this image?"
})
response = requests.post("http://localhost:8000/query/text", json={
"text": "Your long prompt here...",
"question": "Summarize this",
"mode": "auto"
})
response = requests.get("http://localhost:8000/health")
print(response.json())
Async Support
import asyncio
from latent_gate import AsyncLatentGatePipeline, PipelineConfig
async def main():
async with AsyncLatentGatePipeline() as pipeline:
result = await pipeline.query("photo.jpg", "What is this?")
result = await pipeline.query_text("Long prompt...")
results = await pipeline.query_many_images(
["img1.jpg", "img2.jpg", "img3.jpg"],
"Describe each image",
max_concurrent=3,
)
asyncio.run(main())
Video Processing
from latent_gate import LatentGatePipeline, PipelineConfig, VideoProcessor, VideoConfig
config = PipelineConfig(
vision_model="llava:7b",
remote_provider="ollama",
remote_model="llama3:8b",
)
video_config = VideoConfig(
fps=1.0,
max_frames=100,
quality=95,
resize_width=640,
)
with VideoProcessor(config, video_config) as processor:
result = processor.process_video("video.mp4", "Describe the action")
print(f"Frames processed: {result['total_frames']}")
print(f"Unique scenes: {result['statistics']['unique_scenes']}")
print(f"Skip rate: {result['statistics']['skip_rate']}")
Configuration
Config File
ollama_base_url: http://localhost:11434
vision_model: llava:7b
predictor_model: llama3:8b
remote_provider: openai
remote_model: gpt-4o-mini
selective_decoding: true
similarity_threshold: 0.85
use_embeddings: true
enable_caching: true
temperature: 0.1
request_timeout: 120
track_costs: true
cost_db_path: "latentgate_costs.db"
from latent_gate import get_config, LatentGatePipeline
config = get_config("latentgate.yaml")
with LatentGatePipeline(config) as pipeline:
result = pipeline.query("photo.jpg", "Describe this")
Environment Variables
| Variable | Description | Default |
|---|
OPENAI_API_KEY | OpenAI API key | - |
ANTHROPIC_API_KEY | Anthropic API key | - |
GOOGLE_API_KEY | Google API key | - |
LATENTGATE_REMOTE_PROVIDER | Override remote provider | openai |
LATENTGATE_REMOTE_MODEL | Override remote model | provider default (e.g. gpt-4o-mini, claude-sonnet-5) |
LATENTGATE_VISION_MODEL | Override vision model | llava:7b |
LATENTGATE_LOG_LEVEL | Log level | INFO |
LATENTGATE_LOG_FILE | Log file path | - |
LATENTGATE_LOG_JSON | JSON log format | false |
LATENTGATE_TRACK_COSTS | Enable cost analytics | false |
LATENTGATE_COST_DB_PATH | Path to SQLite DB | .latentgate_costs.db |
LATENTGATE_API_KEY | Require Authorization: Bearer <key> on the API (incl. /compress; WebSocket clients may pass ?api_key=) | unset (open) |
LATENTGATE_COMPRESSION_LEVEL | Token optimizer level: lossless, balanced, aggressive | balanced |
LATENTGATE_COMPRESSION_STRATEGY | auto (optimizer + fact-checked local LLM) or deterministic | auto |
LATENTGATE_TARGET_TOKEN_BUDGET | Max tokens for a compressed prompt (0 = no budget) | 0 |
LATENTGATE_MAX_OUTPUT_TOKENS | Max tokens the cloud model may generate per answer | 4096 |
LATENTGATE_PRELOAD | Warm Ollama models in the background at API startup | true |
LATENTGATE_MAX_CONCURRENT_REQUESTS | Max concurrent pipeline calls in the API server | 3 |
LATENTGATE_CORS_ORIGINS | Comma-separated allowed CORS origins | http://localhost:5173 |
Save Config
from latent_gate import PipelineConfig, save_config
config = PipelineConfig(remote_provider="anthropic", remote_model="claude-sonnet-5")
save_config(config, "my_config.yaml")
Docker
docker compose up -d
docker compose --profile setup up ollama-init
docker compose --profile monitoring up -d
docker build -t latent-gate .
docker run -p 8000:8000 latent-gate
The docker-compose setup includes:
- latent-gate API server (port 8000)
- Ollama local LLM server (port 11434)
- ollama-init container that auto-pulls required models (profile:
setup)
- Prometheus metrics collector on port 9090 (profile:
monitoring)
- Grafana dashboard on port 3000 (profile:
monitoring, credentials: admin/latentgate)
Monitoring Stack
Start with monitoring:
docker compose --profile monitoring up -d
Access:
The LatentGate dashboard auto-loads in Grafana with 9 panels covering request rate, latency percentiles (p50/p95/p99), token savings, error rates, pipeline health, and endpoint breakdown.
Customize via environment variables:
GRAFANA_ADMIN β Grafana admin username (default: admin)
GRAFANA_PASSWORD β Grafana password (default: latentgate)
GRAFANA_ANONYMOUS β Enable anonymous access (default: true)
LATENTGATE_ENABLE_METRICS β Enable Prometheus metrics (default: true)
LatentGate works as a Model Context Protocol (MCP) server with every major AI coding tool. Your AI assistant automatically compresses images, long prompts, and documents before they reach the cloud model.
| Tool | Status | Setup |
|---|
| VS Code / Copilot | Supported | Extension |
| Claude Desktop | Supported | MCP Config |
| Claude Code (CLI) | Supported | Skill |
| Cursor | Supported | MCP Config |
| Cline (VS Code) | Supported | MCP Config |
| Continue.dev | Supported | MCP Config |
| Zed Editor | Supported | MCP Config |
VS Code Extension
code --install-extension KathanModh259.latent-gate-vscode
Features:
- Right-click any image to compress with LatentGate
- Select text and press
Ctrl+Shift+Alt+C to compress
- Cost dashboard in activity bar
- Auto-configures MCP for Copilot Chat
- Status bar showing token savings
MCP Setup
For Claude, see Use it in Claude. For Cursor, Cline, Continue, Zed and other
MCP clients, use the same server command:
{
"mcpServers": {
"latent-gate": {
"command": "uvx",
"args": ["--from", "latent-gate[mcp,tokens]", "latent-gate-mcp"]
}
}
}
Or install it into your environment (pip install "latent-gate[mcp,tokens]") and use
"command": "latent-gate-mcp". Note the command is latent-gate-mcp β plain latent-gate
is the CLI and will not speak MCP. The image and compress_* tools additionally need
ollama pull llava:7b and ollama pull phi3:mini.
See integrations/ folder for detailed setup guides per tool.
Speed Optimizations
| Optimization | What It Does | Impact |
|---|
| Connection Pooling | Reuses HTTP connections via requests.Session | ~30-50% faster per call |
| Model Preloading | Warms up Ollama models on init (keep_alive) | Eliminates 5-15s cold start |
| Shorter Prompts | Optimized extraction prompts produce fewer output tokens | ~20% faster generation |
| 3-Tier JSON Parsing | Fast parse, extract from text, LLM fallback | Avoids slow LLM call 90% of time |
| Parallel Processing | Image and text processed simultaneously via ThreadPool | ~40% faster combined queries |
| Content-Hash Caching | Disk cache for repeated images | Instant on cache hit |
| Selective Decoding | Cosine similarity skips redundant API calls | ~2.85x fewer calls |
Cost Benchmarks
Image Queries (by provider, estimated)
Estimates from each provider's published image-token formulas versus a ~150-token local description.
| Provider | Raw Image Tokens | LatentGate Tokens | Savings |
|---|
| OpenAI GPT-4o (high detail) | ~1,105 | ~150 | ~86% |
| Claude 3.5 Sonnet (1MP image) | ~1,334 | ~150 | ~89% |
| Gemini 2.0 Flash | ~258 | ~150 | ~42% |
Text: measured, reproducible
Deterministic optimizer on built-in realistic inputs, counted with tiktoken (o200k_base).
Facts kept = share of numbers, identifiers, URLs, file names and code preserved verbatim.
Reproduce with latent-gate --optimizer-benchmark (no Ollama or API key needed).
| Input | Tokens | lossless | balanced (default) | aggressive |
|---|
| Pretty-printed API JSON | 1,407 | β33%, facts 100% | β33%, facts 100% | β33%, facts 100% |
| 60 repeated log lines + trace | 2,235 | 0% | β92% (ranges kept) | β92% |
| Prompt pasted 3Γ | 121 | β61%, facts 100% | β61%, facts 100% | β61%, facts 100% |
| Verbose spec with 6 requirements | 127 | 0% | β20%, facts 100% | β51%, facts 90% |
| Code review request | 57 | β5% | β16%, facts 100% | β28%, facts 100% |
| 5 RAG chunks + question | 190 | 0% | β58% (2 relevant docs kept) | β58% |
| Total | 4,137 | β13%, facts 100% | β66% | β67% |
Logs lose individual ids/timestamps when folded (the fold keeps first, last and value ranges),
and RAG drops facts from documents irrelevant to the question β both by design.
How it works (safest stage first; see latent_gate/optimizer.py):
- Protect code blocks, inline code, URLs and quoted strings β restored byte-for-byte
- Lossless: whitespace/Unicode cleanup, JSON minification (values untouched), duplicate folding
- Log folding: runs of log lines differing only in numbers/ids β first, last, and ranges
- Filler: pure pleasantries ("Hi!", "Thanks in advance!") and hedging phrases removed
- Selection (only over a budget, or
aggressive): BM25 question-relevance + requirement cues,
original order preserved, the user's actual ask is never dropped
Guarantees: output never has more tokens than input; same input β same output (so provider prompt
caching keeps working); when a local LLM rewrite is used it must be smaller and keep every fact,
otherwise the deterministic result is sent.
from latent_gate import optimize
r = optimize(long_prompt, question="What failed?", level="balanced", max_tokens=2000)
print(r.optimized_tokens, r.savings_pct, r.stages)
Video
Selective decoding skips remote calls for frames similar to the previous one (~2.85x fewer calls
on typical footage).
At Scale (10,000 image queries with gpt-4o-mini, estimated)
| Metric | Traditional | LatentGate | Savings |
|---|
| Input tokens | 12,000,000 | 2,000,000 | 10M tokens |
| Cost | $1.80 | $0.30 | $1.50 (83%) |
Cost Tracking
from latent_gate import CostTracker
tracker = CostTracker()
tracker.record_usage(
query_type="image",
provider="openai",
model="gpt-4o-mini",
input_tokens=150,
output_tokens=200,
tokens_saved=1000,
compression_ratio=6.7,
latency_ms=1500,
)
stats = tracker.get_session_statistics()
print(f"Total cost: ${stats['total_cost']:.4f}")
print(f"Tokens saved: {stats['total_tokens_saved']}")
projection = tracker.get_cost_projection(
daily_queries=1000,
provider="openai",
model="gpt-4o-mini"
)
print(f"Monthly savings: ${projection['savings']['monthly']:.2f}")
tracker.export_report("usage_report.json", fmt="json")
tracker.export_report("usage_report.csv", fmt="csv")
Multi-Language Support
from latent_gate import detect_language, MultiLanguageProcessor
lang = detect_language("Esto es un texto en espaΓ±ol")
print(f"Detected: {lang.name} ({lang.confidence:.0%})")
processor = MultiLanguageProcessor()
text, lang_info = processor.process("Texto en espaΓ±ol para analizar")
print(f"Language: {lang_info.name}, Translated: {text[:100]}...")
Project Structure
latent-gate/
βββ latent_gate/
β βββ __init__.py # Package exports and version
β βββ config.py # PipelineConfig dataclass
β βββ config_loader.py # YAML/TOML/JSON config loading
β βββ payload.py # SemanticPayload (compact representation)
β βββ text_processor.py # TextPayload + TextProcessor (local compression)
β βββ local_processor.py # X-Encoder + Predictor (Ollama vision pipeline)
β βββ remote_decoder.py # Y-Decoder (OpenAI, Anthropic, Google, Ollama)
β βββ selective_decoder.py # Cosine/Jaccard similarity for skip decisions
β βββ fast_client.py # Connection pooling + model preloading
β βββ cache.py # Content-hash disk cache
β βββ pipeline.py # LatentGatePipeline (main orchestrator)
β βββ async_pipeline.py # AsyncLatentGatePipeline
β βββ video_processor.py # Video frame extraction + batch processing
β βββ cost_tracker.py # SQLite-based cost analytics
β βββ mcp_server.py # MCP server (Model Context Protocol)
β βββ api_server.py # FastAPI REST server
β βββ cli.py # Command-line interface
β βββ logging_config.py # Structured logging with rotation
β βββ plugin_system.py # Custom processor plugins
β βββ multilang.py # Multi-language detection and translation
βββ integrations/
β βββ agent_skills/ # Prompt compression skill + scripts
β βββ vscode-extension/ # VS Code extension source
β βββ cursor/ # Cursor rules and MCP config
β βββ continue_dev/ # Continue.dev config
β βββ langchain/ # LangChain integration wrapper
β βββ llamaindex/ # LlamaIndex retriever integration
β βββ openai_functions/ # OpenAI/Anthropic function schemas
βββ tests/ # 240+ tests (unit + integration)
βββ website/ # React-based analytics dashboard & landing page
βββ deployments/ # Kubernetes Helm configs
βββ .github/workflows/ # CI + publish workflows
βββ Dockerfile
βββ docker-compose.yml
βββ pyproject.toml
βββ requirements.txt
Contributing
Contributions welcome! See CONTRIBUTING.md.
Development Setup
git clone https://github.com/KathanModh259/latent-gate.git
cd latent-gate
python -m venv .venv
source .venv/bin/activate
.venv\Scripts\Activate.ps1
pip install -e ".[dev]"
Run Tests
Priority Areas
- Additional vision model support (Florence-2, InternVL, Qwen-VL)
- Custom similarity plugins for domain-specific use cases
- WebSocket support for real-time streaming
- Advanced cost analytics and optimization suggestions
- Plugin development for specialized industries
- Test coverage improvements
- Documentation and examples
Citation
@software{latentgate2026,
author = {Kathan Modh},
title = {LatentGate: Local-First Semantic Compression Pipeline},
year = {2026},
url = {https://github.com/KathanModh259/latent-gate},
version = {1.3.0}
}
Inspired by VL-JEPA (Meta FAIR, 2025).
License
Custom Proprietary License β see LICENSE.
Built by Kathan Modh
Process locally. Send smart. Pay less.