This MCP server classifies prompt injection, jailbreaks, and out-of-scope user input so AI agents can check text safety and scope before acting on it. It exposes a single MCP tool that evaluates untrusted user content for agent use.
🛠️ Key Features
- Provides one tool:
classify_input - Classifies prompt injection, jailbreaks, and out-of-scope user input
- Intended for pre-action safety checks in agents
🚀 Use Cases
- Guardrail step before an agent processes user-provided text
- Filtering or blocking unsafe and out-of-scope requests
- Supporting AI tools with an “AI firewall” classification step
⚡ Developer Benefits
- Simple integration via an MCP tool (
classify_input) - Works with multiple providers (e.g., OpenAI, Anthropic, Google, DeepSeek, Ollama; Google default)
- Configuration via environment variables (e.g.,
KOMA_PROVIDER,GEMINI_API_KEY)
⚠️ Limitations
- Exposes a single classification tool (
classify_input) rather than multiple agent actions
Topics
- ai-security, guardrails, middleware, nodejs, prompt-injection, rag, rate-limiting, typescript, ai-firewall, llm-security, owasp-top-10-llm, ai-governance, ai-tools