Agent♥︎Age
Catalog

FitLLM

Official3 toolsLive

by click6067-ship-it · JavaScript

Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.

run.fitllm/fitllm MCP Server

The run.fitllm/fitllm MCP server provides a read-only LLM sizing check to determine whether a local model will fit on a GPU, multi-GPU rig, or Apple Silicon Mac. It performs architecture-aware memory math focused on exact VRAM and KV-cache requirements, using “will-it-run” style analysis.

🛠️ Key Features

  • Exact VRAM and KV-cache math
  • Architecture-aware fitting for GPUs and Apple Silicon Macs
  • Read-only fit verdicts (fit/on not-fit)
  • Zero-dependency engine behavior (per description)
  • Tool count: 3

🚀 Use Cases

  • Validate whether a specific local LLM can run within available VRAM
  • Compare feasibility across NVIDIA, AMD, and Apple Silicon setups
  • Plan inference configurations for quantization and KV-cache constraints
  • Assess setups using common formats and runtimes (e.g., GGUF, llama-cpp, MLX)

⚡ Developer Benefits

  • CLI-oriented workflow (topics include cli)
  • Focus on memory-calculator-style outputs
  • Helps with inference planning for local-llm usage (local llama)

⚠️ Limitations

  • Described as read-only for fit checking; no mention of model execution, downloads, or tuning

Topics

apple-siliconkv-cachellmmemory-calculatormoequantizationvramggufinferencellama-cpplocal-llmlocalllamamlxnvidiaollamavram-calculatoramdclimlawill-it-run