Agent♥︎Age
Catalog

io.github.RudrenduPaul/inferbench

Official

by RudrenduPaul · Python

Benchmarks local LLM inference speed (tokens/sec) on your own hardware via MCP tools.

io.github.RudrenduPaul/inferbench MCP Server

This MCP server provides tools for benchmarking local LLM inference speed, reporting tokens/sec on your own hardware. It is associated with the InferBench project and includes a CLI with npm and PyPI releases, targeting developers who need repeatable performance measurement across setups.

🛠️ Key Features

  • Benchmarks local LLM inference speed (tokens/sec)
  • Uses MCP tools to run measurements on your own hardware
  • Supports related model/inference ecosystems referenced in topics (e.g., gguf, llama-cpp, mlx, omlx)

🚀 Use Cases

  • Measure tokens-per-second for local LLM inference
  • Compare performance on different hardware or runtimes
  • Run benchmarks via CLI (npm / PyPI packaging indicated)

⚡ Developer Benefits

  • Fits developer tooling workflows for performance benchmarking
  • Provides a CLI distribution (inferbench-cli) for repeatable measurement

⚠️ Limitations

  • Description provided specifies benchmarking output (tokens/sec) but does not detail datasets, configuration options, or measurement methodology.

Topics

apple-siliconbenchmarkbenchmarkingclideveloper-toolsggufinferencellama-cppllmlocal-llmmlxomlxpythontokens-per-secondtypescript