Agentโ™ฅ๏ธŽAge
Catalog

CompletionKit

OfficialLive

by homemade-software-inc ยท Ruby

Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.

This MCP server provides prompt evaluation over MCP by running a prompt against a dataset and scoring each output on a 1โ€“5 scale using an LLM judge. It is positioned for prompt-engineering and prompt-testing workflows, with evaluation-metrics and evaluation-framework concepts reflected in its topic set.

๐Ÿ› ๏ธ Key Features

  • Prompt evals over MCP: run prompts on a dataset
  • LLM-judge scoring for each output
  • 1โ€“5 scoring scale

๐Ÿš€ Use Cases

  • LLM evaluation and evaluation frameworks
  • Prompt testing for prompt-engineering changes
  • Measuring output quality with evaluation metrics

โšก Developer Benefits

  • Supports llm-eval and llm-evaluation workflows
  • Uses LLM-as-judge (llm-as-judge) for automated scoring
  • Aligns with llmops, evaluation-metrics, and evaluation-framework terminology

โš ๏ธ Limitations

  • Limited to scoring outputs on a 1โ€“5 scale via an LLM judge

๐Ÿงฉ Related Topics

  • anthropic, llm, llm-as-judge, llm-eval, llm-evaluation, mcp, ollama, openai
  • prompt-engineering, prompt-testing, rails, rails-engine
  • ruby, ruby-on-rails, evaluation-framework, evaluation-metrics
  • llm-evaluation-framework, llm-evaluation-metrics, llmops

Topics

anthropicllmllm-as-judgellm-evalllm-evaluationmcpollamaopenaiprompt-engineeringprompt-testingrailsrails-enginerubyruby-on-railsevaluation-frameworkevaluation-metricsllm-evaluation-frameworkllm-evaluation-metricsllmops