Agentβ™₯︎Age
Catalog

io.github.RudrenduPaul/agent-eval

Official

by RudrenduPaul Β· Python

Statistical regression testing for LLM agents: p-value, effect size, and CI on behavior change.

io.github.RudrenduPaul/agent-eval (MCP Server)

This MCP server supports statistical regression testing for LLM agents. It evaluates behavior change using metrics such as p-value, effect size, and confidence intervals (CI). The repository is associated with the topic areas crewai, langchain, langgraph, llm-evaluation, and openai-agents-sdk.

πŸ› οΈ Key Features

  • Statistical regression testing for LLM agents
  • P-value computation for behavior change
  • Effect size measurement
  • Confidence intervals (CI) on behavior change

πŸš€ Use Cases

  • Detecting and quantifying behavior change in agent outputs
  • Measuring statistical significance of agent behavior regressions
  • Supporting LLM agent evaluation workflows (e.g., across agent frameworks)

⚑ Developer Benefits

  • Provides statistical outputs (p-value, effect size, CI) for agent evaluation
  • Fits into evaluation pipelines using related tooling topics (e.g., llm-evaluation)

⚠️ Limitations

  • Only the server name, description, and excerpted readme metadata are available here; tool capabilities beyond p-value/effect size/CI are not specified in the provided data.

Topics

crewailangchainlanggraphllm-evaluationopenai-agents-sdkregression-testingstatisticsagent-testingp-valuepromptfoo-alternative
io.github.RudrenduPaul/agent-eval - agentage MCP Catalog