Agent♥︎Age
Catalog

io.github.hidai25/evalview-mcp

Official

by hidai25 · Python

Regression testing for AI agents. Golden baselines, CI/CD, LangGraph, CrewAI, OpenAI, Claude.

EvalView MCP Server: io.github.hidai25/evalview-mcp

This MCP server supports regression testing for AI agents by using golden baselines and snapshot-style detection to flag silent behavioral changes. It is positioned for CI/CD workflows and can be used alongside LangGraph and CrewAI with LLM providers such as OpenAI and Claude.

🛠️ Key Features

  • Regression testing for AI agents
  • Golden baselines / snapshot-style tracking
  • CI/CD-oriented evaluation
  • Compatible with LangGraph and CrewAI
  • Mentions OpenAI and Claude

🚀 Use Cases

  • Detecting changes in agent behavior over time
  • Recording “what the agent does today” for later comparison
  • Running automated agent evaluation in CI/CD pipelines
  • Validating agent outputs using baselines

⚡ Developer Benefits

  • Helps identify silent changes via baseline comparison
  • Supports agent evaluation workflows across common frameworks (LangGraph, CrewAI)
  • Fits developer testing stacks referenced by the topics (e.g., pytest, CLI)

⚠️ Limitations

  • Source details are limited to the provided description, topics, and readme excerpt; tool surface area and exact MCP endpoints are not included in the available data.

Topics

agent-benchmarkagent-evaluationai-agentscrewaievaluationlanggraphopenai-assistantstestingpytestanthropicagentic-ailangchain-agentautogenclillmmcppythonregression-testing