This MCP server supports regression testing for AI agents by using golden baselines and snapshot-style detection to flag silent behavioral changes. It is positioned for CI/CD workflows and can be used alongside LangGraph and CrewAI with LLM providers such as OpenAI and Claude.
🛠️ Key Features
Regression testing for AI agents
Golden baselines / snapshot-style tracking
CI/CD-oriented evaluation
Compatible with LangGraph and CrewAI
Mentions OpenAI and Claude
🚀 Use Cases
Detecting changes in agent behavior over time
Recording “what the agent does today” for later comparison
Running automated agent evaluation in CI/CD pipelines
Validating agent outputs using baselines
⚡ Developer Benefits
Helps identify silent changes via baseline comparison
Supports agent evaluation workflows across common frameworks (LangGraph, CrewAI)
Fits developer testing stacks referenced by the topics (e.g., pytest, CLI)
⚠️ Limitations
Source details are limited to the provided description, topics, and readme excerpt; tool surface area and exact MCP endpoints are not included in the available data.
Your agent returns 200 and looks fine. But a model update, a provider change, or a one-line prompt edit just made it skip a clarification, call the wrong tool, or quietly drop output quality. Your tests still pass. Your users notice before you do.
EvalView snapshots your agent's behavior — the tools it calls, in what order, with what output — and tells you the moment that behavior changes. Like Jest snapshots, but for tool-calling, multi-turn agents.
OpenAI adapter migration: OpenAI shut down the Assistants API on August 26, 2026.
The latest published EvalView release, 0.8.1, still uses that API; the Responses API
migration is currently unreleased source. If you use openai-assistants, follow the
migration guide before running your tests. An assistant_id
alone cannot preserve your agent's configuration. Other adapters are unaffected.
bash
pip install evalview
bash
evalview snapshot # Record your agent's current behavior as the baseline
evalview check # After any change, diff against the baseline
That's the whole loop. check returns one of:
code
✓ login-flow PASSED behavior matches baseline
⚠ refund-request TOOLS_CHANGED called a different tool, or in a different order
✗ billing-dispute REGRESSION score dropped — output quality fell
It diffs the whole trajectory — tool names, parameters, and order — not just the final string. The deterministic tool + sequence diff runs offline, with no API key. Add an LLM judge only when you want output-quality scoring.
Executing your agent can still incur backend API charges: --no-judge skips the
judge, not those calls. Embedding-based semantic comparison is opt-in.
No agent yet? See it work in 30 seconds:
bash
evalview demo
Why snapshot testing (and not assertions)?
Most eval tools ask you to write down what "good" looks like — assertions, metrics, rubrics. That's a lot of upfront work, and you can only catch the failures you thought to assert.
EvalView inverts it: it records what your agent actually does now, and flags any drift from that. You catch regressions you never anticipated, with zero assertions written. When the new behavior is correct, evalview snapshot accepts it as the new baseline — same as updating a snapshot in Jest.
EvalView
Assertion-based eval tools
Setup
Record current behavior
Write assertions/metrics first
Catches
Any drift from baseline
Only what you asserted
Non-determinism
Multi-variant baselines (up to 5 valid paths)
You handle it
Unit of comparison
Full tool-call trajectory
Usually final output
This makes EvalView a merge-time regression gate, which is a different job from observability (Langfuse, LangSmith) or metric scoring (promptfoo, DeepEval, Braintrust). Many teams run one of those for visibility and EvalView as the gate. Honest comparisons →
EvalView tests itself in public, every day
Every day at 09:00 UTC, on pull requests, and on pushes to main,
Core Dogfood exercises the non-live test suite,
type checks, local mock-agent snapshot / check, evalview demo, end-to-end
flows, and an evalview monitor smoke test. It uses no paid API credentials and
makes no paid inference calls. GitHub runner usage is separate.
Live Provider Checks test the real evaluator and
chat assistant only when a maintainer explicitly opts into paid API use on main.
They have no automatic schedule. Their badge records the last manual run; a green
core badge does not establish live-provider health or rule out provider drift.
Package CI, core dogfood, and live checks have separate badges. Failed or incomplete
checks remain visible within their scope, with logs and reports preserved as
artifacts. Rolling issues use separate dogfood-core and dogfood-live labels.
A provider outage, exhausted quota, or missing credential means live health is
unavailable; it does not prove an agent regression.
The historical incident #264
remains available for maintainer review of fresh evidence from both scopes. Neither
workflow automatically closes it. Trust warnings are evidence to investigate,
not proof of gaming or of a particular root cause.
EvalView also does multi-turn testing, statistical/pass@k runs, record/replay cassettes, model-drift canaries, production monitoring with Slack alerts, and auto-generated regression tests from incidents. These are power-user features — start with snapshot and check, reach for the rest when you need them.
An agent that looked successful kept pulling entire documents into its context and made one question cost $42.93. That experience led me to build EvalView. I wrote about it in “I Was Running an AI Casino. Then I Started Writing Tests for My Agents”. The December 2025 post is the origin story; use the current docs for setup and commands.
Contributing
This is a young project built mostly by one developer. Issues, PRs, and "I tried it and X was confusing" feedback are all genuinely valuable.