🚀 Maximize your product's SEO. Submit to 240+ directories in 1-click with DirSubmit. Launch Now
SineFrame M3 logo

SineFrame M3

Test MCP servers and the agents that call them in pytest

2026-10-08

Product Introduction

  1. Definition: SineFrame M3 is an open-source, Apache-2.0 licensed pytest framework and continuous integration (CI) gating system specifically designed for testing Model Context Protocol (MCP) servers. It operates as a specialized testing harness that executes tests against real AI agent binaries like Claude Code, Codex, and OpenCode, rather than just a model API.
  2. Core Value Proposition: It exists to solve the critical gap in MCP server development where unit tests pass, but the AI agent fails to call the correct tool due to issues like prompt misunderstanding or tool description mismatches. Its primary value is ensuring MCP server reliability by validating actual agent-server interaction, preventing broken behavior from reaching production, and enabling robust CI/CD gating for pull requests.

Main Features

  1. Multi-Agent Harness Testing: M3 runs your MCP server tests through the actual binaries of popular coding agents (Claude Code, Codex, OpenCode, Pi). It intercepts and records every MCP tool call made during the test execution. This allows developers to write assertions in plain pytest that verify not just server responses, but that the agent invoked the expected tools with the correct arguments.
  2. Direct Server Testing (No API Key Required): Alongside agent tests, M3 supports direct pytest tests for MCP server schemas, responses, errors, and internal state. These tests run without any model or API key, providing fast, deterministic validation of server logic and contracts independently of agent behavior.
  3. CI/CD Gate with Evidence Upload: The m3 ci test --upload command runs the test suite in a CI environment and enforces a strict gate: if any assertion fails, or if zero tests execute (e.g., due to misconfiguration or test skipping), the CI run fails. Detailed traces, tool calls, and results are uploaded to the hosted platform (app.m3.sineframe.com) for review, providing an audit trail for every pull request.
  4. Agent & Version Comparison: M3 enables side-by-side testing across different agents (Claude Code vs. Codex) and different versions of the same agent (e.g., Codex 0.159.3 vs. 0.160.0). This allows developers to catch regressions or behavioral differences introduced by agent updates before they impact end-users.
  5. Bring Your Own Agent (ACP Support): For agents supporting the Agent Client Protocol (ACP), M3 can be configured via a pytest marker to execute tests using a custom agent binary. This extends the testing framework to in-house or niche agents, recording their tool calls for the same assertions.

Problems Solved

  1. Pain Point: The "silent failure" where an AI agent answers a user's query from its internal knowledge or incorrectly reasoned context without calling the intended MCP tool, despite the server being functional and correctly described.
  2. Target Audience: MCP server developers, AI tooling engineers, platform teams building internal agent ecosystems, and DevOps engineers responsible for CI/CD pipelines of AI-integrated applications.
  3. Use Cases:
    • Pre-merge Validation: Gating pull requests to ensure new MCP server features or changes don't break interaction patterns with Claude Code or other agents.
    • Agent Update Safety Net: Automatically testing an MCP server against the nightly or beta build of an AI agent binary to proactively identify breaking changes.
    • Tool Description Optimization: Iteratively refining tool names, descriptions, and argument schemas by observing which prompts successfully trigger the correct tool calls across different agents.
    • Cross-Agent Compatibility: Ensuring an MCP server provides a consistent and reliable experience for users regardless of whether they use Claude Code, Codex, or another supported agent.

Unique Advantages

  1. Strengths & Limitations (Pros & Cons):

    • Pros:
      • Real-World Fidelity: Tests using actual agent binaries capture the full tool discovery, reasoning, and calling flow, unlike mocked API tests.
      • Developer Experience: Integrates seamlessly into existing Python/pytest workflows; tests are plain Python files stored in the repository.
      • Comprehensive CI: The "zero-executed-tests fail" rule prevents configuration drift and ensures test suites remain valid.
      • Open Source & Extensible: Apache 2.0 license allows for inspection, customization, and integration into private infrastructures.
    • Cons:
      • Execution Speed & Cost: Running tests through large language model (LLM)-based agents is significantly slower and potentially more expensive than direct unit tests.
      • Non-Determinism: Agent outputs can be stochastic, potentially causing flaky tests that require careful prompt and assertion design to mitigate.
      • Infrastructure Complexity: Requires managing the installation and execution of various agent binaries (Claude Code, Codex, etc.) in both local and CI environments.
  2. Key Alternatives & Differentiation:

    • Standard Pytest + MCP SDK Mocking: The typical alternative is writing unit tests with mocked clients. Differentiation: M3 tests the real integration, catching failures in the agent's decision logic that mocks will always pass.
    • Manual Testing with Agents: Developers manually running prompts through an IDE with an agent. Differentiation: M3 automates this, provides reproducible assertions, and integrates it into CI, scaling beyond ad-hoc checks.
    • End-to-End UI Testing (e.g., Playwright): Testing the final user application. Differentiation: M3 operates at the MCP protocol layer, isolating server/agent issues from frontend bugs and providing much more precise, faster feedback for server developers.

Frequently Asked Questions (FAQ)

  1. How does SineFrame M3 testing work with Claude Code? M3 launches the local Claude Code binary, connects it to your MCP server under test, executes the defined prompt, and programmatically records all MCP tool calls and responses for assertion within a pytest function.
  2. Is an API key required to use SineFrame M3 for testing? No, for direct MCP server tests (schema, responses), no API key is needed. For agent tests, you must have the respective agent (e.g., Claude Code, Codex) installed and authenticated locally, as M3 uses the local binary.
  3. Can I use SineFrame M3 in my GitHub Actions CI pipeline? Yes, the m3 ci test command is designed for CI environments. It requires setting up the necessary agent binaries (e.g., Codex CLI) in the CI runner and will pass/fail the pipeline based on test results and evidence upload.
  4. What is the difference between M3 and the MCP Inspector? The MCP Inspector is a debugging tool for observing traffic. SineFrame M3 is an automated testing and assertion framework that uses similar introspection to enable automated, repeatable tests and CI gating.
  5. Does M3 support testing MCP servers written in languages other than Python? Yes, the core function of M3 is to test the MCP server process via stdio. The server can be implemented in any language (Go, TypeScript, Rust) as long as it communicates via the MCP stdio protocol. The test definitions themselves are written in Python/pytest.

Submit to 240+ Directories with 1-Click

Maximize your product's SEO and drive massive traffic by automatically submitting it to over 240 curated startup directories using DirSubmit.

Related Products

Subscribe to Our Newsletter

Get weekly curated tool recommendations and stay updated with the latest product news