Four implementations of one interface drift apart unless something stops them. In the examples repository that something is tests/conformance.mjs: a client written with @modelcontextprotocol/client that connects to any of the servers over Streamable HTTP and runs the same scenario. CI runs it four times, once per language, on every push.

Files in the repository

Every snippet on this page comes from these files; open them to see the whole thing.

What it checks

  1. The catalogue. Exactly search_products, get_product, create_order, cancel_order; each with a description longer than a sentence and an object input schema; search_products annotated read-only and cancel_order annotated destructive.
  2. Reads and pagination. A two-item page for "lamp"; nextCursor present; results summarised (a price string, no priceCents); a category search that needs a second page; following the cursor reaches a page whose nextCursor is null; get_product returns the right product.
  3. Errors that instruct. An unknown id is an error that says "not found"; a limit of 99 is rejected.
  4. Idempotent writes. create_order twice with the same key returns the same order id; the order has a status and total; the same key with a different quantity is a conflict error.
  5. The guarded action. cancel_order without confirm is an error that mentions confirmation and shows the order id; with confirm: true the status becomes cancelled; cancelling again is reported, not repeated.

Run it

node sample-api/server.mjs &
# start any one server, then:
cd tests && npm ci && MCP_URL=http://127.0.0.1:3001/mcp npm test

The scenario is deliberately ordinary JavaScript with a check(condition, label) helper, so it is easy to extend for your own tools.

What it caught

These are the real failures from writing the examples, kept here because each one is a lesson about multi-language integrations:

  • Python used snake_case parameters (product_id) while the other three used camelCase. The schema an assistant sees is the interface, so Python was changed to match, with a comment explaining why.
  • C# dropped nextCursor when it was null, because the SDK's serialiser omits nulls. The assistant would have had no way to know a page was the last one. Fixed with [JsonIgnore(Condition = JsonIgnoreCondition.Never)] on that property.
  • PHP listed zero tools because symfony/finder, which attribute discovery needs, is not a dependency of mcp/sdk. There was no error; the tool list was just empty.
  • PHP rejects bad input differently: as a JSON-RPC invalid-params error instead of an isError tool result. Both are correct under the protocol, so the test accepts either rather than forcing one SDK to imitate another.
  • The Python server had no /health route, so CI's readiness wait timed out while the server was actually fine. The route was added for parity.

None of these would have been found by reading the code. All of them would have surfaced as confusing assistant behaviour in front of a customer.

Adapting it

Replace the four tool names and the scenario steps with yours, keep the structure (catalogue, reads, errors, writes, guarded actions), and run it against every change. If you have one server rather than four, the test is still worth keeping: it is the fastest way to notice that a refactor changed what the assistant sees.