Build an MCP evaluation set from the workflows users need to complete, with explicit starting data, permitted actions, and observable success conditions. Test the server directly, then test the assistant and client using it. A valid tool response and a convincing final sentence are useful signals, but neither alone proves that the user's task was completed correctly.
This guide proposes a starter evaluation set for a fictional room-booking service. The cases are original design examples, not benchmark results or a claim that a specific integration has passed them.
Separate the server contract from the user workflow
Start with direct checks: can the client connect, discover the expected capabilities, submit valid arguments, and interpret results? The official MCP Inspector provides developer tooling for testing and debugging MCP servers, including browser and command-line interfaces. Use tooling compatible with the protocol version and transport your integration supports.
Those checks do not establish that an assistant chooses the right action for an ambiguous request. Add a second layer that runs representative user conversations through your actual client, model, instructions, and tool configuration.
Keep failures attributable to the right layer. A server rejecting a malformed argument may be correct behaviour, while an assistant repeatedly generating that argument is still a workflow problem. Conversely, a helpful refusal by the assistant does not prove the backend rejects unauthorized direct calls.
Start with one outcome you can inspect
For our example, the initial workflow is: find an available room, resolve any required choices, create one authorized reservation, and report its reference. Use a dedicated test environment with synthetic rooms and users.
Define a fixed clock, time zone, and starting state. “Tomorrow afternoon” means something different when the clock changes, so each run needs the same interpretation context. Reset reservations between trials and isolate concurrent runs so one test does not consume another's available room.
A proposed case record could look like this:
Case: book_known_room
Clock: 2026-10-01, Europe/Paris
User: synthetic member of Workspace A
Request: Book Cedar tomorrow from 14:00 to 15:00.
Fixture: Cedar is available; user may book it.
Approval: simulate the required approval in the client.
Expected effect: one matching reservation in Workspace A.
Expected answer: confirmed time and real reservation reference.
Forbidden effects: duplicate booking or another workspace's booking.
This is a planning format, not an MCP message. Adapt it to your product's approval policy. A test runner should not assume that a user has approved an action merely because approval would make the test pass.
Anthropic's agent evaluation guide distinguishes a trial's transcript from its outcome in the environment. It also describes evaluating the model together with the surrounding agent system. That distinction is useful here: inspect the reservation state as well as the conversation, and record which system configuration produced both.
Add cases that change the correct next action
Our recommended starter set varies the decision the assistant must make, rather than merely rewriting the happy-path prompt:
- Known room, available slot: create one reservation after the required approval and return its actual reference.
- Ambiguous room name: two accessible rooms are called Cedar; ask which one before booking.
- Missing time: “Book Cedar tomorrow” requires clarification under this example's policy; do not invent a time.
- No availability: report the lack of a matching slot and offer permitted alternatives without claiming a booking exists.
- Unavailable service: explain that availability could not be established; do not describe an outage as no available rooms.
- Unauthorized workspace: do not expose private room details or create a reservation outside the user's scope.
- Lost response after creation: recover the existing booking through the established operation contract, with no duplicate reservation.
- Request outside tool scope: “Cancel all next month's bookings” must not be treated as permission to use a create-only tool.
These cases are not a universal coverage target. Add scenarios from your own product rules and reviewed support requests, using appropriate consent and data handling. Preserve privacy when translating real incidents into fixtures.
The related guides on error responses and idempotent actions provide concrete contracts for two of these scenarios. An evaluation can only enforce a contract that the team has actually decided.
Write grading rules before running the assistant
For the successful booking case, use deterministic checks wherever the outcome is objective. Query the test database for the workspace, room, start, end, and reservation count. Verify that the reference in the final answer identifies the created reservation. Count secondary effects too if the workflow promises a single confirmation notification.
For ambiguous input, grade whether the assistant asks a question that distinguishes the two rooms and whether it avoids creating a reservation first. Accept different wording that communicates the same decision. Requiring one exact sentence would test phrasing more than behaviour.
Avoid enforcing one exact tool-call sequence unless the order is itself a product requirement. A valid cached room identifier may make an extra lookup unnecessary. Grade the allowed outcome and required checks, while reviewing the trace for prohibited actions or unsupported assumptions.
If you use a model to judge answer quality, provide a narrow rubric and compare its judgments with human review. Keep objective permission and state checks separate from subjective assessments such as clarity. A fluent answer should not compensate for a booking in the wrong workspace.
Record uncertainty and repeated trials honestly
The same scenario may produce different behaviour on separate runs. Anthropic's evaluation guidance recommends multiple trials to account for that variability. Choose a repetition plan you can afford and record the number of trials alongside results; a single successful attempt is weak evidence of consistency. See the evaluation structure discussion.
For our example, report results by case family: completed bookings, correct clarifications, correct handling of empty results, bounded recovery from outages, and access-boundary checks. Keep individual failures inspectable instead of hiding them inside one overall percentage.
Define release rules in advance. A team might require every deterministic authorization check to pass and separately review conversation failures before rollout. That is a proposed release policy, not an industry threshold. Do not choose a passing threshold after seeing the scores.
When a trial cannot run because the test environment is broken, mark it as an infrastructure failure. Do not count it as a successful refusal or quietly remove it from the denominator. If the fixture or expected result was wrong, fix the case and document why.
Preserve enough evidence to reproduce a regression
For each run, save the case version, fixture version, server commit, client version, model identifier, instructions, tool definitions, and relevant configuration. Record tool inputs, safe result summaries, timings, final answer, and final state checks. Redact credentials and keep sensitive traces under appropriate access controls.
Changes to a tool description can alter behaviour even when the backend code is unchanged. Run the same evaluation set before and after such a change, using comparable conditions. Review both improved and regressed cases.
Maintain a stable regression set and a separate set of fresh scenarios used for review. If every failure leads to a prompt tailored to that exact sentence, the known tests may improve without demonstrating broader usability. Add variations in context, permissions, and starting state, not just synonyms.
Make evaluation part of the first integration brief
Before releasing the first workflow, the team should be able to show its case records, fixtures, grading rules, run configuration, and observed failures. Start with a small set you understand deeply, then expand it as the product gains capabilities and real failure evidence.
The Agent Readiness Audit can help frame proposed tools from API documentation. A realistic evaluation set goes further: it establishes what your implementation must demonstrate when an assistant uses those tools, including when clarification or stopping is the correct outcome.
