An assistant should receive a manageable page of records, a precise way to continue, and an unambiguous signal when the search is complete. Without that contract, a useful-looking answer can silently describe only the first slice of your data.

For SaaS teams exposing search or reporting through MCP, pagination is part of answer correctness. A request for “a few recent tickets” and a request for “every unresolved ticket before Friday” need different stopping rules, even when they use the same tool.

Separate tool discovery from business data

MCP has pagination for discovering things such as available tools. In the November 25, 2025 specification, tools/list, resources/list, resources/templates/list, and prompts/list use opaque cursors. A response can include nextCursor, which the client sends back as cursor; the server controls page size. Clients must not infer a fixed page length. These are protocol list operations. They do not automatically paginate the tickets returned by your own search_tickets tool. See the MCP pagination specification.

For business records, define pagination in the tool's input and result contract. Document it alongside the filters, sorting rules, and access boundaries. An assistant should not have to guess whether an array means “everything” or “the first batch.”

Give each page a clear continuation contract

Consider an illustrative support product. Its assistant needs to find unresolved tickets created before a specified time. The following JSON is an example of custom tool arguments, not a new MCP protocol method:

{
  "status": "unresolved",
  "created_before": "2026-10-02T00:00:00Z",
  "sort": "created_at_asc",
  "page_size": 2
}

The tool's business result might contain:

{
  "items": [
    { "id": "ticket_041", "subject": "Export failed" },
    { "id": "ticket_058", "subject": "Invitation expired" }
  ],
  "next_page_token": "opaque-example-token"
}

To continue, the assistant repeats the filters and sort with that token. In this example contract, the final page omits next_page_token. The example token is a placeholder, not a suggested token implementation.

Google's API design guidance provides a useful precedent: continuation tokens are opaque, other query arguments remain consistent between pages, and a short or even empty page does not necessarily mean the collection has ended. It also stresses that a token is not authorization. These are Google API design rules, not universal MCP tool requirements; adopt and document the behavior your tool actually implements. See AIP-158.

For this support tool, we would document a default of 20 records and a maximum of 50, selected after measuring real response sizes. Those numbers are illustrative design choices. A ticket summary and a full conversation transcript have very different costs, so a record limit alone is insufficient. Bound the fields and text length returned per record too.

Make the assistant's stopping rule match the question

Suppose the fixture contains five matching tickets. With pages of two records, a full traversal needs three successful calls. After the first call, an assistant can truthfully say it found two tickets so far. It cannot claim that only two match.

Use different behavior for different requests:

Enforce a traversal budget in the integration: maximum calls, elapsed time, and accumulated response size. If a run reaches that budget, return a partial-result status to the assistant and have it explain the limitation. A suitable answer is: “I reviewed the first 100 matching tickets in creation order. More remain; this summary covers those 100.”

Do not erase a continuation token merely because your wrapper stops fetching. That would turn a local budget into a false end-of-data signal.

Define what happens while records change

Choose deterministic ordering, including a unique tie-breaker. For the example, the backend could order by creation timestamp and then ticket ID, even though the public sort option is simpler. Otherwise, records with identical timestamps leave an ambiguous boundary between pages.

A cursor alone does not promise a frozen dataset. A ticket can change from unresolved to resolved between calls. Decide whether the tool traverses a snapshot or reads the changing collection, then describe that behavior. A cutoff on creation time bounds newly created records; it does not freeze updates to existing ones.

For an exact report, consider a server-generated report or export with a defined snapshot time. For an interactive browse, a documented live view may be sufficient. Deduplicating IDs can remove repeats, but it cannot prove that no records were skipped.

Recheck access and handle broken continuation

Our recommendation is to bind continuation state to the authorized tenant and query, and reject attempts to reuse it with incompatible filters. Recheck record access on every page. A user who loses access between calls should not retain it through an old token. This follows the same separation described in connection authorization versus record access.

Define an explicit expired-token outcome. If recovery requires restarting, the assistant should know that earlier pages may no longer describe the same dataset. Do not silently restart at page one while presenting the output as the next page.

Also distinguish a failed page request from an empty successful page. If page three times out after two pages succeed, the result is incomplete. The assistant should preserve that distinction in its answer, as discussed in MCP tool errors and empty results.

Test completeness, not just the first response

Add pagination scenarios to your MCP evaluation set. Use a known fixture and check both returned records and the assistant's final wording:

These tests should exercise the actual backend boundaries as well as the tool wrapper. A fixture that always fits on one page cannot establish that continuation works.

Before exposing a collection to an assistant, write down three things: the end signal, the consistency promise, and what the user hears when traversal stops early. Those decisions make large result sets usable without disguising partial knowledge as a complete answer.

If you are choosing which workflows to expose first, the Agent Readiness Audit starts from your public API documentation and proposes tools. It can surface pagination questions to investigate; documentation alone cannot prove that deployed continuation, access checks, or complete-answer behavior work correctly.