# WebSocket protocol

Implement this protocol to connect a provider without the SDK. Version `1` carries prices, binding quotes, connector reports, session mode changes, and execution and cache lifecycle notifications. It does not carry inference request bodies or raw engine metrics.

You still need a provider account with an active offering, a connector credential, and a [public endpoint](/providers/reachability). Follow [Getting started](/providers/getting-started) up to installing the connector, and create a credential with **Create token** on the Install page. Before approval your sessions run in staging and receive no routing messages.

Runtime schemas are exported by `@canopyx/provider/protocol` if you want validation without the SDK client. This page describes the wire contract independently of TypeScript.

## Connect

Open a WebSocket upgrade request to the API:

```http
GET /providers/ws HTTP/1.1
Authorization: Bearer canopy_connector_replace_with_your_credential
X-Canopy-Offering: alice/gpt-oss-20b-original
Upgrade: websocket
Connection: Upgrade
```

Your WebSocket library supplies the remaining handshake headers. Use `wss://your-canopy-host/providers/ws` for a remote connection. Use a server-side client that can set the authorization header, and never expose the credential in browser code or logs.

The optional query parameter `cleanCache` accepts `true` or `false`. Omission means no reset. Use `?cleanCache=true` only after physical cache loss; it invalidates the provider's cache accounting before activating the connection.

The optional query parameter `diagnostic=true` opens a diagnostic session: Canopy sends `session.welcome` and closes with `1000`. It is never registered for routing or recorded. `canopy-provider doctor` uses it to test the connection.

| HTTP status | Meaning                                                                    |
| ----------- | -------------------------------------------------------------------------- |
| `400`       | Invalid `cleanCache` or `diagnostic` value                                 |
| `401`       | Missing, invalid or revoked credential, or a rejected or suspended account |
| `403`       | Offering is missing, inactive, or not owned by the provider                |
| `426`       | Request did not ask for a WebSocket upgrade                                |
| `500`       | Upgrade or internal failure                                                |
| `503`       | Server shutting down                                                       |

The server sends `session.welcome`, then `session.state`. Resolve and validate the upstream mapping from welcome, then restore state, then send a [`connector.report`](#connectorreport). In a live session, publish prices and process routing messages. The server restores its state before making the connection eligible for routing; there is no client restoration acknowledgement message.

### Staging and live sessions

The welcome's `mode` is `staging` before approval and `live` after. Canopy decides it from the provider's account status, never from the client. A staging session receives no RFQs, awards, releases or commitment updates, and Canopy rejects its `price.update` messages with `STAGING_SESSION`, so staging prices are never published. Send reports and answer heartbeats.

When the account is approved, Canopy sends [`session.mode`](#sessionmode) with `live` within about a minute, and the session becomes eligible for routing without a reconnect.

## Envelope

Each message is a JSON object:

```json
{
  "v": 1,
  "id": "msg_quote-1",
  "type": "rfq.quote",
  "replyTo": "msg_rfq-1",
  "ts": "2026-08-31T01:00:00.200Z",
  "payload": {
    "requestId": "rfq-1"
  }
}
```

| Field     | Contract                                                                                              |
| --------- | ----------------------------------------------------------------------------------------------------- |
| `v`       | Required integer, currently `1`                                                                       |
| `id`      | Required nonempty message ID, at most 64 characters                                                   |
| `type`    | Required nonempty message type, at most 64 characters                                                 |
| `replyTo` | Optional reply correlation ID, at most 64 characters; inbound `null` is allowed                       |
| `ts`      | Server UTC timestamp; optional on provider messages, at most 40 characters; inbound `null` is allowed |
| `payload` | Message-specific object; required by the price and quote handlers                                     |

Generate a distinct message ID for each message. Server acknowledgements and errors use `replyTo` to identify the provider message. RFQ matching uses `payload.requestId` and the authenticated provider connection, not `replyTo`.

Domain IDs, including RFQ, execution, commitment, and session IDs, are nonempty strings of at most 128 characters. Offering IDs in price updates allow letters, digits, `.`, `_`, `/`, and `-`.

Domain timestamps such as deadlines and expiries use UTC ISO strings ending in `Z`, with seconds and optional one to three fractional digits.

## Message types

| Direction          | Type                 | Purpose                                               |
| ------------------ | -------------------- | ----------------------------------------------------- |
| Provider to server | `price.update`       | Publish standing rates                                |
| Provider to server | `rfq.quote`          | Offer binding terms or decline                        |
| Provider to server | `connector.report`   | Report SDK version and engine checks                  |
| Provider to server | `ping`               | Request an application-level pong acknowledgement     |
| Server to provider | `session.welcome`    | Session identity, mode, models and heartbeat interval |
| Server to provider | `session.state`      | Restored commitments and active executions            |
| Server to provider | `session.mode`       | Switch a staging session to live                      |
| Server to provider | `rfq.request`        | Request a binding quote                               |
| Server to provider | `quote.award`        | Notify an execution award before HTTP dispatch        |
| Server to provider | `quote.release`      | Release an unawarded offer                            |
| Server to provider | `execution.release`  | Finish an execution                                   |
| Server to provider | `commitment.update`  | Maintain cache terms and expiry                       |
| Server to provider | `commitment.release` | Release a cache commitment                            |
| Server to provider | `ack`                | Confirm handling of a provider message                |
| Server to provider | `error`              | Report a rejected message                             |

The following examples show payloads unless explicitly described as full messages.

## `session.welcome`

```json
{
  "userId": "provider-1",
  "sessionId": "session-1",
  "heartbeatIntervalMs": 30000,
  "mode": "staging",
  "protocolVersions": [1],
  "models": [
    {
      "offeringId": "alice/gpt-oss-20b-original",
      "canonicalModelId": "openai/gpt-oss-20b",
      "upstreamModelId": "gpt-oss:20b",
      "status": "active"
    }
  ]
}
```

Read the heartbeat interval from the message rather than assuming the example value. The welcome contains only the offering selected by `X-Canopy-Offering`. Each entry has an `offeringId`, a `canonicalModelId`, an `upstreamModelId`, and an `active` or `inactive` status. The upstream ID is a nonempty string of at most 256 characters and is not subject to canonical-ID character restrictions. It comes from the offering saved on the Setup page and identifies the model sent to inference. Use it to validate the engine and select model-specific metrics on every connection.

## `session.state`

```json
{ "commitments": [], "executions": [] }
```

`commitments` contains [commitment objects](#commitmentupdate). `executions` contains [award objects](#quoteaward) for active work. Restore entries for your configured `offeringId` before handling new messages. Maintained commitment expiries take precedence over provisional commitment values embedded in awards.

## `connector.report`

Send after restoring state, and again whenever its contents change:

```json
{
  "offeringId": "alice/gpt-oss-20b-original",
  "sdkVersion": "0.2.0",
  "host": "gpu-box-1",
  "engine": {
    "kind": "vllm",
    "version": "0.17.1",
    "modelIds": ["gpt-oss:20b"],
    "checks": { "health": true, "models": true, "metrics": true },
    "capacity": { "kvCacheTokens": 524288 }
  }
}
```

`checks.metrics` is `null` for an engine without a metrics endpoint. `capacity` (`maxSequences`, `kvCacheTokens`, `slots`) is optional. Other strings are at most 128 characters. `modelIds` holds at most 200 entries of up to 256 characters each. The report carries no load metrics. The server acknowledges with `{ "stored": true }` and shows the report in the provider's console. In a staging session, the acknowledgement of the first report confirms the connection.

## `session.mode`

```json
{ "mode": "live" }
```

Sent once when an approved account's staging session becomes live. Start publishing prices and handling routing messages.

## `price.update`

Send a full message like:

```json
{
  "v": 1,
  "id": "msg_price-1",
  "type": "price.update",
  "payload": {
    "updates": [
      {
        "offeringId": "alice/gpt-oss-20b-original",
        "inputPerMTok": "0.35",
        "outputPerMTok": "0.61",
        "currency": "USD"
      }
    ]
  }
}
```

`updates` accepts 1 to 100 entries. Each requires an active offering owned by the provider and both rates. `currency` is optional and can only be `"USD"`. Duplicate offering entries in a batch are invalid.

The server acknowledges with `{ "applied": 1 }`, where `applied` is the number of updates. Standing rates describe the offer; routing selects from valid RFQ quotes. This message does not publish cache terms.

## `rfq.request`

```json
{
  "requestId": "rfq-1",
  "offeringId": "alice/gpt-oss-20b-original",
  "canonicalModelId": "openai/gpt-oss-20b",
  "deadline": "2026-08-31T01:00:01.100Z",
  "estimate": {
    "inputTokens": 2400,
    "reusablePrefixTokens": 1800,
    "expectedOutputTokens": 1024,
    "maximumOutputTokens": 2048
  },
  "request": {
    "streaming": true,
    "hasTools": false,
    "hasStructuredOutput": false
  }
}
```

All shown fields are required except `estimate.maximumOutputTokens`. Estimates are finite JSON numbers. The optional top-level `endpoint` is `chatCompletions`, `responses`, or `anthropicMessages`. The server omits it for Chat Completions. Quote only for a protocol and capabilities your configured endpoint supports.

The RFQ contains no prompt or request body. Respond before `deadline` on the same authenticated connection that received it.

## `rfq.quote`

```json
{
  "requestId": "rfq-1",
  "quote": {
    "inputPerMTok": "0.35",
    "cachedInputPerMTok": "0.175",
    "outputPerMTok": "0.61",
    "currency": "USD",
    "expiresAt": "2026-08-31T01:00:01.000Z",
    "cache": { "ttlSeconds": 300, "minCacheableTokens": 1024 }
  }
}
```

Omit `quote` to decline:

```json
{ "requestId": "rfq-1" }
```

Only one reply is accepted per RFQ and authenticated provider, and only from the connection that received it. The RFQ's `requestId` also identifies the pending quote, scoped to that provider.

The server acknowledges a resolved reply with `{ "accepted": true }`, including a decline. This means it handled the reply, not that the quote won. A closed, duplicate, late, or mismatched RFQ receives `UNKNOWN_RFQ`.

### Quote terms

| Field                      | Required                | Contract                        |
| -------------------------- | ----------------------- | ------------------------------- |
| `inputPerMTok`             | Yes                     | Fresh input price               |
| `outputPerMTok`            | Yes                     | Output price                    |
| `cachedInputPerMTok`       | No                      | Cached input price              |
| `currency`                 | No                      | Only `"USD"`                    |
| `expiresAt`                | No                      | UTC quote expiry                |
| `cache.ttlSeconds`         | When `cache` is present | Integer from 1 to 86,400        |
| `cache.minCacheableTokens` | No                      | Integer from 1 to 2,147,483,647 |

Prices are decimal USD strings per million tokens, matching `^\d{1,6}(\.\d{1,6})?$`. Do not send numeric prices or exponent notation. Schema-valid cache terms are not a guarantee of cache routing eligibility; see [How routing works](/routing).

Canopy validates every quote's prices and terms. An invalid quote is declined; accepted offers remain binding until released or expired.

A quote commits the provider to serving the request at those terms if selected before expiry. Track outstanding offers in your admission policy. Expire them locally in case a release is lost. The SDK uses five seconds after the RFQ deadline, bounded by any earlier explicit `expiresAt`, for its pending quote lifetime.

## `quote.release`

```json
{ "requestId": "rfq-1", "reason": "lost" }
```

`reason` is `lost`, `expired`, or `cancelled`. Remove the unawarded offer even if an award notification was lost. Canopy releases losing or invalid offers, replies that miss the RFQ deadline, and offers in cancelled auctions. A winning award replaces its pending quote with an active execution.

## `quote.award`

```json
{
  "executionId": "execution-1",
  "requestId": "rfq-1",
  "quoteId": "rfq-1",
  "offeringId": "alice/gpt-oss-20b-original",
  "canonicalModelId": "openai/gpt-oss-20b",
  "terms": { "inputPerMTok": "0.35", "outputPerMTok": "0.61" },
  "expiresAt": "2026-08-31T01:15:00.500Z",
  "cached": false
}
```

| Field              | Meaning                                                         |
| ------------------ | --------------------------------------------------------------- |
| `executionId`      | Unique execution attempt, used for release and HTTP correlation |
| `requestId`        | Consumer inference request ID                                   |
| `quoteId`          | Original RFQ ID for a new auction; absent for cached follow-ups |
| `canonicalModelId` | Model to execute                                                |
| `terms`            | Accepted quote terms                                            |
| `expiresAt`        | Hard execution deadline, fifteen minutes after issue            |
| `cached`           | Whether this request uses an existing cache commitment          |
| `commitment`       | Optional commitment object with ID, model, terms, and expiry    |

An award is a notification, not a request for acceptance. Record it once per execution ID. Duplicates must not increase the active count or repeat callbacks.

Canopy sends the award, then dispatches the HTTP request to your registered inference endpoint using
your upstream model name. It does not wait for an award acknowledgement.

The HTTP request carries `x-canopy-execution-id` for correlation. For a new auction, `x-canopy-request-id` matches the award's `quoteId`.

A failed award send lets the auction try another valid quote. A failed HTTP connection ends the selected execution. Disconnects, expired quotes, or failed dispatch can cause selection of another quote before an upstream request is sent. Requests already sent upstream are not automatically retried.

## `execution.release`

```json
{ "executionId": "execution-1", "reason": "completed" }
```

`reason` is `completed`, `cancelled`, `failed`, or `expired`. Remove the execution from active accounting. Stream completion, upstream rejection, connection failure, consumer cancellation, and the hard deadline all end executions.

Enforce the supplied deadline locally. It covers the complete HTTP request and response stream, not the time to answer an RFQ. Expiry bounds accounting but does not prove that physical engine work has stopped.

Releasing an execution does not release a maintained cache commitment.

## `commitment.update`

```json
{
  "commitmentId": "commitment-1",
  "offeringId": "alice/gpt-oss-20b-original",
  "canonicalModelId": "openai/gpt-oss-20b",
  "terms": {
    "inputPerMTok": "0.35",
    "cachedInputPerMTok": "0.175",
    "outputPerMTok": "0.61",
    "cache": { "ttlSeconds": 300, "minCacheableTokens": 1024 }
  },
  "expiresAt": "2026-08-31T01:05:10.000Z"
}
```

All four top-level fields are required. This object also appears in awards and restored session state.

An initial award reserves a possible cache commitment until execution expiry plus the cache TTL. After successful cache accounting, `commitment.update` supplies the maintained terms and actual expiry. Each successful cached request refreshes that expiry.

A cached route does not guarantee a physical cache hit. Report cache usage accurately and retain the
physical cache and capacity needed to honor the promise.

A cached follow-up skips the RFQ and receives an award with `cached: true`, the same commitment, and the original terms. Follow-ups can overlap without a per-commitment queue. Account for this potential work when quoting new requests.

Count a commitment as idle only while it is unexpired and has no active executions. A cache promise does not reserve a dedicated execution slot.

## `commitment.release`

```json
{ "commitmentId": "commitment-1", "reason": "expired" }
```

`reason` is `expired` or `invalidated`. Remove the commitment when released and enforce supplied
deadlines locally even if a release notification is delayed or lost.

Failed initial executions release provisional commitments. Provider cache invalidation releases maintained commitments.

## Heartbeats

Respond to WebSocket protocol-level ping frames with pong frames. Most server-side WebSocket libraries do this automatically. The server uses `heartbeatIntervalMs` from the welcome message and terminates connections after two unanswered pings when the next heartbeat check runs.

An optional application message `{ "v": 1, "id": "msg_ping-1", "type": "ping", "payload": {} }` receives an `ack` with `{ "pong": true }`. This acknowledgement does not reset the server's protocol-level missed-pong counter. Do not use it as a substitute for WebSocket pong frames.

## Errors and limits

An error payload has `code`, a human-readable `message`, and optional `details`. Its envelope includes `replyTo` when the server can identify the rejected message.

| Code                  | Meaning                                             |
| --------------------- | --------------------------------------------------- |
| `MALFORMED_MESSAGE`   | Invalid JSON or envelope                            |
| `UNSUPPORTED_VERSION` | `v` is not supported                                |
| `UNSUPPORTED_TYPE`    | Message type has no provider handler                |
| `VALIDATION_FAILED`   | Invalid payload or duplicate model in a price batch |
| `UNKNOWN_MODEL`       | Price update names a model outside the session      |
| `MODEL_INACTIVE`      | Price update names an inactive model                |
| `UNKNOWN_RFQ`         | No open RFQ matches this reply and connection       |
| `STAGING_SESSION`     | A staging session sent `price.update`               |
| `RATE_LIMITED`        | Inbound message queue is full                       |
| `INTERNAL`            | Server failed to handle the operation               |

Messages have a 64 KiB payload limit. Sending messages too quickly can return `RATE_LIMITED` and close
the connection with `1013`. Repeated malformed JSON or envelope messages can close it with `1008`.

State restoration failure closes with `1011`. Shutdown uses `1001`. A revoked credential, or a rejected or suspended account, closes an existing connection with `1008` as soon as the change is made. A changed offering closes it with `1012`; reconnect to receive the current offering. Fix authorization failures instead of retrying indefinitely with an invalid credential.

## Reconnect and cache loss

Reconnect with backoff after transport failures. Restore the new `session.state` before applying subsequent messages. Ordinary reconnects preserve maintained cache commitments. If the engine lost its physical cache, request `cleanCache=true`; do not reset caches on every reconnect.

In-flight work remains bounded by its execution deadline. After a Canopy restart, restore the supplied
session state. Do not replay inference requests from earlier award notifications.

The bearer-credential protocol still needs stronger transport protections before use in an internet-facing marketplace. Use TLS, rotate credentials, and redact secrets, but do not treat those steps as message signing or replay protection.
