# How Canopy serves a request

Canopy connects your application to model providers. It finds a provider that can serve the
requested model while your application keeps using the same API.

## The request journey

1. **The client sends a familiar request.** The application uses the normal OpenAI Chat Completions
   format and places a canonical model ID or a preset ID in `model`.
2. **Canopy resolves the model.** A canonical model ID uses default routing settings. An optional
   organization preset supplies its model and saved constraints.
3. **Providers are checked for eligibility.** A provider must be approved, active, connected, and
   configured for the requested endpoint and model.
4. **Eligible providers receive an RFQ.** The RFQ includes the information needed to estimate the
   work, such as model, size estimates, and request capabilities. It does not include prompt
   content.
5. **Providers quote or decline.** Quotes are valid only for a short period. Canopy ignores late,
   expired, duplicate, or mismatched quotes.
6. **A route is selected.** Canopy considers price, recent provider failures, and generation speed.
   Provider eligibility and any preset constraints still come first.
7. **Only the selected provider receives the request.** Canopy sends the complete request with the
   provider's upstream model ID and streams the response back to the client.
8. **The route can become reusable.** A successful request may create or refresh a time-limited
   commitment for a matching prompt prefix.

The user does not need to know which provider won. The provider does not receive work until its
quote is selected.

## Why prefix reuse matters

Long-running conversations often repeat the same system instructions, history, tools, or other
prefix content. A provider may offer a lower price for input it can reuse from its cache. Canopy can
remember that commercial commitment and use it for later requests with the same beginning.

Prefix reuse can provide:

* less repeated work for the provider;
* more stable pricing during a conversation; and
* a better chance of keeping a warm provider cache.

Canopy uses privacy-preserving matching for this purpose and does not retain raw prompt text.
Reuse stays within the request's organization and model. Each preset has its own cache scope; direct model requests use a separate scope. Applications can use the normal
request fields that identify a user or cacheable session to narrow reuse further.

Cache discounts can influence routing when the offered terms suit the request. Offering a cached
input rate does not guarantee selection or cache reuse.

A commitment is usable only while its terms, cache lifetime, and provider eligibility remain valid.
A changed preset or prompt prefix starts a new route decision when the existing commitment no longer
matches.

## How quotes are compared

Canopy considers the expected cost of serving the request alongside recent provider failures and
observed generation speed. Cache offers can also affect the expected cost of continuing a conversation.
The cheapest quote does not always win.

Canopy prices each valid quote as the cost of this request plus one follow-up. It keeps quotes within
25% of the cheapest, then prefers the provider with fewer recent failures, then the one least
below the model's average generation speed, and finally the lower price. A follow-up uses the cached input rate only when
the quote promises at least 300 seconds of retention, the request's reusable prefix meets the
quote's minimum, and the cached rate is at least 10% below the input rate. The selected provider
is paid the price it quoted.

Routing estimates help compare offers. Final usage charges depend on the accepted prices and reported
usage. Standing prices advertise a provider's rates; each request is served under its accepted quote
or an existing cache commitment.

An eligible existing cache commitment can serve a matching request before a new auction, at its
retained terms. See [Lifecycle and caching](/providers/lifecycle) before offering cache prices.

## What users should expect

* Request a canonical model directly, or use a preset for saved routing settings.
* The full request goes to one selected provider, not to every provider that quotes.
* Successful provider responses keep the expected OpenAI-compatible response format.
* A provider failure before a response is selected returns an error in the normal API flow.
* A provider failure after streaming starts ends that stream. Canopy does not silently switch
  providers halfway through a response.
* Inference requests reserve workspace credits before dispatch. Settlement deducts the final usage
  charge and releases unused reserved credits. Requests return HTTP 402 when available credits cannot
  cover the reservation.
* Provider payouts and automatic provider failover are not yet available.

See [getting started with Canopy](/consumers) for request examples.

## What providers should expect

Providers maintain an outbound WebSocket connection and answer RFQs before their deadlines. They can
change standing prices and decide which model capabilities and workloads to quote.

A request for quote, or RFQ, includes the details needed to estimate cost and check model support.
It does not include the user's prompt. The selected provider then receives the complete request at its configured
endpoint.

See the [provider overview](/providers), [SDK getting started guide](/providers/getting-started),
or [WebSocket protocol reference](/providers/protocol).

## When a new decision is made

Canopy starts a new auction when there is no usable commitment for the request. This happens when:

* no commitment matches the prompt prefix;
* the commitment or cache lifetime has expired;
* the preset or model has changed; or
* the committed provider is no longer eligible.

A connection failure invalidates that provider's active commitments. Client cancellation does not
invalidate a provider, but it also does not create a new commitment.
