# Pricing strategies

The installer writes the policy to `~/.canopy-provider/app/provider.ts`. Review it against your engine and workload before going live. To change live pricing, edit that file and restart the connector. Saving web settings affects future installation, not an existing file or running connector.

The generated policy counts concurrent requests. It treats a short prompt and a long generation as equal work. Choose a limit you have tested; engine capacity metadata is not a workload benchmark.

## Prices and curves

Prices are decimal USD per million tokens. Output and cached input follow input by a ratio unless you unlink them. Defaults come from the model's OpenRouter median rates, or example rates when OpenRouter has none. Cached input starts at an 80% discount. Prices never fall as load rises, and cached input never costs more than fresh input.

| Preset        | Shape                                                         |
| ------------- | ------------------------------------------------------------- |
| Fill capacity | Starting price until 25% load, target at 65%, then busy price |
| Steady        | Target at 50%, then busy price                                |
| Hold price    | Target at 15%, held until 90%, then busy price                |

## Load and admission

The starter policy estimates reserved work as:

```text
max(Canopy active requests, engine running requests)
+ pending quotes
+ weighted idle cached sessions
```

It includes the next request when computing price and admission. Load is this total divided by your concurrent-request limit, not GPU utilisation or KV-cache occupancy.

The larger running count avoids counting the same request twice. Pending quotes can become active if selected. Idle cache commitments can receive follow-ups without a new quote.

The starter declines when metrics are missing or stale, counts are invalid, the engine reports queued requests, or the next request exceeds its limit. The stop-quoting threshold defaults to 100%; lower it to reserve headroom. A threshold of zero declines all new quotes.

A connector tracks commitments for one offering. If several connectors share an engine, their pending quotes are not a shared reservation pool. Customize admission for that arrangement.

:::tip[Did you know?]
Declining a quote does not count as an inference failure. Only accept work you can serve.
:::

## Cache offers

Cache discounts also promise retention. Starter settings enable cache pricing with 300 seconds of retention and a minimum reusable prefix of 1,024 tokens. Each idle cached session reserves 33% of a request. Disable cache pricing if your engine cannot retain the prefix for the offered duration. KV or slot occupancy can limit cache offers; it does not set the starter's request price.

A cached follow-up keeps its accepted terms and may bypass new-quote limits. Existing commitments still reserve capacity when you disable new cache offers. See [Lifecycle and caching](/providers/lifecycle) before changing retention or reservation rules.

## Review and customize the code

The generated `provider.ts` contains the load calculation, price curves, linked ratios, admission rules and pricing callbacks. It uses SDK helpers for exact price arithmetic.

For workload-specific limits, time-of-day rules or other metrics, edit the callbacks directly. Test them without connecting, as described in [Custom pricing code](/providers/policies).

## Preview evaluator

`pricesAtLoad({ strategy, utilisationBps })` returns configured rates for the chart. It does not decide admission or cache eligibility. The pure `evaluateStrategy` function from `@canopyx/provider/strategy` evaluates a request. After local edits, the chart still represents the web draft, not your installed code.

```ts
evaluateStrategy({ strategy, context, rfq, now, utilisation });
// { decision, declineReason, standing, utilisationBps, loadBps, trace }
```

`rfq` and `utilisation` are optional. `now` is epoch milliseconds for quote expiry. A supplied `utilisation` overrides pricing load only; admission uses actual counts. Fractions use basis points: 10,000 means 100%.

Rates interpolate between adjacent points in integer microdollars, rounded up. Linked ratios apply to the rounded input rate. A declined request has no quote. While declining for capacity or unavailable metrics, the connector advertises its highest rates.

### Spec reference

`StrategySpec` is the validated input to the generator and preview.

| Field                                   | Contract                                                                                            |
| --------------------------------------- | --------------------------------------------------------------------------------------------------- |
| `schemaVersion`                         | `1`                                                                                                 |
| `admission.maxConcurrent`               | Integer from 1 to 100,000                                                                           |
| `admission.stopQuotingAtUtilisationBps` | Optional, 0 to 10,000; zero stops all new quotes                                                    |
| `admission.declineWhenQueued`           | Reject work while the engine reports a queue                                                        |
| `admission.idleCachedSessionWeightBps`  | Reserved fraction per idle cached session, 0 to 10,000                                              |
| `admission.maxInputTokens`              | Optional RFQ input-token limit                                                                      |
| `prices.input`                          | 2–32 `{ utilisationBps, price }` points spanning 0–10,000; loads increase and prices never decrease |
| `prices.output`                         | `{ mode: "linked", multiplierBps }` or `{ mode: "curve", points }`                                  |
| `prices.cachedInput`                    | The same rule types as output; required with cache terms and never above fresh input                |
| `cache`                                 | Optional `{ ttlSeconds, minCacheableTokens, maxKvUsageBps }`                                        |
| `quoteExpirySeconds`                    | Optional, 1–3,600; omit for the standard expiry                                                     |

Output ratios range from 0 to 1,000,000 basis points, or 100× input. Linked output rates must fit the wire price limit. Cached-input ratios cannot exceed 10,000 basis points. Independent curves follow the same point rules as input.

## How quotes win

Canopy pays the price you quote. See [How quotes are compared](/routing#how-quotes-are-compared) for how it ranks quotes.
