# Custom pricing code

Edit your installed `provider.ts` to customize pricing or replace its `standingPrice` and `quote` callbacks. Start with the [generated policy](/providers/strategies), test your changes, then restart the connector.

Your callbacks decide prices and admission. The SDK supplies engine metrics and tracks Canopy activity.

For agent-assisted implementation, use the [Canopy provider pricing skill](https://docs.canopyx.ai/skills/canopy-provider/SKILL.md).

## Publish standing prices

`standingPrice(context)` returns your current terms. It can return a value or a promise. The SDK publishes the input and output rates after restoring a connection and every 30 seconds by default.

Standing prices describe your offer. Request selection uses valid, time-limited RFQ quotes. Cache terms belong to individual quotes; the standing-price wire message sends only the model, input rate, output rate, and currency.

## Accept or decline a request

`quote(context)` receives the same activity snapshot plus `rfq`, the request for quote. Return one of:

```ts
// Decline new work.
{ accept: false }

// Commit to serving it at these prices if selected before expiry.
{ accept: true, terms: { inputPerMTok: "0.35", outputPerMTok: "0.61" } }
```

The callback can return a promise, but it must finish before `rfq.deadline`. Throwing, returning invalid terms, or missing the deadline declines the request.

The RFQ includes a canonical model ID, token estimates, streaming and tool requirements, and structured-output requirements. It never includes prompt content. `endpoint` identifies the native protocol; its omission means Chat Completions. See the [RFQ fields](/providers/protocol#rfqrequest).

Only accept requests whose terms you can honor. Sending a quote makes a binding offer; `onAward` is a notification, not another admission decision.

## Read activity counts

Both policy callbacks receive:

| Field                | Meaning                                                                          |
| -------------------- | -------------------------------------------------------------------------------- |
| `activeRequests`     | Awarded Canopy requests until completion or release, including engine queue time |
| `idleCachedSessions` | Unexpired cache commitments with no active Canopy request                        |
| `pendingQuotes`      | Binding quotes awaiting award, release, or expiry                                |
| `metrics`            | Fresh engine metrics, or `undefined`                                             |
| `collectedAt`        | UTC timestamp for those metrics, or `undefined`                                  |

Do not add `activeRequests` to `metrics.runningRequests`. They can count the same work. Aggregate engine metrics cannot identify which Canopy requests are running rather than queued.

Pending quotes are not active requests, but they can become active if selected. Include them when deciding whether you can promise more work. [Starter policies](/providers/strategies#load-and-admission) use the larger of the active and engine-running counts, then add pending quotes and weighted idle cached sessions. This is an estimate, not a physical scheduling guarantee.

Metric fields depend on the adapter. Limited adapters report none of them; see [Engines and metrics](/providers/engines).

## Offer cached input

Add a cached input rate and retention terms to an accepted quote:

```ts
const terms = {
  inputPerMTok: "0.35",
  cachedInputPerMTok: "0.175",
  outputPerMTok: "0.61",
  cache: {
    ttlSeconds: 300,
    minCacheableTokens: 1024,
  },
};
```

A cache offer promises availability at these prices through its expiry. Successful cacheable requests refresh that expiry. Follow-ups can skip the RFQ, keep the original terms, and overlap under the same commitment.

New-quote limits do not stop those follow-ups. Account for possible cached traffic when deciding how much new work to accept. A cache commitment does not reserve a dedicated execution slot; your engine still owns scheduling and physical cache retention.

`idleCachedSessions` counts a commitment once, only while none of its requests are active. It is not a count of tokens or occupied GPU slots.

The wire format allows cache TTLs from 1 to 86,400 seconds. Schema validity alone does not guarantee that routing will treat an offer as useful caching. See [How routing works](/routing) for an overview of routing and [Lifecycle and caching](/providers/lifecycle) for retention obligations.

## Price format

Use decimal USD strings per million tokens. Prices allow one to six integer digits and up to six fractional digits. Zero is allowed; negative values, exponent notation, and JSON numbers are not.

`inputPerMTok` and `outputPerMTok` are required. `cachedInputPerMTok`, `cache`, `currency`, and `expiresAt` are optional. If supplied, `currency` must be `"USD"`. Use a UTC timestamp for `expiresAt` to bound how long the quote remains available for selection. This is separate from an awarded request's execution deadline.

For exact integer scaling, use [`scalePrice`](/providers/sdk-reference#scaleprice). For fractional
changes use [`multiplyPrice`](/providers/sdk-reference#price-interpolation). A `QuoteTerms` type annotation checks the object shape,
not the price string format. Validate returned terms with the public runtime schema when testing.

## Test a policy offline

Keep `provider.ts` limited to configuration and callbacks. Importing a default-exported
`defineProvider(...)` configuration does not connect to Canopy. Call its callbacks with synthetic
contexts to test admission and prices without making binding offers.

This example tests a vLLM `provider.ts`. Adapt the expected decisions and prices to your policy.

Save as `provider.test.ts` next to `provider.ts`:

```ts
import { expect, test } from "bun:test";
import { Schema } from "effect";
import { QuoteTerms } from "@canopyx/provider/protocol";
import provider from "./provider.ts";

type QuoteContext = Parameters<typeof provider.quote>[0];
const context: QuoteContext = {
  metrics: { runningRequests: 0, waitingRequests: 0, kvCacheUsageRatio: 0.1 },
  collectedAt: "2099-01-01T00:00:00.000Z",
  activeRequests: 0,
  idleCachedSessions: 0,
  pendingQuotes: 0,
  rfq: {
    requestId: "policy-test",
    offeringId: provider.offeringId,
    canonicalModelId: "Qwen/Qwen3-32B",
    deadline: "2099-01-01T00:00:01.000Z",
    estimate: { inputTokens: 2048, reusablePrefixTokens: 1024, expectedOutputTokens: 256 },
    request: { streaming: true, hasTools: false, hasStructuredOutput: false },
  },
};

const validateTerms = Schema.decodeUnknownSync(QuoteTerms);

test("an idle engine quotes valid terms", async () => {
  const decision = await provider.quote(context);
  expect(decision.accept).toBe(true);
  if (decision.accept) validateTerms(decision.terms);
});

test("standing prices rise with load", async () => {
  const idle = validateTerms(await provider.standingPrice(context));
  const busy = validateTerms(
    await provider.standingPrice({ ...context, activeRequests: 6, pendingQuotes: 1 }),
  );
  expect(Number(busy.outputPerMTok)).toBeGreaterThanOrEqual(Number(idle.outputPerMTok));
});

for (const overrides of [
  { metrics: undefined, collectedAt: undefined },
  { metrics: { ...context.metrics!, waitingRequests: 3 } },
  { activeRequests: 1_000 },
] satisfies ReadonlyArray<Partial<QuoteContext>>) {
  test(`declines ${JSON.stringify(overrides)}`, async () => {
    expect(await provider.quote({ ...context, ...overrides })).toEqual({ accept: false });
  });
}
```

Run `bun test ./provider.test.ts` and your project's type check. Use the Effect version installed
with the SDK for the runtime schema import. For time-dependent rules, control the clock and use
RFQ deadlines relative to it. To test one input precisely, call `evaluateStrategy` directly with
your spec and read its `trace`, which explains each step.

Add cases for every rule you implement, including threshold equality, overlapping activity counts,
maximum prices, request restrictions, and cache pressure. Test repeated samples and counter resets
if your code keeps history. Check standing and binding prices separately so an advertisement change
cannot hide unchanged RFQ prices.

`canopy-provider check` is an engine diagnostic. It does not execute pricing callbacks or validate
the saved Canopy mapping. Offline callback tests also cannot prove physical cache retention or
prevent cached follow-ups from bypassing admission.
