# Use Canopy with your LLM

Use Canopy from your existing LLM application by setting its base URL, API key, and model.
You can request a canonical model directly. No preset setup is required.

If you use LiteLLM, the [LiteLLM plugin](/litellm) lets you configure Canopy routing from your
proxy's dashboard while keeping your existing endpoint and credentials.

## Set up your application

1. Sign in to the [Canopy web app](https://app.canopyx.ai) and create an API key for your organization.
2. Set the API base URL to `https://api.canopyx.ai/v1`.
3. Choose a canonical model ID, such as `openai/gpt-oss-20b`.

```text
Base URL:  https://api.canopyx.ai/v1
API key:   <your Canopy API key>
Model:     openai/gpt-oss-20b
```

Keep the key in your application's secret settings. The application may call these settings
endpoint, API base URL, token, or model. Override the base URL if you use a local or self-hosted Canopy.

To discover model IDs, use your client's model picker or request the model list:

```sh
$ curl https://api.canopyx.ai/v1/models \
  -H "Authorization: Bearer $CANOPY_API_KEY"
```

This lists active canonical models and your organization's presets. A listing does not guarantee
that a provider is currently connected and able to serve a request.

## Send a request

```sh
$ curl https://api.canopyx.ai/v1/chat/completions \
  -H "Authorization: Bearer $CANOPY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-oss-20b","messages":[{"role":"user","content":"Hello"}],"stream":true}'
```

Canopy supports streaming, tools, and conversation history. Clients using Responses or Anthropic
Messages can use the corresponding `/v1/responses` or `/v1/messages` endpoint.

## Default routing and optional presets

Direct model requests use normal immediate routing, with no extra price ceiling, minimum speed,
time-to-first-token limit, or required cache duration. Canopy checks the request's protocol,
context size, tools, and other capabilities, then considers price, recent provider failures, and
generation speed when selecting an offer. Continuing conversations can reuse an eligible provider cache commitment.

You can also use an organization's preset ID, such as `preset/coding`, as `model` to apply its
saved model and routing settings. Copy the preset ID from the web app or select it from the model list.

If your Canopy operator enables the OpenRouter backup, text Chat Completions can use the same model
when local routing cannot produce a valid winner. Backup requests debit the actual reported
OpenRouter cost from your workspace credits. Other protocols and paid media or search features
still require a local provider. Canopy never switches providers after a response starts.

Fund your workspace on the Billing page before using inference. Requests reserve credits for the
selected model's maximum cost, then charge the actual reported input, cached input, and output at
the accepted provider prices. Unused reserved credits return to your balance. HTTP 402 means the
workspace needs more credits. Provider payouts remain outside the current user-facing flow.

## If something goes wrong

* **Unknown model or preset:** check `/v1/models` for available IDs.
* **No provider is available:** try again later. Providers must be connected and offer a valid quote.
* **Your key is rejected:** create or request a new API key.
* **A response stops while streaming:** the selected provider may have failed. Canopy does not
  silently switch providers halfway through a response.

When contacting support, include the `x-canopy-request-id` response header if your application
shows it. Do not share your API key or prompt.

You can add workspace credits and view payments from Billing in the web app. Provider payouts are
not yet available.

If you host AI models, [become a model provider](/providers).

## Model precision

A direct model ID selects the publisher's original weights. To use a quantized offering, create a
preset and choose its exact precision under **Model and precision**. Canopy shows how many providers
offer each choice and routes only to matching offerings. We do not auto substitute quantized
weights for an original-model request.
