# Public endpoint

Canopy sends inference requests directly to your engine's API over HTTPS. The connector's WebSocket only carries prices and quotes; it does not tunnel requests. Your endpoint therefore needs:

* **HTTPS** with a certificate valid for its host name.
* **A public address.** Canopy refuses host names that resolve to loopback, private, link-local or carrier-grade NAT addresses, including Tailscale's `100.64.0.0/10`.
* **No redirects.** Canopy does not follow them, so serve `/v1/...` directly rather than redirecting to it.
* **Authentication.** A public endpoint is reachable by anyone. Require a key, and save it in Canopy as **Authentication** under **Connection setup** on **Setup**.
* **An OpenAI-compatible API** whose `/v1/models` lists your engine model ID, and whose responses return that ID as `model` and report token `usage`. Set the ID with `--served-model-name` for vLLM and SGLang, or `--alias` for llama.cpp.

Enter the base URL, including its API version, such as `https://inference.example.com/v1`, on the **Setup** page.

## Expose an engine with Caddy

If the machine has a public IP and ports 80 and 443 are open, [Caddy](https://caddyserver.com/) obtains and renews a certificate automatically. This configuration forwards authenticated `/v1/` requests to a local vLLM server and answers everything else, including `/metrics`, with 404:

```text
inference.example.com {
	@canopy {
		path /v1/*
		header Authorization "Bearer {$CANOPY_UPSTREAM_KEY}"
	}
	handle @canopy {
		reverse_proxy 127.0.0.1:8000 {
			flush_interval -1
		}
	}
	handle {
		respond 404
	}
}
```

Point a DNS record for `inference.example.com` at the machine, set `CANOPY_UPSTREAM_KEY` in Caddy's environment, and reload Caddy. In Canopy, choose **Bearer token** and save the same key. `flush_interval -1` streams responses as the engine writes them.

Change the upstream port for other engines: llama.cpp usually listens on `8080` and SGLang on `30000`. Keep the engine itself bound to `127.0.0.1` so only Caddy reaches it.

## Expose an engine with Cloudflare Tunnel

Behind NAT or a firewall, [Cloudflare Tunnel](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/) connects outbound to Cloudflare and serves your engine on a Cloudflare-managed certificate. You need a domain on Cloudflare.

```sh
cloudflared tunnel login
cloudflared tunnel create canopy-inference
cloudflared tunnel route dns canopy-inference inference.example.com
```

Then write `~/.cloudflared/config.yml`:

```yaml
tunnel: canopy-inference
credentials-file: /home/you/.cloudflared/TUNNEL-ID.json
ingress:
  - hostname: inference.example.com
    path: ^/v1/
    service: http://127.0.0.1:8000
  - service: http_status:404
```

Run it with `cloudflared tunnel run canopy-inference`, or install it as a service with `cloudflared service install`. The tunnel does not authenticate requests, so start the engine with an API key (`--api-key` for vLLM, SGLang and llama.cpp) and save that key in Canopy.

Cloudflare closes proxied requests that send no response bytes for 100 seconds. Streaming responses are unaffected once the first token arrives, but very long prompts on a slow engine can exceed that before the first token.

## What Canopy checks

Canopy checks your endpoint from its inference API, the same place real requests come from:

| Check        | What passes                                                                                                        |
| ------------ | ------------------------------------------------------------------------------------------------------------------ |
| Reachability | A `GET` of your models endpoint with your saved upstream authentication succeeds over verified TLS                 |
| Models       | The returned list includes each offering's engine model ID                                                         |
| Canary       | For Chat Completions offerings, a one-token completion returns your engine model ID as `model` and reports `usage` |

Canopy runs the checks when a connector connects, unless it checked in the last 15 minutes, when your account is approved, and when you choose **Check endpoint now** on **Overview**. Results feed the readiness checklist and `canopy-provider doctor`.

Canopy records the outcome, latency and error class, never request or response bodies. The canary costs your engine one output token.

### Error classes

| Error class        | Meaning and fix                                                                                        |
| ------------------ | ------------------------------------------------------------------------------------------------------ |
| `policy_rejected`  | The URL is not public HTTPS. Use an `https://` URL on a public host name                               |
| `dns_failed`       | The host name does not resolve. Check its DNS record                                                   |
| `private_address`  | The host name resolves to a private or loopback address. Expose it publicly, for example with a tunnel |
| `tls_error`        | The certificate is invalid, expired or for another name. Caddy and Cloudflare manage this for you      |
| `connect_failed`   | Nothing accepted the connection. Check the firewall, port and proxy                                    |
| `timeout`          | No complete response within the time limit                                                             |
| `redirect`         | The endpoint redirected. Serve the path directly                                                       |
| `auth_rejected`    | Your endpoint returned 401 or 403. Save the matching key under **Authentication**                      |
| `auth_unavailable` | Canopy could not read the saved upstream credential. Save it again                                     |
| `http_<code>`      | Your endpoint returned another HTTP error, such as `http_404` for a wrong models path                  |
| `invalid_response` | The response was not valid JSON in the expected shape, or was too large                                |
| `model_missing`    | `/v1/models` does not list the offering's engine model ID                                              |
| `model_mismatch`   | The completion's `model` differs from the engine model ID. Set the served model name or alias          |
| `usage_missing`    | The completion did not report `usage.prompt_tokens` and `usage.completion_tokens`                      |
