# Engines and metrics

On each connection, the connector receives the saved engine model ID from Canopy, then validates engine health, model discovery, and metrics before it reports in or publishes prices. It polls health and metrics outside your quote callbacks, so pricing reads the most recent valid snapshot without a network request per quote. Metrics stay on your machine; Canopy never receives them.

## Support levels

| Level     | Engines                 | What the connector does                                                                           |
| --------- | ----------------------- | ------------------------------------------------------------------------------------------------- |
| Supported | vLLM, llama.cpp, SGLang | Checks health and models, reads load and KV or slot metrics, and discovers capacity               |
| Limited   | Ollama, MLX, TabbyAPI   | Checks health and models. There are no load metrics, so pricing uses Canopy's own activity counts |

Each adapter is tested against one pinned engine version, serving from a single engine.

## vLLM

```ts
import { engineApiKey, vllm } from "@canopyx/provider";

const engine = vllm({
  baseUrl: "http://127.0.0.1:8000",
  apiKey: engineApiKey(),
  labels: { engine: "0" },
});
```

The contract target is vLLM `0.17.1`, V1, with one selected model and engine.

| Metric              | Source                      | Meaning                                       |
| ------------------- | --------------------------- | --------------------------------------------- |
| `runningRequests`   | `vllm:num_requests_running` | Requests currently running                    |
| `waitingRequests`   | `vllm:num_requests_waiting` | Requests waiting in the engine queue          |
| `kvCacheUsageRatio` | `vllm:kv_cache_usage_perc`  | KV cache usage as a fraction from zero to one |

The adapter selects the `model_name` label from Canopy's engine model ID. Add `labels` to select exactly one engine when needed. Ambiguous samples fail validation. The SDK does not add worker percentages together or substitute zero for missing metrics.

On connection it reads the KV cache size from `vllm:cache_config_info` (GPU blocks × block size) and the version from `/version`, where exposed. vLLM does not publish its sequence limit. An engine maximum is a ceiling, not your strategy's admission target.

See the [vLLM 0.17.1 metrics documentation](https://docs.vllm.ai/en/v0.17.1/usage/metrics/).

## llama.cpp

```ts
import { engineApiKey, llamaCpp } from "@canopyx/provider";

const engine = llamaCpp({
  baseUrl: "http://127.0.0.1:8080",
  apiKey: engineApiKey(),
});
```

The contract target is llama.cpp `b10795`, started with `--metrics`.

| Metric                 | Source                            | Meaning                                               |
| ---------------------- | --------------------------------- | ----------------------------------------------------- |
| `runningRequests`      | `llamacpp:requests_processing`    | Requests currently processing                         |
| `waitingRequests`      | `llamacpp:requests_deferred`      | Requests waiting in the engine queue                  |
| `slotUsageRatio`       | Processing requests ÷ total slots | Slot occupancy from zero to one, when slots are known |
| `promptTokensTotal`    | `llamacpp:prompt_tokens_total`    | Raw cumulative prompt token count                     |
| `generatedTokensTotal` | `llamacpp:tokens_predicted_total` | Raw cumulative generated token count                  |

The connector reads `total_slots` (set by `--parallel`) and the build from `/props` on each connection, and pricing uses slot occupancy where a strategy would use KV usage. Pass `slots` if `/props` is unavailable.

For a router-mode server, which serves several models, pass `router: true`. The adapter then selects the model with `?model=` on `/metrics` and `/props`.

Counters can decrease after an engine restart. If you calculate rates, discard a sampling window when a counter decreases. The SDK does not derive rates.

See the [llama.cpp server documentation](https://github.com/ggml-org/llama.cpp/blob/b10795/tools/server/README.md) for the pinned contract.

## SGLang

```ts
import { engineApiKey, sglang } from "@canopyx/provider";

const engine = sglang({
  baseUrl: "http://127.0.0.1:30000",
  apiKey: engineApiKey(),
});
```

The contract target is SGLang `v0.5.21`, serving one model with metrics enabled (`--enable-metrics`).

| Metric              | Source                    | Meaning                                        |
| ------------------- | ------------------------- | ---------------------------------------------- |
| `runningRequests`   | `sglang:num_running_reqs` | Requests currently running                     |
| `waitingRequests`   | `sglang:num_queue_reqs`   | Requests waiting in the scheduler queue        |
| `kvCacheUsageRatio` | `sglang:token_usage`      | Share of the KV token pool in use, zero to one |

The adapter accepts both the `sglang:` and `sglang_` metric prefixes, but not both for the same model. It selects the `model_name` label from Canopy's engine model ID; add `labels` to narrow further. On connection it reads `max_running_requests`, the KV token pool and the version from `/server_info`, falling back to `/get_server_info`.

## Limited support: Ollama, MLX and TabbyAPI

```ts
import { mlx, ollama, tabbyApi } from "@canopyx/provider";

const engine = ollama({ baseUrl: "http://127.0.0.1:11434" });
```

| Adapter    | Target          | Health check   |
| ---------- | --------------- | -------------- |
| `ollama`   | Ollama `0.33.1` | `/api/version` |
| `mlx`      | `mlx_lm.server` | `/health`      |
| `tabbyApi` | TabbyAPI        | `/health`      |

These adapters check health and `/v1/models`, and supply fresh metrics that contain no engine load. A strategy therefore prices from Canopy's own counts of active requests, pending quotes and idle cached sessions, and never sees engine-wide load or KV usage. Work the engine serves outside Canopy is invisible to it, so set conservative admission limits.

## Endpoint configuration

Every adapter accepts these options:

| Option       | Description                                                                                           |
| ------------ | ----------------------------------------------------------------------------------------------------- |
| `baseUrl`    | Engine origin used for health, `/v1/models`, discovery and, by default, `/metrics`. Paths are ignored |
| `apiKey`     | Optional bearer credential for collection. `engineApiKey()` reads it without putting it in source     |
| `metricsUrl` | Optional override for the metrics scrape URL                                                          |

These are the connector's collection settings, on the machine beside the engine. The engine model ID, the public endpoint Canopy calls and its credentials are set on the **Setup** page. See [Credentials](/providers/installer#credentials) for where `engineApiKey()` reads the key.

Use a metrics endpoint for the engine you actually serve. A load-balanced scrape that combines unrelated workers does not meet the single-engine contract.

## Freshness and failures

| Provider option    | Default | Purpose                                  |
| ------------------ | ------- | ---------------------------------------- |
| `pollIntervalMs`   | `1000`  | Delay between collection cycles          |
| `requestTimeoutMs` | `3000`  | Collection and ordinary callback timeout |
| `maxMetricsAgeMs`  | `5000`  | Maximum age of a usable snapshot         |

A collection failure immediately invalidates the snapshot. When collection fails or the sample becomes stale, callbacks receive both `metrics: undefined` and `collectedAt: undefined`, and a Canopy strategy declines. **Overview** shows which check failed.

Startup requires successful health, model discovery and, for supported engines, metrics validation.

Engine counts can include traffic outside Canopy and can overlap with `activeRequests`. See [Pricing strategies](/providers/strategies#load-and-admission) for how a strategy combines them.
