# Serving Stack Requirements

What an existing cluster must provide to run the serving stack itself.

Source: /platform/serving-stack-requirements/

<!-- Generated by functions/compose-serving-stack/requirements_doc.py
     from the serving stack component lists. Do not edit; run
     `nix run .#requirements-doc` after changing the stack data. -->
An `InferenceCluster` with `spec.cluster.existing.components: Provided`
installs no serving stack components. The cluster provides the whole
substrate itself, and Modelplane only verifies it and composes its own
configuration on top. This page lists what that cluster must provide,
per component Modelplane would otherwise install. It applies on top of
the general [requirements for an existing
cluster]({{< ref "/platform/inference-cluster.md" >}}).

Modelplane checks the **checked** entries continuously and reports what
is missing through the `RequirementsMet` condition on the
`InferenceCluster`. A check verifies presence and served API versions,
not the installed release: the **not checked** entries, including
component versions, are yours to meet. The version listed per component
is the one Modelplane installs in Managed mode and tests against; stay
close to it.

## On every stack

### `cert-manager`

Managed mode installs chart `cert-manager` `v1.20.2`.

Checked:

- CRD `certificates.cert-manager.io` serving `v1`

Not checked:

- The cert-manager controller and webhook are running and issue Certificates.

### `kube-prometheus-stack`

Managed mode installs chart `kube-prometheus-stack` `84.4.0`.

Checked:

- CRD `podmonitors.monitoring.coreos.com` serving `v1`
- CRD `servicemonitors.monitoring.coreos.com` serving `v1`

Not checked:

- Prometheus discovers `PodMonitor` objects in every namespace: with the chart, set `podMonitorSelectorNilUsesHelmValues` to false and `podMonitorNamespaceSelector` to empty. Modelplane's scrape targets are `PodMonitor` objects in workload namespaces, and the chart's default release label selector never matches them.
- A scrape job for the Envoy Gateway proxy pods' stats endpoint, if you want request metrics at the proxy level.

Not needed:

- Modelplane disables Grafana and Alertmanager.

### `node-feature-discovery`

Managed mode installs chart `node-feature-discovery` `0.19.0`.

Checked:

- CRD `nodefeatures.nfd.k8s-sigs.io` serving `v1alpha1`

Not checked:

- The NFD worker runs on the GPU nodes and labels them with `feature.node.kubernetes.io/pci-10de` and friends. If GPU nodes are tainted, the worker must tolerate the taint, or the DRA driver's `kubelet` plugin never schedules there and every GPU ResourceClaim stays pending with all components looking healthy.

### `nvidia-dra-driver-gpu`

Managed mode installs chart `dra-driver-nvidia-gpu` `0.4.1`.

Checked:

- `DeviceClass` `gpu.nvidia.com` (`resource.k8s.io/v1`)

Not checked:

- The NVIDIA kernel driver and Container Toolkit on every GPU node, from the node image or the GPU Operator. The requirements for an existing cluster state the versions.
- The driver's `kubelet` plugin publishes each GPU node's devices as ResourceSlices.
- No device plugin advertising `nvidia.com/gpu`: a second allocator would hand out the same GPUs behind DRA's back.

Not needed:

- Modelplane disables `ComputeDomains` (multi-node NVLink) and their prerequisites.

### `envoy-gateway`

Managed mode installs chart `gateway-helm` `v1.8.4`.

Checked:

- CRD `gatewayclasses.gateway.networking.k8s.io` serving `v1`
- CRD `gateways.gateway.networking.k8s.io` serving `v1`
- CRD `httproutes.gateway.networking.k8s.io` serving `v1`
- CRD `envoyproxies.gateway.envoyproxy.io` serving `v1alpha1`
- CRD `backends.gateway.envoyproxy.io` serving `v1alpha1`
- CRD `clienttrafficpolicies.gateway.envoyproxy.io` serving `v1alpha1`

Not checked:

- The Envoy Gateway controller runs with the `extensionManager` wired exactly as the values above: external processing delegated to the AI Gateway controller's Service, with the Backend API enabled and InferencePool declared a backend resource. Without it, HTTPRoute to InferencePool `backendRefs` never route, with every component looking healthy.

The exact values Modelplane installs the chart with. The `extensionManager` wiring is the one coupling no check can verify. Without it, routes to an `InferencePool` never route while every component looks healthy:

```yaml
config:
  envoyGateway:
    extensionApis:
      enableBackend: true
    extensionManager:
      hooks:
        xdsTranslator:
          translation:
            listener:
              includeAll: true
            route:
              includeAll: true
            cluster:
              includeAll: true
            secret:
              includeAll: true
          post:
          - Translation
          - Cluster
          - Route
      service:
        fqdn:
          hostname: ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local
          port: 1063
      backendResources:
      - group: inference.networking.k8s.io
        kind: InferencePool
        version: v1
```

### `ai-gateway-crds`

Managed mode installs chart `ai-gateway-crds-helm` `v1.1.0`.

Checked:

- CRD `aigatewayroutes.aigateway.envoyproxy.io` serving `v1alpha1`

Not checked:

- The AI Gateway APIs are v1alpha1 and move with the controller. A provided install tracks the pinned v1.1.0 release.

### `ai-gateway`

Managed mode installs chart `ai-gateway-helm` `v1.1.0`.

Not checked:

- The AI Gateway controller is reachable at `ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local:1063`, the address Modelplane's Envoy Gateway `extensionManager` values point at. A controller installed elsewhere never receives the external processing traffic.

### `gaie-crds`

Managed mode applies manifests Modelplane vendors from the upstream release.

Checked:

- CRD `inferencepools.inference.networking.k8s.io` serving `v1`

### `trust-manager`

Managed mode installs chart `trust-manager` `v0.25.0`.

Checked:

- CRD `bundles.trust.cert-manager.io` serving `v1alpha1`

Not checked:

- trust-manager watches modelplane-system as its trust namespace, where the gateway PKI composes its Bundle. An install watching another namespace never syncs the Bundle, and the control plane can't read the cluster CA.

### `dra-driver-critical-pods-quota`

Managed mode applies manifests Modelplane vendors from the upstream release.

Not checked:

- On clusters that restrict `system-node-critical` pods by namespace quota (GKE does), the DRA driver's namespace needs a `ResourceQuota` admitting them, or the `kubelet` plugin `DaemonSet` never starts.

## Standard stack

### `leader-worker-set`

Managed mode installs chart `lws` `v0.8.0`.

Checked:

- CRD `leaderworkersets.leaderworkerset.x-k8s.io` serving `v1`

Not checked:

- The LeaderWorkerSet controller is running. Modelplane composes against the v0.8 line.

## Dynamo stack

### `grove`

Managed mode installs chart `grove-charts` `v0.1.0-alpha.12-rc2`.

Checked:

- CRD `podcliquesets.grove.io` serving `v1alpha1`

Not checked:

- Grove's API is v1alpha1 with no compatibility promise between versions, so a provided install must run the exact pinned version. A CRD check can't tell alpha revisions apart. Prefer Managed for the Dynamo stack until Grove stabilizes.

### `kai-scheduler`

Managed mode installs chart `kai-scheduler` `v0.16.8`.

Checked:

- CRD `queues.scheduling.run.ai` serving `v2`

Not checked:

- The scheduler answers to the `schedulerName` value `kai-scheduler`, the name Modelplane's engine pods request. Its Queue webhook must be serving. Modelplane composes against the v0.16 line.

## What Modelplane still installs

Provided mode only skips the substrate. Modelplane still composes its
own configuration and workloads: the `modelplane-system` namespace, the
`EnvoyProxy`, `GatewayClass` and `Gateway` for the inference gateway,
and on the Dynamo stack its KAI `Queue` hierarchy and the ModelExpress
server with its CRDs. Their substrate dependencies gate on the checks
above, so none of them is applied before the cluster serves the APIs
they need.

During deletion, Modelplane can't order its configuration ahead of a
substrate it doesn't own. Keep your controllers (the gateway
controller, KAI) running while an `InferenceCluster` deletes, so they
can process finalizers on Modelplane's configuration.
