Modelplane Modelplane docs

Serving Stack Requirements

This document is for an unreleased version of Modelplane.

This document applies to the Modelplane main branch and not to the latest release v0.4.

An InferenceCluster with spec.cluster.existing.components: Provided installs no serving stack components. The cluster provides the whole substrate itself, and Modelplane only verifies it and composes its own configuration on top. This page lists what that cluster must provide, per component Modelplane would otherwise install. It applies on top of the general requirements for an existing cluster .

Modelplane checks the checked entries continuously and reports what is missing through the RequirementsMet condition on the InferenceCluster. A check verifies presence and served API versions, not the installed release: the not checked entries, including component versions, are yours to meet. The version listed per component is the one Modelplane installs in Managed mode and tests against; stay close to it.

On every stack

cert-manager

Managed mode installs chart cert-manager v1.20.2.

Checked:

  • CRD certificates.cert-manager.io serving v1

Not checked:

  • The cert-manager controller and webhook are running and issue Certificates.

kube-prometheus-stack

Managed mode installs chart kube-prometheus-stack 84.4.0.

Checked:

  • CRD podmonitors.monitoring.coreos.com serving v1
  • CRD servicemonitors.monitoring.coreos.com serving v1

Not checked:

  • Prometheus discovers PodMonitor objects in every namespace: with the chart, set podMonitorSelectorNilUsesHelmValues to false and podMonitorNamespaceSelector to empty. Modelplane’s scrape targets are PodMonitor objects in workload namespaces, and the chart’s default release label selector never matches them.
  • A scrape job for the Envoy Gateway proxy pods’ stats endpoint, if you want request metrics at the proxy level.

Not needed:

  • Modelplane disables Grafana and Alertmanager.

node-feature-discovery

Managed mode installs chart node-feature-discovery 0.19.0.

Checked:

  • CRD nodefeatures.nfd.k8s-sigs.io serving v1alpha1

Not checked:

  • The NFD worker runs on the GPU nodes and labels them with feature.node.kubernetes.io/pci-10de and friends. If GPU nodes are tainted, the worker must tolerate the taint, or the DRA driver’s kubelet plugin never schedules there and every GPU ResourceClaim stays pending with all components looking healthy.

nvidia-dra-driver-gpu

Managed mode installs chart dra-driver-nvidia-gpu 0.4.1.

Checked:

  • DeviceClass gpu.nvidia.com (resource.k8s.io/v1)

Not checked:

  • The NVIDIA kernel driver and Container Toolkit on every GPU node, from the node image or the GPU Operator. The requirements for an existing cluster state the versions.
  • The driver’s kubelet plugin publishes each GPU node’s devices as ResourceSlices.
  • No device plugin advertising nvidia.com/gpu: a second allocator would hand out the same GPUs behind DRA’s back.

Not needed:

  • Modelplane disables ComputeDomains (multi-node NVLink) and their prerequisites.

envoy-gateway

Managed mode installs chart gateway-helm v1.8.4.

Checked:

  • CRD gatewayclasses.gateway.networking.k8s.io serving v1
  • CRD gateways.gateway.networking.k8s.io serving v1
  • CRD httproutes.gateway.networking.k8s.io serving v1
  • CRD envoyproxies.gateway.envoyproxy.io serving v1alpha1
  • CRD backends.gateway.envoyproxy.io serving v1alpha1
  • CRD clienttrafficpolicies.gateway.envoyproxy.io serving v1alpha1

Not checked:

  • The Envoy Gateway controller runs with the extensionManager wired exactly as the values above: external processing delegated to the AI Gateway controller’s Service, with the Backend API enabled and InferencePool declared a backend resource. Without it, HTTPRoute to InferencePool backendRefs never route, with every component looking healthy.

The exact values Modelplane installs the chart with. The extensionManager wiring is the one coupling no check can verify. Without it, routes to an InferencePool never route while every component looks healthy:

yaml
config:
  envoyGateway:
    extensionApis:
      enableBackend: true
    extensionManager:
      hooks:
        xdsTranslator:
          translation:
            listener:
              includeAll: true
            route:
              includeAll: true
            cluster:
              includeAll: true
            secret:
              includeAll: true
          post:
          - Translation
          - Cluster
          - Route
      service:
        fqdn:
          hostname: ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local
          port: 1063
      backendResources:
      - group: inference.networking.k8s.io
        kind: InferencePool
        version: v1

ai-gateway-crds

Managed mode installs chart ai-gateway-crds-helm v1.1.0.

Checked:

  • CRD aigatewayroutes.aigateway.envoyproxy.io serving v1alpha1

Not checked:

  • The AI Gateway APIs are v1alpha1 and move with the controller. A provided install tracks the pinned v1.1.0 release.

ai-gateway

Managed mode installs chart ai-gateway-helm v1.1.0.

Not checked:

  • The AI Gateway controller is reachable at ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local:1063, the address Modelplane’s Envoy Gateway extensionManager values point at. A controller installed elsewhere never receives the external processing traffic.

gaie-crds

Managed mode applies manifests Modelplane vendors from the upstream release.

Checked:

  • CRD inferencepools.inference.networking.k8s.io serving v1

trust-manager

Managed mode installs chart trust-manager v0.25.0.

Checked:

  • CRD bundles.trust.cert-manager.io serving v1alpha1

Not checked:

  • trust-manager watches modelplane-system as its trust namespace, where the gateway PKI composes its Bundle. An install watching another namespace never syncs the Bundle, and the control plane can’t read the cluster CA.

dra-driver-critical-pods-quota

Managed mode applies manifests Modelplane vendors from the upstream release.

Not checked:

  • On clusters that restrict system-node-critical pods by namespace quota (GKE does), the DRA driver’s namespace needs a ResourceQuota admitting them, or the kubelet plugin DaemonSet never starts.

Standard stack

leader-worker-set

Managed mode installs chart lws v0.8.0.

Checked:

  • CRD leaderworkersets.leaderworkerset.x-k8s.io serving v1

Not checked:

  • The LeaderWorkerSet controller is running. Modelplane composes against the v0.8 line.

Dynamo stack

grove

Managed mode installs chart grove-charts v0.1.0-alpha.12-rc2.

Checked:

  • CRD podcliquesets.grove.io serving v1alpha1

Not checked:

  • Grove’s API is v1alpha1 with no compatibility promise between versions, so a provided install must run the exact pinned version. A CRD check can’t tell alpha revisions apart. Prefer Managed for the Dynamo stack until Grove stabilizes.

kai-scheduler

Managed mode installs chart kai-scheduler v0.16.8.

Checked:

  • CRD queues.scheduling.run.ai serving v2

Not checked:

  • The scheduler answers to the schedulerName value kai-scheduler, the name Modelplane’s engine pods request. Its Queue webhook must be serving. Modelplane composes against the v0.16 line.

What Modelplane still installs

Provided mode only skips the substrate. Modelplane still composes its own configuration and workloads: the modelplane-system namespace, the EnvoyProxy, GatewayClass and Gateway for the inference gateway, and on the Dynamo stack its KAI Queue hierarchy and the ModelExpress server with its CRDs. Their substrate dependencies gate on the checks above, so none of them is applied before the cluster serves the APIs they need.

During deletion, Modelplane can’t order its configuration ahead of a substrate it doesn’t own. Keep your controllers (the gateway controller, KAI) running while an InferenceCluster deletes, so they can process finalizers on Modelplane’s configuration.