Serving Stack Requirements
This document applies to the Modelplane main branch and not to the latest release v0.4.
An InferenceCluster with spec.cluster.existing.components: Provided
installs no serving stack components. The cluster provides the whole
substrate itself, and Modelplane only verifies it and composes its own
configuration on top. This page lists what that cluster must provide,
per component Modelplane would otherwise install. It applies on top of
the general requirements for an existing
cluster
.
Modelplane checks the checked entries continuously and reports what
is missing through the RequirementsMet condition on the
InferenceCluster. A check verifies presence and served API versions,
not the installed release: the not checked entries, including
component versions, are yours to meet. The version listed per component
is the one Modelplane installs in Managed mode and tests against; stay
close to it.
On every stack
cert-manager
Managed mode installs chart cert-manager v1.20.2.
Checked:
- CRD
certificates.cert-manager.ioservingv1
Not checked:
- The cert-manager controller and webhook are running and issue Certificates.
kube-prometheus-stack
Managed mode installs chart kube-prometheus-stack 84.4.0.
Checked:
- CRD
podmonitors.monitoring.coreos.comservingv1 - CRD
servicemonitors.monitoring.coreos.comservingv1
Not checked:
- Prometheus discovers
PodMonitorobjects in every namespace: with the chart, setpodMonitorSelectorNilUsesHelmValuesto false andpodMonitorNamespaceSelectorto empty. Modelplane’s scrape targets arePodMonitorobjects in workload namespaces, and the chart’s default release label selector never matches them. - A scrape job for the Envoy Gateway proxy pods’ stats endpoint, if you want request metrics at the proxy level.
Not needed:
- Modelplane disables Grafana and Alertmanager.
node-feature-discovery
Managed mode installs chart node-feature-discovery 0.19.0.
Checked:
- CRD
nodefeatures.nfd.k8s-sigs.ioservingv1alpha1
Not checked:
- The NFD worker runs on the GPU nodes and labels them with
feature.node.kubernetes.io/pci-10deand friends. If GPU nodes are tainted, the worker must tolerate the taint, or the DRA driver’skubeletplugin never schedules there and every GPU ResourceClaim stays pending with all components looking healthy.
nvidia-dra-driver-gpu
Managed mode installs chart dra-driver-nvidia-gpu 0.4.1.
Checked:
DeviceClassgpu.nvidia.com(resource.k8s.io/v1)
Not checked:
- The NVIDIA kernel driver and Container Toolkit on every GPU node, from the node image or the GPU Operator. The requirements for an existing cluster state the versions.
- The driver’s
kubeletplugin publishes each GPU node’s devices as ResourceSlices. - No device plugin advertising
nvidia.com/gpu: a second allocator would hand out the same GPUs behind DRA’s back.
Not needed:
- Modelplane disables
ComputeDomains(multi-node NVLink) and their prerequisites.
envoy-gateway
Managed mode installs chart gateway-helm v1.8.4.
Checked:
- CRD
gatewayclasses.gateway.networking.k8s.ioservingv1 - CRD
gateways.gateway.networking.k8s.ioservingv1 - CRD
httproutes.gateway.networking.k8s.ioservingv1 - CRD
envoyproxies.gateway.envoyproxy.ioservingv1alpha1 - CRD
backends.gateway.envoyproxy.ioservingv1alpha1 - CRD
clienttrafficpolicies.gateway.envoyproxy.ioservingv1alpha1
Not checked:
- The Envoy Gateway controller runs with the
extensionManagerwired exactly as the values above: external processing delegated to the AI Gateway controller’s Service, with the Backend API enabled and InferencePool declared a backend resource. Without it, HTTPRoute to InferencePoolbackendRefsnever route, with every component looking healthy.
The exact values Modelplane installs the chart with. The extensionManager wiring is the one coupling no check can verify. Without it, routes to an InferencePool never route while every component looks healthy:
config:
envoyGateway:
extensionApis:
enableBackend: true
extensionManager:
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- Translation
- Cluster
- Route
service:
fqdn:
hostname: ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local
port: 1063
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1ai-gateway-crds
Managed mode installs chart ai-gateway-crds-helm v1.1.0.
Checked:
- CRD
aigatewayroutes.aigateway.envoyproxy.ioservingv1alpha1
Not checked:
- The AI Gateway APIs are v1alpha1 and move with the controller. A provided install tracks the pinned v1.1.0 release.
ai-gateway
Managed mode installs chart ai-gateway-helm v1.1.0.
Not checked:
- The AI Gateway controller is reachable at
ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local:1063, the address Modelplane’s Envoy GatewayextensionManagervalues point at. A controller installed elsewhere never receives the external processing traffic.
gaie-crds
Managed mode applies manifests Modelplane vendors from the upstream release.
Checked:
- CRD
inferencepools.inference.networking.k8s.ioservingv1
trust-manager
Managed mode installs chart trust-manager v0.25.0.
Checked:
- CRD
bundles.trust.cert-manager.ioservingv1alpha1
Not checked:
- trust-manager watches modelplane-system as its trust namespace, where the gateway PKI composes its Bundle. An install watching another namespace never syncs the Bundle, and the control plane can’t read the cluster CA.
dra-driver-critical-pods-quota
Managed mode applies manifests Modelplane vendors from the upstream release.
Not checked:
- On clusters that restrict
system-node-criticalpods by namespace quota (GKE does), the DRA driver’s namespace needs aResourceQuotaadmitting them, or thekubeletpluginDaemonSetnever starts.
Standard stack
leader-worker-set
Managed mode installs chart lws v0.8.0.
Checked:
- CRD
leaderworkersets.leaderworkerset.x-k8s.ioservingv1
Not checked:
- The LeaderWorkerSet controller is running. Modelplane composes against the v0.8 line.
Dynamo stack
grove
Managed mode installs chart grove-charts v0.1.0-alpha.12-rc2.
Checked:
- CRD
podcliquesets.grove.ioservingv1alpha1
Not checked:
- Grove’s API is v1alpha1 with no compatibility promise between versions, so a provided install must run the exact pinned version. A CRD check can’t tell alpha revisions apart. Prefer Managed for the Dynamo stack until Grove stabilizes.
kai-scheduler
Managed mode installs chart kai-scheduler v0.16.8.
Checked:
- CRD
queues.scheduling.run.aiservingv2
Not checked:
- The scheduler answers to the
schedulerNamevaluekai-scheduler, the name Modelplane’s engine pods request. Its Queue webhook must be serving. Modelplane composes against the v0.16 line.
What Modelplane still installs
Provided mode only skips the substrate. Modelplane still composes its
own configuration and workloads: the modelplane-system namespace, the
EnvoyProxy, GatewayClass and Gateway for the inference gateway,
and on the Dynamo stack its KAI Queue hierarchy and the ModelExpress
server with its CRDs. Their substrate dependencies gate on the checks
above, so none of them is applied before the cluster serves the APIs
they need.
During deletion, Modelplane can’t order its configuration ahead of a
substrate it doesn’t own. Keep your controllers (the gateway
controller, KAI) running while an InferenceCluster deletes, so they
can process finalizers on Modelplane’s configuration.