Skip to main content

Kubernetes

The Kubernetes provider maps one durable replica intent to either a Deployment plus Service, or one standard KServe InferenceService. Kubernetes owns scheduling and Pods. InferCrane owns revision lifecycle, rollout policy, routing membership, evidence, and deletion of its exact labeled resources.
The adapter has hermetic and Kind lifecycle qualification. Real GPU, vLLM, SGLang, and custom OCI compatibility remains deferred to consolidated manual qualification. Registration is not a real-cluster support claim.

Choose the workload API

InferCrane does not create child Pods directly. It does not install KServe, Gateway API, a GPU device plugin, model storage, or a Kubernetes distribution.

Bootstrap a namespace

Review the manifests before applying them:
The included Role is namespace-scoped and contains no wildcard, Secret, Pod, cluster-wide, or RBAC mutation permission. It has read-only get/list access to PVC metadata for cache verification. If the control plane runs outside the cluster, bind the same Role to the user or identity in its kubeconfig instead of the included service account subject. Create the worker credential without putting it in Git or shell history. Its value must equal the control-plane worker credential used for health checks and routing:

Configure the control plane

Configuration is all-or-nothing. The image must be immutable. The adapter always passes an explicit context and namespace to kubectl; it never relies on or changes the current context.

Pre-pull the immutable runtime image

The default image policy is prefer, which renders imagePullPolicy: IfNotPresent. For a latency-critical pool, first use an operator-owned DaemonSet to pull the exact digest onto every eligible GPU node. A minimal container template is:
Wait until the DaemonSet is ready on the complete intended node pool, then set INFERCRANE_KUBERNETES_IMAGE_CACHE_POLICY=required. InferCrane renders imagePullPolicy: Never in that mode, so a missing prewarm fails immediately with ErrImageNeverPull instead of silently turning deploy into a registry transfer. InferCrane does not create this cluster-wide DaemonSet or change node taints; those remain explicit operator actions. Digest changes require rolling the pre-pull DaemonSet before deploying the new workload revision.

Reuse an immutable model cache

For clusters with a shared filesystem or CSI driver that supports multi-reader claims, an administrator can prepare one PVC containing the Hugging Face cache tree. InferCrane adopts the claim; it does not create storage or download model bytes through the control API. The claim must be Bound, support ReadOnlyMany or ReadWriteMany, and carry the exact model identity digest:
At launch, the provider verifies the claim before server-side apply, mounts it at /models read-only, and points the Hugging Face cache environment at /models/huggingface. A missing, unbound, single-node-only, or wrongly annotated required claim fails before workload mutation. prefer falls back to the runtime’s ordinary immutable download only when no claim is mapped; a configured but invalid claim never silently falls back. Record the verified locality through the common artifact API:
This is a shared-volume path, not a node-local cache or image-streaming claim. CSI performance, mount propagation, model-read throughput, and real GPU startup remain cluster-specific evidence. Validate the Kubernetes API, optional KServe CRD, and every required namespaced permission without creating a workload:

Deploy

The GPU value maps to the configured product-label value. It is never silently substituted.
Provisioning uses strict server-side dry-run followed by server-side apply with the infercrane-provider-v1 field manager. InferCrane does not use --force-conflicts; ownership drift fails visibly. A lost apply response is safe to retry because resource names and ownership metadata derive from the durable replica key.

KServe Standard mode

Install and operate a compatible KServe release separately, apply deploy/kubernetes/kserve/provider-rbac.yaml, then set:
The provider verifies the InferenceService API and required permissions during doctor. KServe owns generated Deployments, Services, and Pods; InferCrane inventories and deletes only the parent InferenceService. Raw KServe condition metadata remains available through infercrane inspect. KServe LLMInferenceService and llm-d remain separate integration boundaries. Dynamo uses a distinct kubernetes-dynamo adapter with one explicit mutation owner; it is never enabled implicitly by workload_api.

Expose the logical endpoint

deploy/kubernetes/gateway-api/httproute.yaml is an optional placeholder. Replace the Gateway and hostname, then apply it only if your cluster already has Gateway API and a controller. It routes to the InferCrane gateway, not directly to revisions or workers - so safe rollout ownership stays singular.

Local conformance

The disposable Kind test proves strict apply, restart observation, lost-state repair, foreign field ownership rejection, idempotent deletion, and zero remaining run-owned resources. It schedules no GPU and sends no paid provider request.