DeploymentSpec
New files declare the stable v1 file contract:main are resolved and persisted as immutable Hugging Face commits
before provisioning.
name identifies the deployment lifecycle. endpoint identifies the stable application-facing
route and defaults to name when omitted. A new deployment revision can therefore replace the
serving plan underneath support-production without changing application configuration.
SGLang uses a built-in immutable workload profile. Custom OCI images declare the full contract in
runtime.workload; see Custom OCI workloads. Mutable image tags and
underspecified probes are rejected before provisioning.
An optional serving object describes an advanced backend topology without exposing provider CRD
shape to the core schema. Today backend: dynamo requires one outer graph replica, Kubernetes, and
the kubernetes-dynamo adapter. See NVIDIA Dynamo for the simple command and
the complete reviewable YAML form.
compute.mode defaults to elastic; elastic replica bounds default to 1..1. Serverless requires
min_replicas: 0 and defaults max_replicas to one when omitted. Revisions are immutable; changing
a field creates a candidate. provider.adapter persists the exact implementation when more than one
profile can serve a cloud/runtime pair; it is optional for an unambiguous default. Current executable
elastic defaults include configured AWS EC2 and GCP Compute, while other provider-product profiles
remain registered but deferred until their exact combination is qualified. AWS and GCP require an
explicit provider.region. Cost is omitted unless a trustworthy provider measurement exists.
resources.gpu_count is the number of accelerators allocated to each runtime replica and defaults
to one. It is part of immutable intent, pricing, capacity history, benchmarks, recommendations, and
inference passports. SkyPilot, GCP Compute, and Kubernetes compile the exact count into provider
resources. AWS EC2 binds the count to INFERCRANE_AWS_GPU_COUNT for the configured instance type and
rejects a mismatch before creating paid capacity. Serverless currently supports one accelerator per
worker and rejects any other count explicitly.