Provider setup
Provider integrations are registered control-plane adapters. Each provider owns its credentials, infrastructure semantics, and billable resources; InferCrane owns durable intent, reconciliation, and cleanup. A provider is supported only after its adapter combination appears in the release qualification matrix.Choose the least disruptive path
You do not need cloud-administrator access to evaluate InferCrane. Start with the smallest ownership boundary that proves value:
For BYOC, a cloud administrator prepares the boundary once. Day-to-day operators then use a
short-lived control-plane identity; deployment specifications never contain provider credentials.
Common onboarding sequence
The provider-specific pages contain the exact prerequisites and configuration fields. The safe sequence is the same across providers:- Choose a boundary. Select one account or project, region, private network, runtime image, and GPU profile. Do not start with account-wide administrator credentials.
- Create workload identity. Give the InferCrane control plane only the lifecycle permissions it needs. Give workers a separate identity that can read only their worker secret and required artifacts.
- Confirm quota and cost controls. Check GPU quota and regional availability. Configure a provider budget alert, while remembering that an alert does not stop spend.
- Configure InferCrane. Store provider settings outside deployment YAML and keep private files
mode
0600. - Run a read-only preflight.
infercrane doctor --aws,infercrane doctor --gcp, orinfercrane doctor --kubernetesvalidates identity and required API reads without provisioning a GPU. - Plan before mutation. Review
infercrane planoutput, the exact model revision, immutable runtime digest, GPU, region, and cache policy. - Deploy with a recovery handle. Use an idempotency key and retain the operation ID. Closing the terminal does not cancel the durable operation.
- Prove cleanup. Delete through InferCrane, check
infercrane orphans, and confirm provider inventory returns to the recorded baseline.
doctor proves configuration and identity reads. It cannot prove GPU stock, IAM propagation,
private routing, driver compatibility, model fit, or deletion semantics. Those boundaries are
qualified by the first guarded deployment and remain specific to the exact provider/runtime/model
combination.Data residency and production qualification
A region flag is placement intent, not a residency guarantee. InferCrane cannot prove where a provider stores control metadata, model artifacts, logs, backups, or failed requests unless the provider and customer configuration expose evidence for each boundary. Credentials must remain in the provider-specific workload identity or referenced secret path; region selection does not weaken that rule. Before approving a residency-constrained production path:SkyPilot provider breadth and launch preflight
SkyPilot is the portable capacity driver for clouds that do not need a native InferCrane lifecycle adapter. Provider registration is declarative and fail closed: a validated manifest appears in the planning catalog even before credentials exist, with stateconnection-required. This lets an
exact current provider quote inform a recommendation and return connection remediation without
starting provider processes. A cloud appears as executable only when its manifest is present and
every named credential environment variable exists in the control-plane process.
credential_env contains environment-variable names, never secret values. The underlying SkyPilot
installation must support the declared cloud and the process must receive the corresponding
credentials. A registered runtime is a routing capability, not real-infrastructure qualification;
the exact provider, region, accelerator, runtime, image, and model tuple must still pass the release
qualification workflow.
Before a billable launch, an authenticated client can compare the current catalog quote with
provider capacity evidence:
catalog_quoteis a timestamped catalog observation; it does not reserve inventory.launch_evidencereports connection, availability, quota, and deployability independently. Missing credentials and unsupported provider probes returnunknownorconnection-required, never an optimistic success.- Only an accepted provider launch reserves capacity. The resulting operation and qualification evidence remain the durable proof used by Release Guard.
unknown until
the actual launch or a provider-specific read-only probe supplies evidence.
RunPod
Set a scopedRUNPOD_API_KEY on the control plane. Native RunPod Pods are the default elastic
worker path and do not depend on a third-party catalog. Declare a SkyPilot provider manifest only
when intentionally using its long-tail execution path. Run infercrane doctor --cloud before
provisioning.
For Serverless, create a RunPod vLLM template with MODEL_NAME, immutable MODEL_REVISION, and RAW_OPENAI_OUTPUT=1, then set INFERCRANE_RUNPOD_SERVERLESS_TEMPLATE_ID. infercrane doctor --serverless reads and validates the template without creating an endpoint.
Set INFERCRANE_URL to a URL reachable from AIPerf and clients. Keep provider and InferCrane credentials out of specifications, logs, issue reports, and benchmark artifacts. Always inspect existing pods/endpoints before retrying manual acceptance and delete paid resources after the test.
AWS EC2 BYOC
AWS elastic support uses a separately registered EC2 adapter rather than provider conditionals in the lifecycle engine. It requires a complete role, private network, AMI, instance profile, worker secret, instance type/GPU, region, and immutable image configuration. See AWS EC2 BYOC and runinfercrane doctor --aws before provisioning.
ASG, EKS, SageMaker, and Bedrock have separate registered profiles. Registration documents their
ownership boundary; it is not executable qualification. Inspect infercrane integrations.
GCP Compute BYOC
Thegcp-compute adapter launches private, digest-pinned workers with an attached service account
and deterministic adoption identity. Configuration is all-or-nothing. See
GCP Compute BYOC. MIG, GKE, and Vertex remain separate registered,
deferred profiles rather than implicit aliases. Run infercrane doctor --gcp before provisioning;
it performs only identity and Compute API reads.
Google’s public Cloud Billing Catalog requires a consumer API key. To add current GCP GPU billing
components to the price market, create a key restricted to cloudbilling.googleapis.com and set it
as GOOGLE_CLOUD_BILLING_API_KEY on the control plane. This key is used only for the official
Compute Engine SKU catalog. The resulting rows are GPU-device rates, not complete VM quotes or
capacity evidence; CPU, memory, disk, network, quota, and launchability remain separate facts.
CoreWeave
Thecoreweave-cks profile is CKS-first: InferCrane reuses its namespaced Kubernetes lifecycle and
does not install or own the provider-managed GPU operator. The profile is registered but not yet
executable or locally qualified; real CKS qualification remains deferred.
Kubernetes
The Kubernetes adapter uses an explicit kubeconfig context, one namespace, an immutable default runtime image, and a worker Secret reference. It owns a bounded Deployment/Service set or one optional KServe InferenceService. Apply the reviewed namespace and RBAC manifests, then runinfercrane doctor --kubernetes. See Kubernetes for exact configuration,
security boundaries, local Kind qualification, and current real-GPU limitations.