Skip to main content

RunPod

RunPod is one registered infrastructure provider with separate executable adapters. Built-in vLLM elastic deployments use SkyPilot. Immutable custom OCI workloads use the native runpod-pods REST adapter, which does not require SSH inside the runtime image. Serverless deployments use RunPod’s native endpoint lifecycle. InferCrane’s base production stack does not require RunPod; enable it explicitly with compose.production.runpod.yaml.
The commands in this guide can create billable GPU resources. The RunPod adapters have hermetic contract coverage, but the current elastic path remains experimental until the exact release, image, model, GPU, and RunPod environment pass the credentialed qualification suite. A successful plan or local simulation is not real-provider qualification.

Before the first deployment

1

Store a scoped RunPod key

Give the key only the Pod and Serverless permissions you intend this control plane to use. Keep it outside the repository and restrict the file:
2

Start the explicit RunPod production profile

Copy .env.production.example to a private path, replace its placeholder values, and add the key-file path from .env.runpod.example. The Serverless template ID is optional unless you use provider-native Serverless.
The base production profile is provider-neutral. The additional overlay mounts the key read-only and enables native RunPod Pods and optional provider-native Serverless. SkyPilot is not enabled implicitly by a RunPod credential.Large custom OCI workloads may need more than the default 100 GiB container disk. Set, for example, INFERCRANE_RUNPOD_CONTAINER_DISK_GIB=500 in the same private environment file. Disk size is validated before the control plane starts and may affect provider cost.Native Pods can instead reuse a model-specific RunPod network volume. Create the volume in the intended data center and name it infercrane-artifact- followed by the first 20 hexadecimal characters of SHA-256 over the exact model@commit identity. Then configure the exact mapping:
For gated repositories, create a scoped RunPod secret and configure only its name:
Native Pod requests contain RunPod’s secret reference syntax, never the Hugging Face token. InferCrane fails closed on an invalid secret name and cannot verify the secret value without launching a workload.To populate a new identity-named volume with a resumable transfer, first set the exact model/commit, data center, size, current provider rates, retention window, and watchdog, then run make plan-runpod-artifact-cache. Review its JSON worst-case cost before setting INFERCRANE_RUNPOD_MAX_COST_USD and using make build-runpod-artifact-cache. The build command is billable; its hermetic safety test is make test-runpod-artifact-cache.Before creating a Pod, InferCrane performs a read-only volume lookup, verifies its ID, exact identity-derived name, positive size, and data center, then mounts it at /workspace. Standard Hugging Face downloads and filesystem-materialized model profiles use that persistent path. The first deployment may still download the model; required proves the persistent volume was attached, not that its bytes were already warm. Runtime readiness remains the proof that the exact model became serveable. Deleting a Pod does not delete this operator-owned volume.
3

Connect and run read-only checks

Use the same URL and API key configured in the private environment file:
doctor checks local dependencies and credentials without creating a Pod. Stop here if a required check fails.
4

Plan before paying

Confirm the model, immutable artifact identity, GPU, compute mode, and replica bounds. RunPod stock can change after a read-only availability check, and InferCrane never fabricates a price when the provider has not supplied trustworthy cost data.

Deploy and follow durable state

The command returns a durable operation ID. Closing the terminal does not cancel the operation. Use the ID from the deploy response to reattach, or inspect the deployment timeline:
Do not submit another deployment because allocation is slow. Reuse the same idempotency key and inspect RunPod inventory before retrying an unresolved create response.

Delete and verify cleanup

Deletion is also a durable operation. Preview it, confirm explicitly, then follow provider cleanup:
The final orphans check proves InferCrane has no known unmanaged resource. For a credentialed qualification, also verify the RunPod inventory itself: zero run-owned Pods and zero run-owned Serverless endpoints. Never erase the PostgreSQL state while provider cleanup is unresolved. See provider setup for the credential boundary, production operations for self-hosting, serverless lifecycle for worker-zero behavior, and compatibility and qualification for the current real-provider evidence boundary.