RunPod
RunPod is one registered infrastructure provider with separate executable adapters. Built-in vLLM elastic deployments use SkyPilot. Immutable custom OCI workloads use the nativerunpod-pods REST
adapter, which does not require SSH inside the runtime image. Serverless deployments use RunPod’s
native endpoint lifecycle. InferCrane’s base production stack does not require RunPod; enable it
explicitly with compose.production.runpod.yaml.
Before the first deployment
1
Store a scoped RunPod key
Give the key only the Pod and Serverless permissions you intend this control plane to use. Keep
it outside the repository and restrict the file:
2
Start the explicit RunPod production profile
Copy The base production profile is provider-neutral. The additional overlay mounts the key
read-only and enables native RunPod Pods and optional provider-native Serverless. SkyPilot is
not enabled implicitly by a RunPod credential.Large custom OCI workloads may need more than the default 100 GiB container disk. Set, for
example, For gated repositories, create a scoped RunPod secret and configure only its name:Native Pod requests contain RunPod’s secret reference syntax, never the Hugging Face token.
InferCrane fails closed on an invalid secret name and cannot verify the secret value without
launching a workload.To populate a new identity-named volume with a resumable transfer, first set the exact
model/commit, data center, size, current provider rates, retention window, and watchdog, then run
.env.production.example to a private path, replace its placeholder values, and add the
key-file path from .env.runpod.example. The Serverless template ID is optional unless you use
provider-native Serverless.INFERCRANE_RUNPOD_CONTAINER_DISK_GIB=500 in the same private environment file. Disk
size is validated before the control plane starts and may affect provider cost.Native Pods can instead reuse a model-specific RunPod network volume. Create the volume in the
intended data center and name it infercrane-artifact- followed by the first 20 hexadecimal
characters of SHA-256 over the exact model@commit identity. Then configure the exact mapping:make plan-runpod-artifact-cache. Review its JSON worst-case cost before setting
INFERCRANE_RUNPOD_MAX_COST_USD and using make build-runpod-artifact-cache. The build command is
billable; its hermetic safety test is make test-runpod-artifact-cache.Before creating a Pod, InferCrane performs a read-only volume lookup, verifies its ID, exact
identity-derived name, positive size, and data center, then mounts it at /workspace. Standard
Hugging Face downloads and filesystem-materialized model profiles use that persistent path.
The first deployment may still download the model; required proves the persistent volume was
attached, not that its bytes were already warm. Runtime readiness remains the proof that the
exact model became serveable. Deleting a Pod does not delete this operator-owned volume.3
Connect and run read-only checks
Use the same URL and API key configured in the private environment file:
doctor checks local dependencies and credentials without creating a Pod. Stop here if a
required check fails.4
Plan before paying
Deploy and follow durable state
- Elastic
- Custom OCI
- Serverless
Delete and verify cleanup
Deletion is also a durable operation. Preview it, confirm explicitly, then follow provider cleanup:orphans check proves InferCrane has no known unmanaged resource. For a credentialed
qualification, also verify the RunPod inventory itself: zero run-owned Pods and zero run-owned
Serverless endpoints. Never erase the PostgreSQL state while provider cleanup is unresolved.
See provider setup for the credential boundary, production operations
for self-hosting, serverless lifecycle for worker-zero behavior, and
compatibility and qualification for the current real-provider evidence boundary.