Custom OCI workloads
InferCrane can launch a user-supplied inference image without becoming an image builder or container engine. The deployment revision records the exact image digest, argv and observable lifecycle contract before a provider mutation is submitted.command[0] and preserve every remaining argument.
${MODEL}, ${MODEL_REVISION}, and ${PORT} are replaced without losing argument boundaries. The workload receives the worker credential as
INFERCRANE_WORKER_API_KEY; it must read that environment variable directly. InferCrane rejects
credentials in command arguments because runtimes commonly print parsed argv during startup. The
value is not persisted in revision or provider metadata.
InferCrane resolves model.id through the Hugging Face ModelArtifact adapter before launch; a
custom image does not imply arbitrary artifact protocols.
Runtime contract
The image must be pinned by@sha256. The supported protocol and probe paths are deliberately
narrow: OpenAI-compatible HTTP, /health, /v1/models, and /metrics. Client disconnect must
cancel upstream HTTP work, and InferCrane withdraws the router generation and drains active
connections before provider deletion.
Custom OCI is simulated-qualified on the AWS EC2, GCP Compute, and namespaced Kubernetes elastic
paths. RunPod custom OCI uses the native runpod-pods REST adapter instead of SkyPilot, so the
upstream image does not need an SSH daemon or rsync. Each adapter compiles the same immutable image,
executable argv, probe, shutdown, secret, and accelerator-count contract into its native resource.
RunPod additionally maps the runtime port to its HTTPS Pod proxy and uses the provider API for
deterministic adoption and deletion.
Real GPU evidence remains bound to the exact image, model revision, provider, accelerator topology,
and workload. Serverless, arbitrary probe scripts, mutable tags, image builds and autoscaling
signals for custom runtimes are rejected rather than silently assumed. Configure enough RunPod
container disk for the image, checkpoint, and runtime overhead with
INFERCRANE_RUNPOD_CONTAINER_DISK_GIB; the default is 100 GiB and the accepted range is 50–2048.
Inspect the executable support boundary with: