Skip to main content

Custom OCI workloads

InferCrane can launch a user-supplied inference image without becoming an image builder or container engine. The deployment revision records the exact image digest, argv and observable lifecycle contract before a provider mutation is submitted.
The command is complete executable argv, not image-entrypoint arguments and not a shell program. Providers override the image entrypoint with command[0] and preserve every remaining argument. ${MODEL}, ${MODEL_REVISION}, and ${PORT} are replaced without losing argument boundaries. The workload receives the worker credential as INFERCRANE_WORKER_API_KEY; it must read that environment variable directly. InferCrane rejects credentials in command arguments because runtimes commonly print parsed argv during startup. The value is not persisted in revision or provider metadata. InferCrane resolves model.id through the Hugging Face ModelArtifact adapter before launch; a custom image does not imply arbitrary artifact protocols.

Runtime contract

The image must be pinned by @sha256. The supported protocol and probe paths are deliberately narrow: OpenAI-compatible HTTP, /health, /v1/models, and /metrics. Client disconnect must cancel upstream HTTP work, and InferCrane withdraws the router generation and drains active connections before provider deletion. Custom OCI is simulated-qualified on the AWS EC2, GCP Compute, and namespaced Kubernetes elastic paths. RunPod custom OCI uses the native runpod-pods REST adapter instead of SkyPilot, so the upstream image does not need an SSH daemon or rsync. Each adapter compiles the same immutable image, executable argv, probe, shutdown, secret, and accelerator-count contract into its native resource. RunPod additionally maps the runtime port to its HTTPS Pod proxy and uses the provider API for deterministic adoption and deletion. Real GPU evidence remains bound to the exact image, model revision, provider, accelerator topology, and workload. Serverless, arbitrary probe scripts, mutable tags, image builds and autoscaling signals for custom runtimes are rejected rather than silently assumed. Configure enough RunPod container disk for the image, checkpoint, and runtime overhead with INFERCRANE_RUNPOD_CONTAINER_DISK_GIB; the default is 100 GiB and the accepted range is 50–2048. Inspect the executable support boundary with: