Skip to main content

Treat model locality as evidence

Large model artifacts often dominate readiness time. InferCrane records cache state through a neutral contract while each provider adapter retains its native mechanism: a volume or provider cache, an AWS volume or snapshot, a Kubernetes volume or node cache, or another qualified implementation.
The result separates observations from requested work and shows whether an observation is still fresh.

Request prefetch

This persists durable intent. When an adapter for the named provider is configured, InferCrane ensures that intent with a stable idempotency key and checkpoints the provider operation identity. If a provider response is lost, a retry may repeat the API call with the same key and adopt the same logical operation; it must not create a second cache resource. If no adapter is configured, the response says execution: not_configured instead of implying that warming began. Submission is not proof that the artifact is present. The adapter must publish a fresh observation before a planner can rely on locality.

Record an adapter observation

Valid states are present, prefetching, missing, and unknown. Expiry prevents stale provider state from becoming a permanent planning assumption.
Cache population remains adapter-specific and requires qualification on the selected infrastructure. requested or running describes work, while only a fresh present observation describes locality.
For AWS, the configured adapter verifies and adopts an existing encrypted EBS snapshot whose tags bind it to the exact immutable model identity. A completed provider operation triggers a fresh adapter observation automatically. Snapshot creation remains an explicit operator build boundary; the included bounded builder prepares it on a temporary CPU instance. The word prefetch still does not imply that the control API copied model bytes. For Kubernetes, the adapter can verify and adopt an administrator-prepared PVC bound to one exact model@commit digest. The claim must support multi-reader access. InferCrane mounts it read-only and sets the runtime’s Hugging Face cache root, while Kubernetes and the CSI driver retain storage ownership. See the Kubernetes cache setup. For GCP Compute, the adapter verifies an administrator-prepared zonal Persistent Disk whose exact description binds it to one immutable model@commit digest. InferCrane attaches it read-only with deletion disabled, mounts it as the runtime’s Hugging Face cache, and leaves its lifecycle with the customer. See the GCP cache setup. For RunPod Pods, the repository includes a guarded, resumable network-volume population workflow. plan-runpod-artifact-cache is read-only and refuses to quote a cost without operator-supplied current GPU and storage rates. build-runpod-artifact-cache additionally requires explicit paid resource approval, a hard maximum cost, an immutable Hugging Face commit, and a bounded watchdog. It deletes the temporary preparer Pod but preserves the identity-named volume for deliberate reuse; volume deletion is a separate explicit command. Hermetic tests qualify these safety contracts only, not RunPod transfer speed or model correctness.

Why prefetch is not a universal download command

Providers expose different ownership boundaries. For example, RunPod Serverless configures one cached model on an endpoint, while an EC2 or Kubernetes deployment commonly obtains locality from a persistent volume, snapshot, node cache, or image strategy owned by that workload. Those are not interchangeable APIs. InferCrane therefore does not download model weights through a custom protocol and does not mark a prefetch request as a cache hit. An adapter may consume the durable intent using its native mechanism, but it must publish a bounded fresh observation before planning or placement can rely on locality. Until then, infercrane plan reports artifact cache and startup time as unknown and unavailable.

Container image prewarming is separate

AWS and GCP bootstrap reuse an exact digest already present in a customer-maintained VM image and emit bounded startup-stage timestamps. This can remove a large container transfer, but it does not mean the model artifact is local. Keep these evidence boundaries separate: InferCrane never converts the first row into the second or third. AWS Fast Snapshot Restore and similar paid acceleration remain explicit operator choices rather than silent defaults.