Treat model locality as evidence
Large model artifacts often dominate readiness time. InferCrane records cache state through a neutral contract while each provider adapter retains its native mechanism: a volume or provider cache, an AWS volume or snapshot, a Kubernetes volume or node cache, or another qualified implementation.Request prefetch
execution: not_configured instead of implying that warming began.
Submission is not proof that the artifact is present. The adapter must publish a fresh
observation before a planner can rely on locality.
Record an adapter observation
present, prefetching, missing, and unknown. Expiry prevents stale provider
state from becoming a permanent planning assumption.
Cache population remains adapter-specific and requires qualification on the selected
infrastructure.
requested or running describes work, while only a fresh present
observation describes locality.prefetch still does
not imply that the control API copied model bytes.
For Kubernetes, the adapter can verify and adopt an administrator-prepared PVC bound to one exact
model@commit digest. The claim must support multi-reader access. InferCrane mounts it read-only and
sets the runtime’s Hugging Face cache root, while Kubernetes and the CSI driver retain storage
ownership. See the Kubernetes cache setup.
For GCP Compute, the adapter verifies an administrator-prepared zonal Persistent Disk whose exact
description binds it to one immutable model@commit digest. InferCrane attaches it read-only with
deletion disabled, mounts it as the runtime’s Hugging Face cache, and leaves its lifecycle with the
customer. See the GCP cache setup.
For RunPod Pods, the repository includes a guarded, resumable network-volume population workflow.
plan-runpod-artifact-cache is read-only and refuses to quote a cost without operator-supplied
current GPU and storage rates. build-runpod-artifact-cache additionally requires explicit paid
resource approval, a hard maximum cost, an immutable Hugging Face commit, and a bounded watchdog.
It deletes the temporary preparer Pod but preserves the identity-named volume for deliberate reuse;
volume deletion is a separate explicit command. Hermetic tests qualify these safety contracts only,
not RunPod transfer speed or model correctness.
Why prefetch is not a universal download command
Providers expose different ownership boundaries. For example, RunPod Serverless configures one cached model on an endpoint, while an EC2 or Kubernetes deployment commonly obtains locality from a persistent volume, snapshot, node cache, or image strategy owned by that workload. Those are not interchangeable APIs. InferCrane therefore does not download model weights through a custom protocol and does not mark a prefetch request as a cache hit. An adapter may consume the durable intent using its native mechanism, but it must publish a bounded fresh observation before planning or placement can rely on locality. Until then,infercrane plan reports artifact cache and startup time as unknown and
unavailable.
Container image prewarming is separate
AWS and GCP bootstrap reuse an exact digest already present in a customer-maintained VM image and emit bounded startup-stage timestamps. This can remove a large container transfer, but it does not mean the model artifact is local. Keep these evidence boundaries separate:
InferCrane never converts the first row into the second or third. AWS Fast Snapshot Restore and
similar paid acceleration remain explicit operator choices rather than silent defaults.