Know what a deployment is waiting for
Provisioning is a durable operation. The CLI can disconnect while the control plane continues to reconcile one deterministic provider identity per replica intent. Reconnect with the operation ID printed bydeploy:
Choose and check an adapter before provisioning
integrations distinguishes registration, local evidence, and real-system qualification. plan
validates the requested intent without allocating capacity. When at least three successful,
tenant-scoped observations match the exact provider adapter, runtime, region, and accelerator,
the plan also shows an observed p50 from durable replica intent through runtime readiness. P95 is
withheld until twenty successes. These values are historical evidence, never a stock reservation
or readiness guarantee. Follow the exact adapter guide for
credentials, networking, immutable images, runtime contract, preflight, cleanup, and current limits:
- RunPod elastic and Serverless
- AWS EC2 BYOC
- GCP Compute BYOC
- Kubernetes and KServe
- NVIDIA Dynamo
- Exact-combination compatibility check
You do not need to read every provider guide. Select one only after the inputs above identify the
operator-owned boundary:
If these requirements do not select exactly one adapter, stop and collect workload, security,
residency, capacity, and cost constraints; InferCrane will not silently choose infrastructure.
status shows whether the deployment is serving and whether desired capacity is still converging.
operation watch follows the blocking mutation. logs follows durable events and is the best view
when you need the transition history rather than the latest snapshot.
Read the current stage
Not every provider exposes container download, artifact transfer, model load, and runtime
initialization separately. InferCrane reports the narrowest boundary supported by fresh evidence.
For example,
provider_capacity_or_worker_initialization means exactly that the product cannot
separate those phases from the available observation.
Diagnose a long wait
- Copy the blocking operation ID from
status. - Run
operation watchand note the last changed stage, provider message, attempt, and retry time. - Run
logs --followin a second terminal to distinguish a repeated observation from a real state transition. - Run
infercrane inspect DEPLOYMENT --output jsonfor the persisted provider identity and non-secret infrastructure details. - If the provider retains a failed resource - for example after an interrupted image pull - cancel the durable operation before replacing it. Do not start a second deploy with a different identity.
Capacity evidence is advisory
Before the first create, InferCrane discovers the deterministic resource. An existing resource is adopted before mutable stock is consulted. If no resource exists, a provider adapter may reportavailable, constrained, unavailable, or unknown:
unavailabledefers creation with a retryable error;constrainedproceeds with an explicit warning because stock is not a reservation;unknownpermits the normal provider attempt but cannot support a capacity claim.
Reuse prewarmed runtime images
AWS and GCP worker bootstrap checks for the immutable runtime image locally before pulling it. A customer-maintained AMI or VM image may therefore prewarm the exact qualified digest and avoid retransferring those container layers. Startup logs exposeidentity_start, identity_ready,
image_check, image_cache_hit, image_pull_start, image_pull_complete, runtime_start, and
runtime_container_started timestamps without credentials. AWS observations parse only this
closed marker grammar, discard all other console output, and persist it with the independently
observed runtime_ready_at timestamp. infercrane inspect DEPLOYMENT renders the measured
waterfall.
AWS can additionally enforce INFERCRANE_AWS_IMAGE_CACHE_POLICY=required. A miss then fails before
the registry pull instead of silently violating a fast-start pool’s contract. The default prefer
policy retains the safe fallback to an immutable pull.
This optimization concerns container layers only. It does not prove model weights are cached.
Model-artifact locality remains a separate, provider-native observation through the
artifact cache contract. InferCrane does not silently enable billed
snapshot acceleration or accept a mutable image tag.
Idempotency and cleanup
The provider resource key is stored before any external create call. If the create succeeds but its response is lost, retry performs discovery and adopts the resource. Delete re-observes asynchronous provider removal and cannot mark the replica deleted while the resource remains visible.Reconcile an interrupted create
- Do not create a new deployment, revision, or idempotency key.
-
Reattach to the persisted operation and export the stored identities:
-
In the provider’s read-only inventory, locate the exact
provider_resource_id, deterministic external key, ownership tags/labels, revision, and replica ordinal frominspect. A same-name resource with mismatched ownership is not adoptable. -
If the identities match, let the existing operation resume. Reissuing the original deploy/apply
request is safe only with its original idempotency key and identical intent. Replaying
rollout provision DEPLOYMENT REVISION_ID --waitis safe for the same persisted revision; do not runrollout createagain. - If provider inventory is unavailable, incomplete, or mismatched, stop mutation and preserve PostgreSQL. Resolve ownership manually before create or delete.
orphans
response cannot prove a provider credential can see every external resource.
After a paid qualification run, use both views: