GCP Compute BYOC
Thegcp-compute profile realizes one durable replica intent as one private Compute Engine VM. A
deterministic resource name lets reconciliation adopt the VM after an uncertain create response.
InferCrane does not store a service-account key, assign a public IP, or infer provider pricing.
Configure
The production image contains a checksum-pinnedgcloud client. Authenticate with Application
Default Credentials using an identity allowed to create, describe, list, and delete only the
intended project resources. For Docker Compose, set GCLOUD_CONFIG_DIR to a read-only gcloud
configuration directory. The entrypoint copies that bootstrap profile into container-local /tmp
storage because gcloud requires a writable credential database; refreshed tokens are therefore
never persisted back to the host mount. Prefer workload identity when the control plane runs on GCP;
do not put a service-account key in a DeploymentSpec. Configure all fields together:
One-time administrator handoff
Keep the control plane and worker identities separate:
For a self-hosted control plane outside Google Cloud, use service-account impersonation or Workload
Identity Federation. For a control plane running on Google Cloud, attach the control identity to the
runner. Do not download a service-account JSON key for either path.
Prepare one private subnet with Private Google Access and allow TCP
8000 only from the
control-plane network or identity boundary to the worker service account. A GPU worker has no public
address, so the control plane must have a real private route into that subnet. Enabling an API or
passing doctor does not create that route.
Before configuring InferCrane, verify these non-mutating boundaries with the same effective control
identity:
infercrane_startup markers and discards every other console line.
GPUS_ALL_REGIONS quota. An observed insufficient limit fails before VM
creation; a missing or provider-unknown metric remains unknown rather than becoming a fabricated
pass or rejection. It fails when Private Google Access is not enabled because workers intentionally
have no public IP. Quota is not stock: current zone capacity, the private route from the control
plane, and runtime readiness remain real-infrastructure checks.
The attached service account reads the worker credential from Secret Manager at startup. Grant it
only secret-version access to that secret. The subnet must allow the control plane to reach the
runtime port over private addressing.
The boot disk defaults to 200 GiB and is created as pd-balanced. Do not reduce it below the
measured sum of the immutable runtime image, model artifacts, temporary download space, and safety
headroom. InferCrane refuses values below 50 GiB; the small default disk inherited from a base VM
image is not sufficient evidence that an inference runtime can materialize successfully.
For Container-Optimized OS, worker startup installs the image-compatible NVIDIA driver with
cos-extensions, verifies the GPU device nodes, and mounts the driver libraries and devices into
the runtime container before it can be marked started. A custom immutable VM image may provide a
preinstalled driver instead, but nvidia-smi must pass before the container is launched. This
driver stage is part of startup evidence; a running VM is not treated as runtime readiness.
Reuse an immutable model cache
For vLLM and SGLang, InferCrane can adopt a pre-populated zonal Persistent Disk rather than downloading the same immutable model on every VM. Storage creation and population remain an administrator-controlled build step. InferCrane verifies the disk, attaches it read-only, keeps it when a worker is deleted, mounts it at the Hugging Face cache root, and runs the container in offline artifact mode. Prepare a filesystem containing the Hugging Face cache for one exactmodel@commit. The disk must
be in the configured VM zone and its description must contain only the exact identity binding:
prefer falls back to normal artifact materialization only when no disk is mapped. A configured
disk that is missing, in another zone, not READY, or bound to a different digest always fails
before VM creation. required also rejects an unmapped model. InferCrane never creates or deletes
the shared disk implicitly, and it does not claim a cache hit until the verified disk mounts on the
worker. Real attach, mount, and startup-time improvement still require paid GCP qualification.
Deploy
The advancedprovider.adapter field selects an exact infrastructure profile. It is optional while
only one default adapter exists for a cloud/runtime pair.
infercrane-managed=true; real cleanup must still be confirmed during manual
qualification. Artifact cache disks are deliberately excluded from that ownership label and remain
after worker deletion.
The guarded real-provider harness is available when an isolated paid project is ready. The default
GCP path includes the generic buffered, streaming, long-context, long-generation, and overload
performance matrix; the larger concurrency sweep remains explicit. Final inventory compares the
baseline and after-state for InferCrane-owned VMs, boot disks, addresses, and forwarding rules.
Other GCP profiles
gcp-mig, gcp-gke, and gcp-vertex have independent registered capability boundaries. They are
not aliases for Compute Engine and are not executable or qualified until their own lifecycle
contracts pass. infercrane integrations --output json is authoritative.