Skip to main content

GCP Compute BYOC

The gcp-compute profile realizes one durable replica intent as one private Compute Engine VM. A deterministic resource name lets reconciliation adopt the VM after an uncertain create response. InferCrane does not store a service-account key, assign a public IP, or infer provider pricing.
The adapter is hermetically qualified. Real GCP GPU execution is deferred to consolidated manual qualification and must not be described as production-qualified yet.

Configure

The production image contains a checksum-pinned gcloud client. Authenticate with Application Default Credentials using an identity allowed to create, describe, list, and delete only the intended project resources. For Docker Compose, set GCLOUD_CONFIG_DIR to a read-only gcloud configuration directory. The entrypoint copies that bootstrap profile into container-local /tmp storage because gcloud requires a writable credential database; refreshed tokens are therefore never persisted back to the host mount. Prefer workload identity when the control plane runs on GCP; do not put a service-account key in a DeploymentSpec. Configure all fields together:

One-time administrator handoff

Keep the control plane and worker identities separate: For a self-hosted control plane outside Google Cloud, use service-account impersonation or Workload Identity Federation. For a control plane running on Google Cloud, attach the control identity to the runner. Do not download a service-account JSON key for either path. Prepare one private subnet with Private Google Access and allow TCP 8000 only from the control-plane network or identity boundary to the worker service account. A GPU worker has no public address, so the control plane must have a real private route into that subnet. Enabling an API or passing doctor does not create that route. Before configuring InferCrane, verify these non-mutating boundaries with the same effective control identity:
The final command should be empty for a new boundary. A billing budget alert must still be checked separately; budget alerts notify and do not stop Compute Engine spend. Allow read-only serial-port output if startup waterfall evidence is required. InferCrane retains only its closed infercrane_startup markers and discards every other console line.
Validate identity and every configured dependency without creating capacity:
The preflight describes the zone, machine type, accelerator type, private subnet, attached worker service account, worker secret, and immutable VM image. It also reads the matching regional GPU quota and project-wide GPUS_ALL_REGIONS quota. An observed insufficient limit fails before VM creation; a missing or provider-unknown metric remains unknown rather than becoming a fabricated pass or rejection. It fails when Private Google Access is not enabled because workers intentionally have no public IP. Quota is not stock: current zone capacity, the private route from the control plane, and runtime readiness remain real-infrastructure checks. The attached service account reads the worker credential from Secret Manager at startup. Grant it only secret-version access to that secret. The subnet must allow the control plane to reach the runtime port over private addressing. The boot disk defaults to 200 GiB and is created as pd-balanced. Do not reduce it below the measured sum of the immutable runtime image, model artifacts, temporary download space, and safety headroom. InferCrane refuses values below 50 GiB; the small default disk inherited from a base VM image is not sufficient evidence that an inference runtime can materialize successfully. For Container-Optimized OS, worker startup installs the image-compatible NVIDIA driver with cos-extensions, verifies the GPU device nodes, and mounts the driver libraries and devices into the runtime container before it can be marked started. A custom immutable VM image may provide a preinstalled driver instead, but nvidia-smi must pass before the container is launched. This driver stage is part of startup evidence; a running VM is not treated as runtime readiness.

Reuse an immutable model cache

For vLLM and SGLang, InferCrane can adopt a pre-populated zonal Persistent Disk rather than downloading the same immutable model on every VM. Storage creation and population remain an administrator-controlled build step. InferCrane verifies the disk, attaches it read-only, keeps it when a worker is deleted, mounts it at the Hugging Face cache root, and runs the container in offline artifact mode. Prepare a filesystem containing the Hugging Face cache for one exact model@commit. The disk must be in the configured VM zone and its description must contain only the exact identity binding:
Then adopt and observe the cache through the same provider-neutral operation surface:
prefer falls back to normal artifact materialization only when no disk is mapped. A configured disk that is missing, in another zone, not READY, or bound to a different digest always fails before VM creation. required also rejects an unmapped model. InferCrane never creates or deletes the shared disk implicitly, and it does not claim a cache hit until the verified disk mounts on the worker. Real attach, mount, and startup-time improvement still require paid GCP qualification.

Deploy

The advanced provider.adapter field selects an exact infrastructure profile. It is optional while only one default adapter exists for a cloud/runtime pair.
The operation remains durable if the terminal disconnects. Inventory and deletion are restricted to resources labeled infercrane-managed=true; real cleanup must still be confirmed during manual qualification. Artifact cache disks are deliberately excluded from that ownership label and remain after worker deletion. The guarded real-provider harness is available when an isolated paid project is ready. The default GCP path includes the generic buffered, streaming, long-context, long-generation, and overload performance matrix; the larger concurrency sweep remains explicit. Final inventory compares the baseline and after-state for InferCrane-owned VMs, boot disks, addresses, and forwarding rules.

Other GCP profiles

gcp-mig, gcp-gke, and gcp-vertex have independent registered capability boundaries. They are not aliases for Compute Engine and are not executable or qualified until their own lifecycle contracts pass. infercrane integrations --output json is authoritative.