Skip to main content

AWS EC2 BYOC

The AWS adapter provisions one EC2 instance per durable replica intent through Provider Contract V1. InferCrane assumes a customer role, uses explicit private networking, and adopts resources by idempotency token and ownership tags after uncertain responses. This is intentionally narrow. It is not EKS, SageMaker, automatic instance selection, public-IP bootstrap, or a general AWS abstraction.

Prerequisites

  • AWS CLI v2 on the control-plane host
  • a source identity permitted to call sts:AssumeRole
  • a role trust policy scoped to that source identity and, preferably, an external ID
  • an explicit private subnet and security group path between the control plane and worker port 8000
  • a GPU-compatible AMI containing NVIDIA drivers, Docker, and AWS CLI v2
  • an EC2 instance profile permitted to read exactly the worker API-key secret
  • an OCI runtime image pinned by sha256 digest
The assumed control-plane role needs a deliberately narrow permission set:
  • ec2:DescribeInstances, ec2:DescribeImages, ec2:DescribeVolumes, ec2:DescribeSnapshots, and ec2:GetConsoleOutput in the selected region;
  • ec2:RunInstances for the exact AMI, subnet, security group, network interface, volume, instance type, and managed instance tag boundary;
  • ec2:CreateTags for both instance/* and volume/*, restricted to ec2:CreateAction=RunInstances and aws:RequestTag/infercrane:managed=true;
  • ec2:TerminateInstances restricted to ec2:ResourceTag/infercrane:managed=true;
  • iam:PassRole for exactly the configured worker role and only to EC2.
Volume tagging is required because a root volume can remain billable independently of its instance. Restrict account IDs, resources, regions, tags, and instance types according to your AWS policy.

Control-plane configuration

Configure the complete set. Partial configuration fails startup rather than silently disabling safety controls.
INFERCRANE_AWS_SUBNET_IDS is an ordered list of approved private capacity boundaries. When AWS returns a definitive insufficient-capacity response, InferCrane tries the next subnet with a placement-specific idempotency token. It does not cross a boundary after timeouts, authentication failures, or ambiguous responses; the durable retry first discovers any owned instance by tag. INFERCRANE_AWS_SUBNET_ID remains supported for a single subnet. INFERCRANE_AWS_GPU and INFERCRANE_AWS_GPU_COUNT describe the exact accelerator topology of the configured instance type. A DeploymentSpec requesting another type or count fails before the EC2 create call. Configure a separate control-plane/provider profile when you need a different instance topology; InferCrane does not guess instance-family GPU counts. The Secrets Manager value referenced by INFERCRANE_AWS_WORKER_SECRET_ARN must contain the same worker credential configured as INFERCRANE_API_KEY on this self-hosted control plane. InferCrane uses that value for private worker health checks, routing, and explicit candidate validation; it persists only the secret ARN in provider metadata. INFERCRANE_AWS_IMAGE_DIGEST is the configured vLLM default and remains required for the adapter’s complete startup configuration. An SGLang or custom OCI revision supplies its own immutable image and argv; the EC2 adapter uses that revision workload instead of the vLLM default. It still uses the same private network, instance profile, worker secret and ownership tags. INFERCRANE_AWS_ROOT_VOLUME_GIB defaults to 200 GiB. InferCrane creates an encrypted gp3 root volume that is deleted with the worker. Keep enough headroom for the GPU runtime image, model artifacts, tokenizer data, and temporary download files; values below 50 GiB are rejected before AWS mutation. The 200 GiB default follows a real failure in which a current vLLM CUDA image exhausted a 100 GiB Deep Learning AMI root filesystem while Docker was extracting its FlashInfer layer. A smaller explicit value remains available for a qualified slim image, but must not inherit evidence from the default vLLM profile. The root volume defaults to the gp3 baseline of 3,000 IOPS and 125 MiB/s. Increase INFERCRANE_AWS_GP3_IOPS or INFERCRANE_AWS_GP3_THROUGHPUT_MIBPS only from measured evidence: higher provisioned performance can add AWS charges, and storage tuning cannot remove model engine initialization time. InferCrane validates AWS’s documented bounds and throughput-to-IOPS ratio before mutation and includes the selected values in RunInstances. InferCrane reads the AMI metadata before every create or adoption and overrides the AMI’s declared root device, not a hard-coded Linux device name. It rejects an invalid root mapping or a configured volume smaller than the source snapshot before RunInstances. The actual root-device name, size, and encrypted intent are included in the durable ownership tags so a retry cannot adopt an instance whose large disk was accidentally attached as an unused secondary volume. INFERCRANE_AWS_IMAGE_CACHE_POLICY defaults to prefer: a prewarmed exact digest is reused and a missing digest is pulled. Set it to required for a qualification or production pool whose AMI is contractually prewarmed. A required-cache miss fails the replica before any registry pull, is visible as image_cache_miss_required, and is cleaned up by the durable operation. This prevents a supposedly fast-start pool from silently becoming a long image pull; it does not claim that model weights are cached.

Reuse an immutable model-artifact snapshot

An operator-prepared encrypted EBS snapshot can remove the registry download from each replica boot. Its filesystem root must contain the qualified Hugging Face cache tree and the ext4 filesystem label must be INFERCRANE_ART. Tag the completed snapshot with:
InferCrane validates the configured snapshot before creating a worker, attaches it as an encrypted gp3 volume during RunInstances, discovers it by filesystem label, mounts it read-only, and starts vLLM or SGLang in offline artifact mode. The instance and volume carry the snapshot ID plus identity digest, allowing a retry after a lost create response to adopt only the same immutable cache. Use the durable artifact API to verify and record a fresh provider observation:
This command adopts and validates an existing snapshot; it does not build one. Cache population, regional copy, KMS policy, and retention remain explicit operator steps. required prevents a worker from launching without an exact mapping. prefer retains ordinary immutable model download when no mapping exists. A mapped snapshot that is incomplete, unencrypted, incorrectly tagged, or not mountable fails closed under either policy.

Build a cache snapshot without GPU billing

The repository includes a bounded AWS builder. It launches a temporary CPU instance, downloads one exact model commit into an encrypted EBS volume, stops the instance, creates the tagged snapshot, and deletes the temporary instance and volume. It never enables Fast Snapshot Restore or a provisioned initialization rate.
Use a current x86-64 Linux AMI with cloud-init, Python 3, and outbound HTTPS through the selected subnet. Public models need no Hugging Face token. For a repository that contains redundant weight formats, an operator may set a reviewed JSON ignore list; do not exclude files without qualifying the resulting runtime. The output snapshot-mapping.json is the exact value for INFERCRANE_AWS_ARTIFACT_SNAPSHOTS_JSON. The run ID is the recovery handle. A terminal or network interruption initiates best-effort cleanup; after a laptop power loss, repeat:
The immutable snapshot intentionally survives builder cleanup and remains billable. Delete it only after it is no longer referenced:

Prove a real cache hit

The cache qualification uses required, verifies the startup marker, sends real buffered and streaming requests, stores a measured benchmark row, deletes the GPU worker, and checks provider inventory against the pre-run baseline. It skips the already-completed full load matrix because this gate proves startup locality rather than throughput.
Do not describe a snapshot-backed deployment as a cache hit unless the archived inspect payload contains startup_evidence.artifact_cache: hit and final provider inventory is unchanged. Volumes created from EBS snapshots may initialize lazily. Set INFERCRANE_AWS_ARTIFACT_VOLUME_INITIALIZATION_RATE_MIBPS to an explicit value from 100 through 300 only when you accept AWS provisioned initialization charges. Fast Snapshot Restore is never enabled by InferCrane. Both options need separate cost and latency evidence. Validate role assumption without creating a resource:
This proves role assumption and caller identity. It intentionally does not call RunInstances, so resource-scoped CreateTags, quota, AMI/subnet, and capacity failures remain plan/deploy boundaries.

Deploy

The deploy command below can create a billable EC2 instance. InferCrane currently reports AWS price as unknown. The documented vLLM, SGLang, and portable custom-OCI L40S paths have narrow real qualification evidence, but that does not qualify another AMI, model, instance type, region, or network. Run doctor and plan first, obtain the required quota, and use explicit cost approval.
The requested region and GPU must exactly match the configured, qualified instance profile:
Closing the terminal does not stop provisioning. Resume with the operation ID printed by deploy. Deletion terminates only instances carrying InferCrane ownership tags and the persisted replica key. Inspect the measured startup waterfall after provisioning:
When ec2:GetConsoleOutput is permitted, InferCrane extracts only its closed startup markers and persists their timestamps plus the runtime-ready time. Arbitrary EC2 console output is discarded; runtime logs, model output, and credentials are never copied into provider metadata. inspect reports image cache and artifact cache separately; an artifact hit still does not prove the runtime is ready until health and served-model identity checks pass. For the portable-runtime path, apply examples/sglang.yaml or replace the placeholder image in examples/custom-oci.yaml in this repository. Every changed image/model/GPU combination remains its own real qualification gate. Do not treat Qwen as the only supported or tested workload. The repository includes immutable AWS qualification intents for three permissively licensed chat/reasoning families:
  • examples/aws-mistral-7b.yaml: Mistral 7B Instruct, Apache-2.0;
  • examples/aws-deepseek-r1-distill-7b.yaml: DeepSeek R1 Distill Qwen 7B, MIT;
  • examples/aws-granite-8b.yaml: IBM Granite 3.3 8B Instruct, Apache-2.0.
Qwen3-8B remains the cross-runtime tool-calling and structured-output regression fixture. BGE-M3 is the embedding protocol fixture, while gated Llama and Gemma recipes require the operator to accept their upstream terms. Each real run must archive the exact model revision and must not transfer a benchmark claim from one model family to another.

Security and accounting behavior

STS credentials are short lived and exist only in the child AWS CLI process environment. InferCrane does not persist or return them. Workers retrieve their API key directly from Secrets Manager through the instance profile. EC2 is launched without a public IP. Cost is reported as unknown. InferCrane does not ship a live AWS pricing catalog and will not infer cost from an instance-type name.

Qualification state

Hermetic contract tests cover create-response loss, adoption, replay, deletion, tag-scoped instance and volume inventory, private networking, immutable images, normalized provider failures, and credential redaction. The narrow real AWS evidence is recorded in AWS real-infrastructure evidence; inspect infercrane integrations for the current contract state.