AWS EC2 BYOC
The AWS adapter provisions one EC2 instance per durable replica intent through Provider Contract V1. InferCrane assumes a customer role, uses explicit private networking, and adopts resources by idempotency token and ownership tags after uncertain responses. This is intentionally narrow. It is not EKS, SageMaker, automatic instance selection, public-IP bootstrap, or a general AWS abstraction.Prerequisites
- AWS CLI v2 on the control-plane host
- a source identity permitted to call
sts:AssumeRole - a role trust policy scoped to that source identity and, preferably, an external ID
- an explicit private subnet and security group path between the control plane and worker port
8000 - a GPU-compatible AMI containing NVIDIA drivers, Docker, and AWS CLI v2
- an EC2 instance profile permitted to read exactly the worker API-key secret
- an OCI runtime image pinned by
sha256digest
ec2:DescribeInstances,ec2:DescribeImages,ec2:DescribeVolumes,ec2:DescribeSnapshots, andec2:GetConsoleOutputin the selected region;ec2:RunInstancesfor the exact AMI, subnet, security group, network interface, volume, instance type, and managed instance tag boundary;ec2:CreateTagsfor bothinstance/*andvolume/*, restricted toec2:CreateAction=RunInstancesandaws:RequestTag/infercrane:managed=true;ec2:TerminateInstancesrestricted toec2:ResourceTag/infercrane:managed=true;iam:PassRolefor exactly the configured worker role and only to EC2.
Control-plane configuration
Configure the complete set. Partial configuration fails startup rather than silently disabling safety controls.INFERCRANE_AWS_SUBNET_IDS is an ordered list of approved private capacity boundaries. When AWS
returns a definitive insufficient-capacity response, InferCrane tries the next subnet with a
placement-specific idempotency token. It does not cross a boundary after timeouts, authentication
failures, or ambiguous responses; the durable retry first discovers any owned instance by tag.
INFERCRANE_AWS_SUBNET_ID remains supported for a single subnet.
INFERCRANE_AWS_GPU and INFERCRANE_AWS_GPU_COUNT describe the exact accelerator topology of the
configured instance type. A DeploymentSpec requesting another type or count fails before the EC2
create call. Configure a separate control-plane/provider profile when you need a different instance
topology; InferCrane does not guess instance-family GPU counts.
The Secrets Manager value referenced by INFERCRANE_AWS_WORKER_SECRET_ARN must contain the same
worker credential configured as INFERCRANE_API_KEY on this self-hosted control plane. InferCrane
uses that value for private worker health checks, routing, and explicit candidate validation; it
persists only the secret ARN in provider metadata.
INFERCRANE_AWS_IMAGE_DIGEST is the configured vLLM default and remains required for the adapter’s
complete startup configuration. An SGLang or custom OCI revision supplies its own immutable image
and argv; the EC2 adapter uses that revision workload instead of the vLLM default. It still uses the
same private network, instance profile, worker secret and ownership tags.
INFERCRANE_AWS_ROOT_VOLUME_GIB defaults to 200 GiB. InferCrane creates an encrypted gp3 root
volume that is deleted with the worker. Keep enough headroom for the GPU runtime image, model
artifacts, tokenizer data, and temporary download files; values below 50 GiB are rejected before
AWS mutation. The 200 GiB default follows a real failure in which a current vLLM CUDA image exhausted
a 100 GiB Deep Learning AMI root filesystem while Docker was extracting its FlashInfer layer. A
smaller explicit value remains available for a qualified slim image, but must not inherit evidence
from the default vLLM profile.
The root volume defaults to the gp3 baseline of 3,000 IOPS and 125 MiB/s. Increase
INFERCRANE_AWS_GP3_IOPS or INFERCRANE_AWS_GP3_THROUGHPUT_MIBPS only from measured evidence:
higher provisioned performance can add AWS charges, and storage tuning cannot remove model engine
initialization time. InferCrane validates AWS’s documented bounds and throughput-to-IOPS ratio
before mutation and includes the selected values in RunInstances.
InferCrane reads the AMI metadata before every create or adoption and overrides the AMI’s declared
root device, not a hard-coded Linux device name. It rejects an invalid root mapping or a configured
volume smaller than the source snapshot before RunInstances. The actual root-device name, size,
and encrypted intent are included in the durable ownership tags so a retry cannot adopt an instance
whose large disk was accidentally attached as an unused secondary volume.
INFERCRANE_AWS_IMAGE_CACHE_POLICY defaults to prefer: a prewarmed exact digest is reused and a
missing digest is pulled. Set it to required for a qualification or production pool whose AMI is
contractually prewarmed. A required-cache miss fails the replica before any registry pull, is
visible as image_cache_miss_required, and is cleaned up by the durable operation. This prevents a
supposedly fast-start pool from silently becoming a long image pull; it does not claim that model
weights are cached.
Reuse an immutable model-artifact snapshot
An operator-prepared encrypted EBS snapshot can remove the registry download from each replica boot. Its filesystem root must contain the qualified Hugging Face cache tree and the ext4 filesystem label must beINFERCRANE_ART. Tag the completed snapshot with:
gp3 volume during RunInstances, discovers it by filesystem label, mounts it read-only, and starts
vLLM or SGLang in offline artifact mode. The instance and volume carry the snapshot ID plus identity
digest, allowing a retry after a lost create response to adopt only the same immutable cache.
Use the durable artifact API to verify and record a fresh provider observation:
required prevents a worker
from launching without an exact mapping. prefer retains ordinary immutable model download when no
mapping exists. A mapped snapshot that is incomplete, unencrypted, incorrectly tagged, or not
mountable fails closed under either policy.
Build a cache snapshot without GPU billing
The repository includes a bounded AWS builder. It launches a temporary CPU instance, downloads one exact model commit into an encrypted EBS volume, stops the instance, creates the tagged snapshot, and deletes the temporary instance and volume. It never enables Fast Snapshot Restore or a provisioned initialization rate.snapshot-mapping.json is the exact value for
INFERCRANE_AWS_ARTIFACT_SNAPSHOTS_JSON.
The run ID is the recovery handle. A terminal or network interruption initiates best-effort cleanup;
after a laptop power loss, repeat:
Prove a real cache hit
The cache qualification usesrequired, verifies the startup marker, sends real buffered and
streaming requests, stores a measured benchmark row, deletes the GPU worker, and checks provider
inventory against the pre-run baseline. It skips the already-completed full load matrix because this
gate proves startup locality rather than throughput.
inspect payload
contains startup_evidence.artifact_cache: hit and final provider inventory is unchanged.
Volumes created from EBS snapshots may initialize lazily. Set
INFERCRANE_AWS_ARTIFACT_VOLUME_INITIALIZATION_RATE_MIBPS to an explicit value from 100 through 300
only when you accept AWS provisioned initialization charges. Fast Snapshot Restore is never enabled
by InferCrane. Both options need separate cost and latency evidence.
Validate role assumption without creating a resource:
RunInstances, so
resource-scoped CreateTags, quota, AMI/subnet, and capacity failures remain plan/deploy boundaries.
Deploy
The requested region and GPU must exactly match the configured, qualified instance profile:deploy.
Deletion terminates only instances carrying InferCrane ownership tags and the persisted replica key.
Inspect the measured startup waterfall after provisioning:
ec2:GetConsoleOutput is permitted, InferCrane extracts only its closed startup markers and
persists their timestamps plus the runtime-ready time. Arbitrary EC2 console output is discarded;
runtime logs, model output, and credentials are never copied into provider metadata.
inspect reports image cache and artifact cache separately; an artifact hit still does not prove the
runtime is ready until health and served-model identity checks pass.
For the portable-runtime path, apply
examples/sglang.yaml or
replace the placeholder image in
examples/custom-oci.yaml in this repository.
Every changed image/model/GPU combination remains its own real qualification gate.
Do not treat Qwen as the only supported or tested workload. The repository includes immutable AWS
qualification intents for three permissively licensed chat/reasoning families:
examples/aws-mistral-7b.yaml: Mistral 7B Instruct, Apache-2.0;examples/aws-deepseek-r1-distill-7b.yaml: DeepSeek R1 Distill Qwen 7B, MIT;examples/aws-granite-8b.yaml: IBM Granite 3.3 8B Instruct, Apache-2.0.
Security and accounting behavior
STS credentials are short lived and exist only in the child AWS CLI process environment. InferCrane does not persist or return them. Workers retrieve their API key directly from Secrets Manager through the instance profile. EC2 is launched without a public IP. Cost is reported asunknown. InferCrane does not ship a live AWS pricing catalog and will not
infer cost from an instance-type name.
Qualification state
Hermetic contract tests cover create-response loss, adoption, replay, deletion, tag-scoped instance and volume inventory, private networking, immutable images, normalized provider failures, and credential redaction. The narrow real AWS evidence is recorded in AWS real-infrastructure evidence; inspectinfercrane integrations for
the current contract state.