Build a production service safely
This recipe defines one logical model, separate staging and production endpoints, one version-controlled deployment, endpoint admission, a distributed tenant request ceiling, bounded autoscaling, and a fail-closed candidate test. It does not make an experimental provider/runtime combination production-qualified.1. Keep the serving intent in version control
Createcoder-runtime.yaml with the exact reviewed model, runtime, provider, region, GPU, and
min=1,max=2. Do not depend on implicit CLI defaults:
max_replicas: 2 is a capacity ceiling,
not a monetary budget. Stop if provider price, stock, quota, ownership, or required qualification is
unknown.
2. Create stable environment identities
Create the domain objects before provider mutation:model="coder-production"; they never need a provider address or revision
ID.
3. Deploy and bind one reviewed runtime
4. Bound new requests and tenant volume
Persist endpoint-wide admission first:PUT --data-binary @previous-quota.json; never reconstruct it from memory.
Qualify in a dedicated tenant just after a UTC-minute boundary: send exactly the approved nonzero
number of requests, require the next request to return 429 before upstream transmission, and
confirm the audited quota.update. Missing 429 or audit evidence blocks opening the endpoint.
5. Prove failure handling away from production
The GPU-free local proof first verifies control flow without cost:coder-staging-runtime deployment and bind it only
to coder-staging. Use a deliberately invalid candidate revision there, never invalid production
configuration and never a second deployment after an uncertain response:
REJECT; the staging active revision
must remain routed. Preserve evidence, then remove only the exact candidate:
6. Qualify autoscaling and active-stream drain
Against the exact valid staging combination, start one long SSE request and a bounded AIPerf load in separate terminals:1 → 2 → 1, fresh vLLM signals, no capacity above max=2, one
persisted provider identity per replica intent, an intact stream with terminal data: [DONE], one
request ID with no replay, generation-safe drain, full cooldown/recovery, and canonical direct
provider inventory returning to baseline after staging cleanup. Missing metrics or real capacity is
INCONCLUSIVE, not permission to promote. The stream/list behavior and provider timing must be
proven in the exact environment; local fake workers do not qualify it.
7. Open production only after real qualification
Run the selected provider’s credentialed product gate with an isolated account/project/namespace, least-privilege credentials, a fixed run ID, explicit paid-resource approval, and a recorded canonical inventory baseline:PROVIDER only with the selected maintained gate (runpod, aws, gcp, or kubernetes)
and satisfy that gate’s documented credential/network prerequisites first. The exact model commit,
runtime/container, provider adapter, region, accelerator, network, replica bounds, and InferCrane
release must match this service. Rerun a disconnected gate with the same run ID; a new ID risks
duplicate resources. The report must prove readiness, buffered/streaming protocols, cancellation,
benchmarking, durable recovery, deletion, empty InferCrane orphan state, and provider inventory equal
to baseline. Local fixtures, Kind, and provider HTTP simulations prove control logic but cannot prove
GPU readiness, IAM/networking, capacity, billing, or deletion semantics.
Before application traffic: