Release Guard
Release Guard is a deterministic, persisted comparison between an active revision and its candidate. It never uses an LLM to decide whether a candidate should be promoted.--evaluate compares the active and candidate plans and
never promotes traffic. Use infercrane endpoint inspect coder-production to retrieve plan IDs
before a separate, explicit promotion.
For deployment revisions, promotion is an authenticated, traffic-changing operation and requires the
latest persisted evaluation to be ACCEPT for the same candidate and still-current active revision:
promote atomically changes the active revision and target generation, then drains the old
generation. Validation traffic may incur the explicit cost and privacy impact acknowledged by
rollout validate; promotion itself does not duplicate requests. Release Guard still requires real
active/candidate evidence for the exact environment.
An evaluation returns one of three decisions:
ACCEPT: enough evidence exists and every measured regression is within policy.REJECT: readiness failed, proven compatibility mismatches, or a measured regression exceeds policy.WAIT: the candidate is ready but there is not enough trustworthy evidence yet.
infercrane rollout inspect DEPLOYMENT renders the latest persisted active/candidate comparison as
a metric table and shows unavailable measurements explicitly. JSON output retains the complete
policy, measurements, reason codes, revision identities, and evaluation timestamp.
Missing measurements are not fabricated. Output throughput is compared only when both revisions report token usage. TTFT is required before acceptance. A candidate with no healthy ready replica is rejected immediately.
Endpoint Release Guard never applies one primary deployment’s metrics to an arbitrary routing
graph. It currently qualifies single-primary comparisons, a primary change with an unchanged
primary-fallback graph, and an append-only managed fallback whose immutable policy includes
explicit privacy consent and hard request/cost reservations. Weighted changes or other unmeasured
topology changes persist INCONCLUSIVE with serving_plan_topology_unqualified; they cannot be
promoted until per-binding evidence exists.
Infrastructure can be healthy while model behavior has regressed. For quality-sensitive changes,
run a task-specific customer-owned evaluation and attach its
signed aggregate evidence. Release Guard can require comparable
suite/evaluator versions, a minimum score, and a bounded regression. InferCrane still does not select
an LLM judge, retain evaluation prompts or outputs, or let an evaluator promote a revision.
Because a candidate is deliberately absent from the logical route, operators gather bounded
candidate evidence explicitly with AIPerf rather than duplicating production requests:
This implements candidate → measure → gate → active without silently replaying live prompts.
Teams that require production-shape validation can use content-free Inference Replay fingerprints or
run their own reviewed evaluator dataset against the isolated candidate. InferCrane does not make
silent shadowing the default: raw prompts may be sensitive, duplicate calls add cost, and tool calls
may have side effects.
rollout create persists an immutable candidate and returns its revision ID. provision creates
revision-scoped capacity without adding it to the active route. Configure policy, run the bounded
validation, and inspect the persisted result. Execute promote only when the decision is ACCEPT
for that same candidate and active revision; WAIT and REJECT are not authorization to move
traffic.
Release Guard uses a persisted active/candidate benchmark pair only when tool version and workload
parameters match. The evaluation snapshot records aiperf_benchmark plus both benchmark IDs as its
evidence source. Otherwise it waits for comparable evidence instead of mixing measurements.
Change model, provider, and scaling bounds in one candidate
Replica bounds are immutable revision configuration. A candidate withmin < max enables bounded
autoscaling, but Release Guard validation alone does not prove a real provider can complete
1 → N → 1 or drain a long stream. Qualify that dynamic behavior first on an isolated staging
deployment with the exact intended model, runtime, provider, GPU, region, and bounds.
candidate-staging and retain its benchmark, stream, scaling-decision, provider-inventory, and
cleanup evidence.
min=1,max=4; the Guard snapshot must refer to that candidate and the still-current active revision.
Topology evidence is INCONCLUSIVE for unmeasured weighted graphs or arbitrary binding changes.
Reject on mismatch or regression. Promote only on ACCEPT, then keep the automatic rollback window
or operator recovery ready and re-observe the real scaling policy. If policy requires dynamic
autoscaling proof before any production promotion, the matching staging qualification is mandatory;
Release Guard does not silently substitute a synthetic benchmark for it.
InferCrane does not silently duplicate inference requests. rollout validate prints a cost/privacy
notice, enforces persisted request and concurrency ceilings, runs AIPerf against each revision, and
then queues the durable evaluation. It does not retain prompts or generated output.
Retry provisioning or validation safely
Transient capacity, transport, or readiness failures do not authorize a second candidate. Keep the same deployment, candidate revision ID, and durable operation:rollout provision re-observes and adopts capacity for the same immutable revision. Do not run
rollout create again or change provider identity after an uncertain response. Compare the
persisted provider resource IDs with read-only provider inventory before retrying when ownership is
ambiguous.
Validation sends explicit inference to both revisions and may incur provider cost and expose the
validation workload to the configured runtime. Retry it only after reviewing the persisted policy
ceilings and acknowledging that boundary again:
REJECT, followed by
candidate rejection and candidate-only cleanup, not a loop that creates new capacity. Preserve the
Guard evaluation, operation timeline, and provider identities until direct inventory confirms no
orphan.
When automatic rollback is enabled, promotion retains the previous healthy capacity until the
persisted observation monitor reaches ACCEPT or REJECT. The monitor snapshots its policy and
deadline at creation, so a concurrent policy edit cannot silently weaken an in-flight decision. Rejection atomically restores the old
revision and target set, waits for its router generation, drains active streams safely, and deletes
only failed revision capacity. A restart resumes the same monitor and deadline.
Recover manually and preserve evidence
An evaluationREJECT does not move traffic. Record the decision before cleaning candidate
capacity, then reject only that candidate:
rollout inspect; never guess it or delete the
deployment to recover a release.
After acceptance, issue an Inference Passport to make the exact
revision and release evidence independently verifiable.
For CI collection and provider-outcome failure semantics, use the
release evidence runbook.
For one reviewer-facing quality, performance, stream, recovery, rollback, and cleanup gate, use the
production release approval checklist.