Evidence-gated optimization
InferCrane separates proposing a configuration from proving it. The optimizer starts with reviewed model metadata and locally or really qualified provider/runtime compatibility. When the exact optional estimator toolchain is available, it uses AIConfigurator to reduce the candidate space. Otherwise it creates conservative reviewed catalog candidates. Neither path contacts a control plane or provisions paid capacity. Check the optional estimator once:plotext pin is required: AIConfigurator 0.11.0 is not compatible with plotext 6.0.0 even
though its published dependency range permits it. optimize doctor validates the complete tuple,
not just the top-level package.
--source auto tries the pinned AIConfigurator adapter and visibly falls back to the
reviewed catalog when the optional tool is unavailable. Use --source aiconfigurator to require the
estimator or --source catalog for deterministic reviewed starting profiles only.
The result is a bounded candidate set. Each spec pins:
- the exact model revision;
- runtime engine and version;
- provider adapter, region, and accelerator;
- reviewed runtime arguments;
- compute mode, scale bounds, and routing strategy.
Start from a public workload prior
When customer traffic is not available yet, choose one of InferCrane’s immutable content-free public priors. The default interactive lane is available directly to the optimizer:GET /api/v1/workload-profiles.
Each result includes a ready-to-submit benchmark_request, its exact optimizer binding, source
provenance, and the evidence digest. Available lanes cover interactive, long-prefill, and
decode-heavy shapes.
The public trace chooses token-shape experiments only. It does not supply a customer arrival rate,
quality target, or production SLO, so InferCrane marks it promotion_eligible: false. Every
candidate requires a representative customer replay before qualification or promotion. Use the
remote Modal profiler in the public trace PoC only to
reproduce or update the embedded artifact; normal product use does not download or rescan the
dataset.
Start with any pinned open-weight model
A model does not need to be in the reviewed catalog to enter the proof loop. Supply its repository, immutable commit, intended GPU count, and one or more executable runtimes:main is rejected; the revision must be a 40–64 character
hexadecimal commit.
The command refuses to overwrite a non-empty candidate directory. Use --output json to archive the
input digest, source, compatibility evidence, required evidence, and limitations with the specs.
The simple and advanced paths
The common path needs an objective, not a collection of runtime flags:--allow-model-metadata-network.
Create a durable campaign
The authenticated console exposes the same boundaries under Optimize:- choose a reviewed model, objective, infrastructure target, and optional SLO;
- review the candidate boundary without allocating provider resources;
- create an immutable campaign;
- grant a hard spend cap with a maximum 24-hour expiry;
- follow durable execution from Activity, even after closing the browser;
- inspect modeled and measured evidence separately;
- explicitly publish a qualified new endpoint or promote a Release Guard-approved evolution.
propose for a completely offline preview. Use create to persist the same immutable proposal
without provisioning anything:
new_endpoint campaign. To optimize a candidate revision for an endpoint
that already serves production, bind the campaign to its stable deployment identity:
qualified after measured performance, sourced cost, and signed semantic
quality evidence. It then waits for explicit human activation. An evolve_endpoint campaign must
create a candidate revision against the named deployment, compare it with active production through
Release Guard, and wait for explicit human promotion. InferCrane never invents a baseline for a new
endpoint and never lets an evolution bypass Release Guard.
Activation is itself a durable, idempotent operation. A new-endpoint campaign may activate only a
qualified candidate. It atomically publishes the endpoint name from the immutable proposal and
binds it to the selected isolated candidate deployment; an existing alias is never silently rebound.
The activation operation waits until the request-path route generation is published. An evolution
campaign may promote only a guard_passed candidate and delegates the traffic mutation to the
normal guarded rollout workflow. Approval expiry is checked before mutation, while a retry after a
completed activation remains a successful no-op.
Approval first verifies every candidate against a fresh sourced price for its exact cloud, region,
GPU, and maximum replica count. It then records a bounded authority window and queues one durable
execution operation. The worker composes the existing deployment/rollout lifecycle, sends the
versioned AIPerf workload directly to the exact candidate revision, adopts revision-bound signed
quality evidence, ranks persisted results, evaluates Release Guard for evolution campaigns, and
cleans rejected candidates. It never promotes automatically.
The execution capability is fail-closed. The control plane must find AIPerf and must have an exact
price snapshot valid for the entire approval window. Configure price evidence before approving:
POST /api/v1/optimization/proposals. This endpoint performs no provider mutation and makes no
performance claim. Pass the returned immutable proposal to
POST /api/v1/optimization/campaigns when you are ready to persist the review boundary. Paid
execution still requires the separate approval endpoint.
The restart-safe execution path is locally qualified for lost child-operation responses,
cancellation, expired authority, evidence identity, cleanup, and the human promotion boundary.
Canceling a campaign queues durable cleanup even when its execution operation already finished.
Closing the CLI does not stop the control-plane operation; reattach with infercrane operation watch. A qualified candidate keeps that parent operation in a bounded waiting state. If explicit
activation or promotion does not happen before authority expires, InferCrane marks the candidate
failed and cleans its resources rather than allowing an unbounded GPU bill.
Execution uses a measured-evidence barrier: all viable candidates reach AIPerf measurement and
revision-bound quality validation before ranking begins. Inference Lab then compares only the exact
same workload shape. Proposal order cannot choose the winner. Missing sourced cost, mismatched
workloads, or incomplete metrics produce inconclusive; explicit SLO or cost violations produce
rejected. For evolution campaigns, the selected candidate still has to pass
active-versus-candidate Release Guard and wait for human promotion. The persisted campaign links
benchmark, quality, Lab, and, where a real baseline exists, Release Guard evidence separately.
Plan an optimized artifact
Builders run outside InferCrane. The common path produces a strict immutable plan:--plan-file.
Propose, prove, release
Use the generated files through the existing lifecycle:Seed candidates from NVIDIA Model Optimizer
InferCrane can discover reviewed NVIDIA Model Optimizer checkpoints as provenance-bound candidate seeds. Each seed pins the publisher repository, immutable revision, artifact identity, algorithm, precision, and source URL before it enters the search space. The console shows that provenance and the remaining qualification gates. These checkpoints are unmeasured candidates, not InferCrane performance claims. InferCrane does not claim that it transformed the publisher artifact or that the artifact is faster on a target GPU. Exact runtime and accelerator compatibility, correctness, semantic quality, workload performance, and cost evidence remain mandatory before Release Guard can accept a candidate. A publisher checkpoint that fails any gate is retained as negative evidence rather than promoted.What the first planner knows
The first source composes:- reviewed model recipes;
- the executable integration compatibility inventory;
- reviewed vLLM configuration profiles for balanced, interactive, and throughput workloads;
- exact runtime versions and provider adapters.
nvidia.com/gpu are capacity resources, not GPU models, and cannot qualify a mechanism. Explicit
aliases such as NVIDIA-L40S normalize to the exact hardware family. Dynamo candidates compile to
one InferCrane-owned parent graph with outer autoscaling fixed at one; Dynamo owns its internal
workers. This prevents two controllers from scaling the same serving graph.
It does not infer that a model fits an accelerator. It does not automatically enable quantization,
speculative decoding, LMCache, disaggregated serving, FlashInfer, or TensorRT-LLM. Those mechanisms
are version- and workload-specific and require their own compatibility and qualification evidence.
Use --include-simulated only for development exploration; simulated candidates are excluded by
default.
Upstream composition boundaries
InferCrane will reuse upstream execution and optimization work behind replaceable adapters:
The durable rule is: the optimizer proposes; InferCrane proves. Candidate sources, builders, and
execution systems remain replaceable. Modeled results cannot satisfy a measured gate, provider
mutation requires bounded approval, optimized artifacts retain immutable provenance, and activation
or promotion remains an explicit action.
Multi-lane optimization evidence
External optimization workers can emit the content-freeinfercrane.dev/optimization-evidence/v1 contract defined in
schemas/optimization-evidence-v1.schema.json.
It keeps three benchmark evidence levels distinct:
screeningis inexpensive candidate-search evidence (at least 12 requests per lane);qualificationrequires at least 100 requests, two independent runs, and 300 seconds per lane;publicrequires at least 300 requests, three independent runs, and 600 seconds per lane, and is reserved for independently repeated baseline-versus-final receipts.
external_unverified inside
InferCrane until the control plane binds it to its own immutable revision and passing signed quality
evidence. A community source remains recorded even after InferCrane independently reproduces its
recipe; community measurements never become qualification evidence.