Skip to main content

Evidence-gated optimization

InferCrane separates proposing a configuration from proving it. The optimizer starts with reviewed model metadata and locally or really qualified provider/runtime compatibility. When the exact optional estimator toolchain is available, it uses AIConfigurator to reduce the candidate space. Otherwise it creates conservative reviewed catalog candidates. Neither path contacts a control plane or provisions paid capacity. Check the optional estimator once:
Install its exact compatible tuple in an isolated Python 3.11–3.13 environment:
The plotext pin is required: AIConfigurator 0.11.0 is not compatible with plotext 6.0.0 even though its published dependency range permits it. optimize doctor validates the complete tuple, not just the top-level package.
The default --source auto tries the pinned AIConfigurator adapter and visibly falls back to the reviewed catalog when the optional tool is unavailable. Use --source aiconfigurator to require the estimator or --source catalog for deterministic reviewed starting profiles only. The result is a bounded candidate set. Each spec pins:
  • the exact model revision;
  • runtime engine and version;
  • provider adapter, region, and accelerator;
  • reviewed runtime arguments;
  • compute mode, scale bounds, and routing strategy.

Start from a public workload prior

When customer traffic is not available yet, choose one of InferCrane’s immutable content-free public priors. The default interactive lane is available directly to the optimizer:
Automation clients can discover the same versioned inputs with GET /api/v1/workload-profiles. Each result includes a ready-to-submit benchmark_request, its exact optimizer binding, source provenance, and the evidence digest. Available lanes cover interactive, long-prefill, and decode-heavy shapes. The public trace chooses token-shape experiments only. It does not supply a customer arrival rate, quality target, or production SLO, so InferCrane marks it promotion_eligible: false. Every candidate requires a representative customer replay before qualification or promotion. Use the remote Modal profiler in the public trace PoC only to reproduce or update the embedded artifact; normal product use does not download or rescan the dataset.

Start with any pinned open-weight model

A model does not need to be in the reviewed catalog to enter the proof loop. Supply its repository, immutable commit, intended GPU count, and one or more executable runtimes:
For an uncataloged model InferCrane creates only an empty-argument vLLM or SGLang baseline. It does not infer architecture support, memory fit, tokenizer behavior, licensing, quantization, speculative decoding, or kernel compatibility. Those gaps are listed as required evidence. The optional AIConfigurator adapter may reduce this generic candidate space, but its output remains modeled and unqualified. A mutable branch such as main is rejected; the revision must be a 40–64 character hexadecimal commit. The command refuses to overwrite a non-empty candidate directory. Use --output json to archive the input digest, source, compatibility evidence, required evidence, and limitations with the specs.
A proposal is not a benchmark result or recommendation. AIConfigurator rows are marked modeled and proposed-modeled-unqualified; catalog rows are marked unmeasured and proposed-unmeasured. InferCrane never presents an estimate as measured or qualified evidence.

The simple and advanced paths

The common path needs an objective, not a collection of runtime flags:
Advanced users can bind the search to an observed content-free Replay digest, choose runtimes, control the workload profile and concurrency, require the estimator, and inspect the complete JSON provenance. Provider credentials are not passed to the estimator. Model metadata network access is off by default and requires --allow-model-metadata-network.

Create a durable campaign

The authenticated console exposes the same boundaries under Optimize:
  1. choose a reviewed model, objective, infrastructure target, and optional SLO;
  2. review the candidate boundary without allocating provider resources;
  3. create an immutable campaign;
  4. grant a hard spend cap with a maximum 24-hour expiry;
  5. follow durable execution from Activity, even after closing the browser;
  6. inspect modeled and measured evidence separately;
  7. explicitly publish a qualified new endpoint or promote a Release Guard-approved evolution.
Cancellation is durable too: it persists the decision and links to the cleanup operation. The console never treats a proposed endpoint name as published until activation completes. Use propose for a completely offline preview. Use create to persist the same immutable proposal without provisioning anything:
That default creates a new_endpoint campaign. To optimize a candidate revision for an endpoint that already serves production, bind the campaign to its stable deployment identity:
InferCrane keeps these proof paths distinct. A new endpoint has no production baseline, so a winning candidate may become qualified after measured performance, sourced cost, and signed semantic quality evidence. It then waits for explicit human activation. An evolve_endpoint campaign must create a candidate revision against the named deployment, compare it with active production through Release Guard, and wait for explicit human promotion. InferCrane never invents a baseline for a new endpoint and never lets an evolution bypass Release Guard. Activation is itself a durable, idempotent operation. A new-endpoint campaign may activate only a qualified candidate. It atomically publishes the endpoint name from the immutable proposal and binds it to the selected isolated candidate deployment; an existing alias is never silently rebound. The activation operation waits until the request-path route generation is published. An evolution campaign may promote only a guard_passed candidate and delegates the traffic mutation to the normal guarded rollout workflow. Approval expiry is checked before mutation, while a retry after a completed activation remains a successful no-op. Approval first verifies every candidate against a fresh sourced price for its exact cloud, region, GPU, and maximum replica count. It then records a bounded authority window and queues one durable execution operation. The worker composes the existing deployment/rollout lifecycle, sends the versioned AIPerf workload directly to the exact candidate revision, adopts revision-bound signed quality evidence, ranks persisted results, evaluates Release Guard for evolution campaigns, and cleans rejected candidates. It never promotes automatically. The execution capability is fail-closed. The control plane must find AIPerf and must have an exact price snapshot valid for the entire approval window. Configure price evidence before approving:
The rate is the total hourly cost of that exact maximum-replica tuple. A missing tuple, expired source window, absent AIPerf binary, or approval below the worst-case execution cost is rejected before provider mutation. Operators must refresh price evidence rather than treating this static example as current pricing. Automation clients can request the same non-mutating candidate set with POST /api/v1/optimization/proposals. This endpoint performs no provider mutation and makes no performance claim. Pass the returned immutable proposal to POST /api/v1/optimization/campaigns when you are ready to persist the review boundary. Paid execution still requires the separate approval endpoint. The restart-safe execution path is locally qualified for lost child-operation responses, cancellation, expired authority, evidence identity, cleanup, and the human promotion boundary. Canceling a campaign queues durable cleanup even when its execution operation already finished. Closing the CLI does not stop the control-plane operation; reattach with infercrane operation watch. A qualified candidate keeps that parent operation in a bounded waiting state. If explicit activation or promotion does not happen before authority expires, InferCrane marks the candidate failed and cleans its resources rather than allowing an unbounded GPU bill. Execution uses a measured-evidence barrier: all viable candidates reach AIPerf measurement and revision-bound quality validation before ranking begins. Inference Lab then compares only the exact same workload shape. Proposal order cannot choose the winner. Missing sourced cost, mismatched workloads, or incomplete metrics produce inconclusive; explicit SLO or cost violations produce rejected. For evolution campaigns, the selected candidate still has to pass active-versus-candidate Release Guard and wait for human promotion. The persisted campaign links benchmark, quality, Lab, and, where a real baseline exists, Release Guard evidence separately.

Plan an optimized artifact

Builders run outside InferCrane. The common path produces a strict immutable plan:
Available plan presets cover FP8, AWQ, GPTQ, NVIDIA-specific NVFP4, EAGLE-3, MTP, DFlash, and TensorRT engine builds. A preset is not a compatibility or speed claim. Every plan requires an exact tool version, digest-pinned builder, hardware constraints, immutable output attestation, exact candidate binding, and passing semantic quality evidence before qualification. Advanced users can provide the complete JSON plan with --plan-file.

Propose, prove, release

Use the generated files through the existing lifecycle:
When model precision or model artifacts differ, attach signed semantic evaluation evidence before promotion. When cost is part of the objective, attach sourced, dated cost evidence. Only an exact model, runtime, accelerator, provider, and canonical workload result supports a measured recipe.

Seed candidates from NVIDIA Model Optimizer

InferCrane can discover reviewed NVIDIA Model Optimizer checkpoints as provenance-bound candidate seeds. Each seed pins the publisher repository, immutable revision, artifact identity, algorithm, precision, and source URL before it enters the search space. The console shows that provenance and the remaining qualification gates. These checkpoints are unmeasured candidates, not InferCrane performance claims. InferCrane does not claim that it transformed the publisher artifact or that the artifact is faster on a target GPU. Exact runtime and accelerator compatibility, correctness, semantic quality, workload performance, and cost evidence remain mandatory before Release Guard can accept a candidate. A publisher checkpoint that fails any gate is retained as negative evidence rather than promoted.

What the first planner knows

The first source composes: Accelerator matching uses concrete hardware identity. Generic Kubernetes resource names such as nvidia.com/gpu are capacity resources, not GPU models, and cannot qualify a mechanism. Explicit aliases such as NVIDIA-L40S normalize to the exact hardware family. Dynamo candidates compile to one InferCrane-owned parent graph with outer autoscaling fixed at one; Dynamo owns its internal workers. This prevents two controllers from scaling the same serving graph. It does not infer that a model fits an accelerator. It does not automatically enable quantization, speculative decoding, LMCache, disaggregated serving, FlashInfer, or TensorRT-LLM. Those mechanisms are version- and workload-specific and require their own compatibility and qualification evidence. Use --include-simulated only for development exploration; simulated candidates are excluded by default.

Upstream composition boundaries

InferCrane will reuse upstream execution and optimization work behind replaceable adapters: The durable rule is: the optimizer proposes; InferCrane proves. Candidate sources, builders, and execution systems remain replaceable. Modeled results cannot satisfy a measured gate, provider mutation requires bounded approval, optimized artifacts retain immutable provenance, and activation or promotion remains an explicit action.

Multi-lane optimization evidence

External optimization workers can emit the content-free infercrane.dev/optimization-evidence/v1 contract defined in schemas/optimization-evidence-v1.schema.json. It keeps three benchmark evidence levels distinct:
  • screening is inexpensive candidate-search evidence (at least 12 requests per lane);
  • qualification requires at least 100 requests, two independent runs, and 300 seconds per lane;
  • public requires at least 300 requests, three independent runs, and 600 seconds per lane, and is reserved for independently repeated baseline-versus-final receipts.
These schema floors prevent evidence relabeling; a workload policy may require larger samples. They do not retroactively promote the existing 12-request Qwen measurements, which remain screening evidence. Every campaign pins one model, hardware tuple, and workload digest. Every candidate adds an immutable serving-recipe digest and every required concurrency lane. Each lane records aggregate output throughput, request throughput, SLO goodput in requests/second, and SLO-qualified output tokens/second as separate metrics. InferCrane applies error, prompt-integrity, quality, sample-size, TTFT, ITL, and SLO-attainment gates to every lane before calculating a score. An ineligible candidate has no score, even when another lane has the highest raw throughput. Evaluate a secret-free document locally without provisioning infrastructure:
The result pins the evaluator algorithm, input digest, policy/version, rejection reasons, Pareto set, and selected candidate. The supported policies cover qualified output throughput, SLO request goodput, qualified throughput per dollar, minimum COGS under the SLO, and latency-sensitive serving. FastPath provisional handoffs have a separate inspection boundary:
Inspection validates immutable model/tokenizer revisions, tokenizer and chat-template hashes, runtime commit, digest-pinned image, CUDA and hardware identity, launch revision, evidence level, quality declaration, and every workload lane. The result remains external_unverified inside InferCrane until the control plane binds it to its own immutable revision and passing signed quality evidence. A community source remains recorded even after InferCrane independently reproduces its recipe; community measurements never become qualification evidence.

Benchmarks and competitor comparisons

AIConfigurator estimates help decide what to test; they are not publishable performance claims. Use the competitive benchmark methodology to compare InferCrane-managed serving with Baseten, Fireworks, or another provider on the same model artifact, precision, request dataset, protocol, concurrency, region class, warm/cold definition, and output token accounting. If an external provider does not expose the same artifact or configuration, label the result as a product comparison rather than an engine comparison.