Skip to main content

Recipes and Inference Lab

Need candidate configurations before measuring? Start with Evidence-gated optimization. The optimizer proposes immutable, explicitly unmeasured DeploymentSpecs; this page covers the measured evidence and comparison stages that follow. Recipes turn one real deployment revision into a content-addressed, reproducible record. InferCrane will only capture a recipe when the active revision has an immutable ModelArtifact, an explicit runtime version, and a matching persisted AIPerf benchmark.
The recipe payload records the immutable model commit, runtime/version/arguments, provider adapter, region, GPU, compute mode, workload contract, benchmark workload, and benchmark ID. Provenance records the deployment/revision, artifact, benchmark tool/version, InferCrane version, and evidence class. The SHA-256 digest covers both canonical objects. Retrying the same name/version is idempotent only for identical content; different content is rejected. The CLI and private browser console expose the reviewed catalog through infercrane models and Models. Selecting a serving profile prefills the model, immutable revision, and runtime while writing its candidate runtime arguments, accelerator hint, compute mode, replica bounds, and routing strategy into an editable DeploymentSpec. Provider and region remain operator choices. A catalog entry remains configuration-verified until that exact serving plan has measured benchmark evidence. See Verified Models for the complete evidence boundary.

Compare evidence

Inference Lab v3 compares existing benchmark history. It does not provision hardware or run another load generator:
Choose interactive, latency, throughput, or cost-efficiency. Interactive and latency minimize measured p95 TTFT, throughput maximizes output tokens per second, and cost efficiency maximizes measured output tokens per second per sourced hourly currency unit. Rows include runtime, provider, GPU, error rate, TTFT, TPOT, goodput, throughput, constraint result, and cost metadata exactly as persisted. Region, GPU-count, error-rate, goodput, throughput, and sourced hourly-cost constraints fail closed when the required evidence is absent. Current Lab results emit only MEASURED rows. The schema reserves modeled and heuristic labels for future evidence sources, but InferCrane never substitutes them silently. With no comparable benchmark, the result is an empty persisted evaluation. Lab recommends a candidate only when every displayed candidate has the same exact canonical workload digest. Mixed concurrency or token shapes are shown as UNRANKED; filter with --profile and, when needed, --workload-digest. A profile is a convenient filter, while the digest is the final comparability boundary. Prompt and output content are not part of recipes, benchmark history, or Lab evaluations.

Qualification boundary

A captured recipe proves provenance, not universal portability. Provider/runtime support remains scoped to the integration matrix, and local benchmark fixtures are not real GPU evidence. Cost remains unavailable unless a trustworthy source and observation timestamp were persisted with the benchmark.