Reproducible benchmarking
InferCrane delegates load generation and measurement to AIPerf. It does not contain a second load generator.constant, poisson, and gamma arrival patterns model different request timing. The request rate
and pattern must be supplied together. InferCrane pools the record-level results, retains the run
count and load shape, and stores Student-t 95% intervals and coefficient of variation for
throughput, tail latency, and goodput when at least two runs contain the metric.
Profiles define load shape, not runtime tuning claims. interactive, balanced, throughput,
buffered, long-context, long-generation, and overload record their immutable
benchmark-profile-v1 identity inside the workload evidence. Explicit token or concurrency flags
override profile defaults and remain visible in the result.
For a bounded qualification sweep across concurrency 1, 8, 32, and 128 plus long prompt,
long-generation, streaming, buffered, and overload shapes:
.infercrane/performance/RUN_ID/matrix.json. Set
INFERCRANE_BENCHMARK_TTFT_SLO_MS, INFERCRANE_BENCHMARK_TPOT_SLO_MS, and/or
INFERCRANE_BENCHMARK_LATENCY_SLO_MS to compute measured goodput: successful requests per second
that satisfy every declared threshold. A zero threshold means no SLO was declared and goodput stays
unavailable rather than being invented.
The profile matrix deliberately changes workload shape by objective, so its rows are not a
scaling curve. Use the separate same-workload campaign when comparing concurrency behavior:
1, 8, 32, and 128. The resulting
.infercrane/performance/RUN_ID-concurrency-sweep/sweep.json is the only one of these two artifacts
that supports a concurrency scaling comparison. Override the bounded defaults with
INFERCRANE_BENCHMARK_SWEEP_* variables; every override remains in persisted workload evidence.
The CLI submits the benchmark through the authenticated control-plane API. The control plane resolves the active immutable revision, runs AIPerf against the logical InferCrane endpoint, and persists the result. A fixed dataset seed defaults to 17 and can be changed with --random-seed. AIPerf 0.12.0 is pinned so flag semantics and exports cannot drift between candidates.
Release Guard validation can benchmark an isolated candidate explicitly:
null. Goodput remains null without SLO thresholds.
Cost metadata becomes available only when fresh, revision-bound deployment_hourly_rate evidence
covers the run; otherwise it contains available: false and a reason. None are fabricated.
What “fast” means
InferCrane does not promise one globally fastest configuration. It measures the exact combination and lets the operator choose the objective:
Runtime arguments are immutable revision data. Tune a candidate, run the same profile against active
and candidate revisions, and let Release Guard compare matching evidence before promotion. This is
how InferCrane can provide a fast qualified default without turning an upstream vLLM or SGLang flag
into an unsupported universal claim.
For reviewed generation models, scaffold two candidate projects and preserve their different runtime
arguments explicitly:
RECOMMENDED only when candidates share one exact workload digest and the required
metric - and sourced hourly cost for cost efficiency - is present. Unlike workload shapes remain
UNRANKED.
Maintainers can qualify one pinned AWS model family at a time without rerunning unrelated runtime
paths:
records export level. This contains per-request measurements but not the raw request/response export, so prompt and generated content are not persisted. Results remain in the InferCrane PostgreSQL database and are never uploaded by default.
The control-plane image pins AIPerf 0.12.0. For a standalone control plane, install it with
pipx install aiperf==0.12.0; the benchmark adapter rejects another version rather than silently
mixing tool semantics.