Skip to main content

AIPerf

InferCrane delegates load generation to AIPerf and persists the workload, runtime, model artifact, provider, GPU, revision, results, and exact reproduction command.
Use --revision candidate only when you intend to benchmark isolated candidate capacity. Benchmark data remains in the InferCrane control plane; it is not uploaded by default. InferCrane pins AIPerf 0.12.0, requests record-only exports, and can drive constant, Poisson, or gamma arrival timing. Independent runs produce 95% confidence intervals in persisted workload evidence; they are not flattened into a single unexplained number.

Benchmarking

Understand measurements, persisted reproduction metadata, and evidence limits.