Claims with receipts
InferCrane binds every benchmark to an exact model, runtime, accelerator, workload, and date. A measured win is not generalized to another provider, GPU, context length, concurrency level, or production network path.Qwen3.8-27B on one H200
The paired run used the same model, image, workload, hardware, and output checks. Deterministic
outputs matched. Streaming and buffered usage, finish reasons, structured output, forced tools,
malformed-request handling, and prompt-token accounting passed.
Rejections are part of the result
The campaign preserved candidates that sounded attractive but lost under the target workload:- FlashInfer GDN paths preserved correctness but regressed throughput and cold readiness.
- DFlash2 won some lower-concurrency screening lanes but did not replace native MTP at the selected saturation boundary.
- The near-262K context run completed correctly but missed its latency SLO.
WAIT or REJECT; it is never converted into a release win.
Read benchmarking, optimization, and
compatibility and qualification for the evidence rules.