Skip to main content

SGLang

SGLang is InferCrane’s second registered inference engine. The runtime profile pins the official lmsysorg/sglang:v0.5.12 multi-platform manifest by digest and launches its OpenAI-compatible server through the same provider-neutral workload contract used by custom OCI images. The image tag and launch form follow the official SGLang release and Docker documentation. InferCrane resolved the manifest digest during implementation; the immutable value is visible through infercrane integrations --output json and the persisted revision.
The built-in profile declares the immutable image and argv, so normal users do not repeat them. The full 0.5.12 image passed real AWS L40S readiness and model identity, buffered and streaming OpenAI transport, AIPerf qualification, immutable model-revision launch, and durable shutdown. Local conformance also covers client cancellation, connection draining, the metrics endpoint, and immutable launch intent. A smaller official runtime image was considered and then rejected by real AWS qualification because its published dependency set could not start the SGLang server. Tool calling, structured output, broader model compatibility, and other providers remain separate qualification surfaces. Autoscaling is rejected for SGLang because normalized SGLang scaling signals are not yet qualified.