Surfaces
The current vLLM protocol documentation
lists these upstream surfaces, but availability depends on both runtime version and model task.
InferCrane therefore treats each capability independently. An unknown or unsupported claim returns
422 unsupported_protocol before any request reaches the workload.
Client compatibility contract
InferCrane qualifies HTTP protocol behavior, not every release of every OpenAI-compatible SDK. Use the InferCrane/v1 base URL, an InferCrane tenant credential, and the stable endpoint name as
model. Disable automatic client retries while qualifying streaming or failure behavior.
The current pinned vLLM profile qualifies Chat Completions, text Completions for compatible models,
and Embeddings for embedding models. Responses and online batch remain gateway-implemented but
unqualified for that profile. LiteLLM or another upstream may accept more parameters, but InferCrane
does not inherit those claims automatically. For any other client library, first run the same
buffered, streaming, error-schema, cancellation, and model-identity fixtures in staging; otherwise
the safe configuration is only the subset proven by the HTTP contract above.
Model identity
Send the stable InferCrane endpoint name asmodel. The gateway selects a pinned route generation
and rewrites only that field to the binding’s upstream model identity.
Qualify tool calling and streaming
Protocol registration is not proof that one runtime/model pair implements every optional behavior. Use an isolated staging endpoint backed by the exact candidate revision; do not send qualification traffic to the production alias. First inspect the registered runtime claim and candidate identity:curl -N disables client buffering; the stream must produce
well-formed SSE data and a terminal [DONE] without replaying a partial request:
X-Request-Id header from each response and inspect the persisted attribution:
Qualify streaming, cancellation, quotas, and retries together
Run this only against an isolated staging tenant and endpoint. First persist bounded admission and record the exact identity under test:- Send a successful buffered request and verify
X-InferCrane-Attemptnever exceeds the configured retry budget. Only connection failures and502,503, or504on InferCrane-managed capacity are retryable. Queue wait, retry backoff, every attempt, and response delivery must all fit inside the single endpointrequest_timeout_ms; retries never receive a fresh deadline. - Start the
curl -Nstreaming request above, capture its response headers, and terminate the client while chunks are arriving. Inspect itsX-Request-Id. The recorded request must remain one request, and no second stream may appear. Client cancellation propagation to the exact runtime is real-runtime evidence, not something a local fake proves. - Issue concurrent buffered requests beyond
max-concurrency + max-queue; excess work must receive an OpenAI-compatible rejection within the configured queue timeout rather than silently exceeding admission. - Exhaust the staging tenant’s request ceiling within one UTC minute. The next request must return
429before upstream transmission. Correlate the available request IDs and thequota.updateaudit event.
endpoint promote or rollout promote. Preserve the candidate and persisted evidence and leave
Release Guard INCONCLUSIVE.
Telemetry and privacy
Every accepted surface records the protocol operation, logical endpoint, selected binding, deployment/revision where known, provider/runtime dimensions, latency, status, token counts where the upstream reports them, and streaming errors. Prompt, response, embedding vectors, and tool arguments are not persisted by default.Current boundary
The default vLLM candidate is pinned to0.22.1. Responses and online chat batch remain unknown,
not advertised as qualified, until the pinned image passes protocol conformance and real GPU
qualification. Runtime availability is not treated as product evidence automatically.