Skip to main content
InferCrane exposes one stable logical endpoint while preserving the selected runtime’s native OpenAI-compatible request and response shape. It does not translate every workload into a lossy universal schema.

Surfaces

The current vLLM protocol documentation lists these upstream surfaces, but availability depends on both runtime version and model task. InferCrane therefore treats each capability independently. An unknown or unsupported claim returns 422 unsupported_protocol before any request reaches the workload.

Client compatibility contract

InferCrane qualifies HTTP protocol behavior, not every release of every OpenAI-compatible SDK. Use the InferCrane /v1 base URL, an InferCrane tenant credential, and the stable endpoint name as model. Disable automatic client retries while qualifying streaming or failure behavior.
The current pinned vLLM profile qualifies Chat Completions, text Completions for compatible models, and Embeddings for embedding models. Responses and online batch remain gateway-implemented but unqualified for that profile. LiteLLM or another upstream may accept more parameters, but InferCrane does not inherit those claims automatically. For any other client library, first run the same buffered, streaming, error-schema, cancellation, and model-identity fixtures in staging; otherwise the safe configuration is only the subset proven by the HTTP contract above.

Model identity

Send the stable InferCrane endpoint name as model. The gateway selects a pinned route generation and rewrites only that field to the binding’s upstream model identity.
All other protocol fields pass through. Upstream status, content type, response body, cancellation, and streaming semantics are preserved. InferCrane never retries a partially transmitted stream.

Qualify tool calling and streaming

Protocol registration is not proof that one runtime/model pair implements every optional behavior. Use an isolated staging endpoint backed by the exact candidate revision; do not send qualification traffic to the production alias. First inspect the registered runtime claim and candidate identity:
Then exercise a buffered tool call. The response must be valid OpenAI-compatible JSON and contain the tool-call structure expected by your application; a healthy HTTP status alone is insufficient:
Exercise streaming independently. curl -N disables client buffering; the stream must produce well-formed SSE data and a terminal [DONE] without replaying a partial request:
Capture the X-Request-Id header from each response and inspect the persisted attribution:
See the Request Inspector field contract for which routing and timing boundaries remain explicit or unavailable. Record the InferCrane release, endpoint serving-plan digest, deployment revision, immutable model commit, runtime version, container digest, provider, accelerator, request fixture, response schema, and streaming outcome. Repeat under the representative load and cancellation behavior required by your production application before promotion.
These commands prove only the exact runtime, model, revision, and environment tested. Local fake runtimes prove gateway control flow, not real tool selection, CUDA execution, streaming behavior, or model quality. Keep Release Guard INCONCLUSIVE when required real-runtime evidence is absent.

Qualify streaming, cancellation, quotas, and retries together

Run this only against an isolated staging tenant and endpoint. First persist bounded admission and record the exact identity under test:
Use the tenant quota API procedure to set a small staging-only requests-per-minute ceiling without guessing the existing production quota. Then perform four independent checks:
  1. Send a successful buffered request and verify X-InferCrane-Attempt never exceeds the configured retry budget. Only connection failures and 502, 503, or 504 on InferCrane-managed capacity are retryable. Queue wait, retry backoff, every attempt, and response delivery must all fit inside the single endpoint request_timeout_ms; retries never receive a fresh deadline.
  2. Start the curl -N streaming request above, capture its response headers, and terminate the client while chunks are arriving. Inspect its X-Request-Id. The recorded request must remain one request, and no second stream may appear. Client cancellation propagation to the exact runtime is real-runtime evidence, not something a local fake proves.
  3. Issue concurrent buffered requests beyond max-concurrency + max-queue; excess work must receive an OpenAI-compatible rejection within the configured queue timeout rather than silently exceeding admission.
  4. Exhaust the staging tenant’s request ceiling within one UTC minute. The next request must return 429 before upstream transmission. Correlate the available request IDs and the quota.update audit event.
Record a qualification result containing InferCrane version, endpoint serving-plan digest, candidate revision, immutable model, runtime/container version, provider, accelerator, admission policy, tenant ceiling, request fixture, response codes, attempt counts, stream cancellation result, and request evidence. Passing local fixture tests proves admission and accounting state machines; only the same procedure against the real runtime proves cancellation after upstream transmission. If any request cannot be correlated, a partial stream is replayed, the retry ceiling is exceeded, the endpoint deadline is reset between attempts, quota accounting cannot be demonstrated, or required real-runtime evidence is absent, do not run endpoint promote or rollout promote. Preserve the candidate and persisted evidence and leave Release Guard INCONCLUSIVE.

Telemetry and privacy

Every accepted surface records the protocol operation, logical endpoint, selected binding, deployment/revision where known, provider/runtime dimensions, latency, status, token counts where the upstream reports them, and streaming errors. Prompt, response, embedding vectors, and tool arguments are not persisted by default.

Current boundary

The default vLLM candidate is pinned to 0.22.1. Responses and online chat batch remain unknown, not advertised as qualified, until the pinned image passes protocol conformance and real GPU qualification. Runtime availability is not treated as product evidence automatically.