Skip to main content

Connect LiteLLM without replacing it

LiteLLM and InferCrane own different problems. LiteLLM translates provider protocols and holds its upstream credentials. InferCrane owns the stable application endpoint, adoption state, persisted request evidence, and safe release decisions.
Before changing an application client, read the client compatibility contract. It is the source of truth for supported protocol surfaces, pass-through fields, InferCrane-interpreted fields, streaming, cancellation, and client retry settings. A LiteLLM connection does not automatically qualify every LiteLLM model, plugin, or OpenAI SDK behavior.

Connect in observe-only mode

InferCrane verifies the OpenAI-compatible discovery surface before persisting the connection. It does not install, configure, upgrade, or fork LiteLLM.
The private console exposes the same path under Create endpoint → Connect existing inference. Choose LiteLLM as the connector. Connections start observe-only; traffic ownership never transfers silently.

Promote traffic ownership explicitly

After you have inspected health and evidence:
The application continues calling:
Use max_retries=0 in Python and maxRetries: 0 in TypeScript while qualifying the route. For a stream, set stream=True or stream: true and consume the SDK’s streaming iterator to completion; InferCrane preserves SSE ordering and terminal [DONE], never retries a partial stream, and records the route under one request ID. Closing the client propagates cancellation best-effort, but whether the exact LiteLLM/upstream pair stops compute or billing is real-environment evidence.

Compatibility-safe client profile

SDKs differ in spelling, defaults, and retry behavior, but they all send the same HTTP contract. Start with this deliberately small profile and expand it only after qualifying the exact LiteLLM route. InferCrane does not publish a fictional SDK-by-SDK support claim for fields it merely passes through. The lowest-risk portable request contains only model, messages, and an application-bounded max_tokens. Parameter compatibility is route-level, not client-library-level: This is the important boundary: InferCrane can enumerate the capability gates it owns, but it cannot truthfully enumerate every rejected LiteLLM parameter without the exact LiteLLM version, route, plugin, provider, and model. Record that matrix from staging qualification rather than copying another route’s result. Use a binding known not to claim the requested capability to verify fail-closed behavior before traffic migration:
If that request reaches LiteLLM, if the error schema changes, or if any pass-through field behaves differently from the recorded staging fixture, keep traffic ownership unchanged and mark the route unqualified. The full protocol matrix and cURL, Python, and TypeScript examples also cover size limits, error schemas, tools, retries, and runtime-specific parameters.

Ownership and qualification

The OpenAI-compatible behavior is locally qualified with hermetic fixtures. Real LiteLLM versions, plugins, upstream providers, and credential policy still require qualification in your environment.

Why not bundle LiteLLM?

Bundling would make InferCrane responsible for another gateway’s security patches, license boundary, provider catalog, and upgrades. The composition contract keeps that dependency replaceable and lets teams use LiteLLM, another gateway, or InferCrane’s native data plane without changing the logical endpoint contract. The open-source tree outside LiteLLM’s separately identified enterprise area is MIT-licensed. The decision not to fork is operational rather than merely legal: provider integrations and security patches should continue arriving from their specialist owner. Any future managed-process adapter must pin an exact upstream image, publish an SBOM, qualify its behavior, and remain replaceable.