Skip to main content

Integration model

InferCrane is a portable production inference platform, not a wrapper around one cloud. Its durable control plane owns deployment, revision, operation, Release Guard, explanation, and telemetry semantics. Those semantics do not belong to RunPod, SkyPilot, or vLLM. External systems enter through capability-specific contracts.
Before selecting an adapter, run the exact-combination compatibility check:
All three commands are read-only. If the release evidence does not name the complete provider, runtime, compute mode, immutable model, accelerator, and environment combination, treat it as unqualified before committing money or a migration plan.

Keep sensitive inputs on self-hosted infrastructure

Use a serving plan that contains only a self-hosted deployment or a directly adopted self-hosted target. Do not add an external provider connection or an external fallback binding. For a new InferCrane-managed deployment, qualify the exact combination before creating billable capacity. plan is read-only; deploy is not:
Stop if those results do not identify the intended adapter, runtime, compute mode, immutable model, accelerator, environment, and required real-system evidence. Otherwise deploy and wait for the durable operation to reach readiness before binding it:
For a vLLM or SGLang endpoint you already operate, begin read-only and transfer traffic ownership only after inspection:
The first plan is a single manual binding, so there is no fallback destination that can receive request content. Prompt and output bodies are not recorded by default; verify network controls, runtime logging, and the adopted service itself according to your security policy.
Do not use a LiteLLM route with external upstreams for a strict self-hosted-only requirement. InferCrane can govern the connection to a user-managed LiteLLM gateway, but LiteLLM owns routing inside that gateway. InferCrane cannot prove that its internal routes keep data local. An endpoint binding that uses --acknowledge-external-data --enable-external explicitly permits request data to leave controlled infrastructure; request and cost limits bound usage, not data residency.
See Stable endpoints and serving plans for the binding model and LiteLLM gateway for the responsibility boundary. This authenticated read returns the compiled Provider Contract, Runtime Contract, and Composition Contract versions, registered adapters, capabilities, ownership boundaries, and separate local versus real-system qualification states. It never upgrades registration or hermetic simulation into a public support claim.

Registration is not qualification

A control-plane process registers concrete adapters during startup. Durable replicas persist the adapter identity used to create them, so restart, rollback, and deletion return to the same implementation without guessing from a cloud name. The release support matrix is a separate policy. An adapter can exist in development without being advertised publicly. Qualification requires configuration, documentation, compatibility records, failure testing, real infrastructure acceptance, and zero leaked billable resources.
The current adapter registry includes RunPod elastic/serverless, narrow AWS EC2, GCP Compute, and Kubernetes elastic and optional Dynamo serving-graph adapters; vLLM, SGLang, custom OCI, and governed external targets; plus LiteLLM, external-sandbox access, and signed-artifact-handoff composition profiles. The compatibility matrix qualifies only exact combinations; registration never implies real-system evidence.

Adding an integration

  1. Implement only the narrow capability contract the external system owns.
  2. Give the adapter a stable durable identity and register it at process composition.
  3. Map trustworthy observations into InferCrane’s normalized state and telemetry.
  4. Add configuration, diagnostics, documentation, and deterministic failure behavior.
  5. Qualify the exact cloud/runtime/compute-mode combination with real lifecycle evidence.
The versioned contract details are documented in Provider Contract V1, Runtime Contract V1, and Integration ownership. See Ecosystem compatibility for the exact user path and maturity of model APIs, gateways, agents, vector databases, workflow engines, training systems, sandboxes, runtimes, Kubernetes, and GPU schedulers. Do not add a provider conditional to a generic workflow, expose registration as support, build a second scheduler, or silently fabricate unavailable provider data. Registration and qualification remain separate states throughout the API, CLI, console, and documentation.