Use the right owner for each layer
InferCrane does not need to replace every component in an inference or agent stack. It provides the stable model endpoint, durable lifecycle, routing policy, and operational evidence. A user-managed gateway can translate provider protocols, and a sandbox can execute agent tools while calling that endpoint like any other application.Connect a user-managed LiteLLM gateway
Keep LiteLLM’s provider configuration and credentials in LiteLLM. Connect its OpenAI-compatible surface to InferCrane in observe-only mode first:Add managed APIs without changing application code
An authenticated OpenRouter or OpenAI-compatible API can be a first-class endpoint binding. Stage it as a fallback while the application continues sendingmodel="coder-production":
Use OpenRouter as deployment overflow
OpenRouter can also be optional deployment-level emergency capacity, not an implicit default. Configure it with a secret reference, explicit model mapping, privacy acknowledgement, and hard request and cost reservation ceilings. InferCrane selects it before sending bytes and never silently duplicates or shadows a request.Call InferCrane from an agent sandbox
The sandbox is the application execution environment. Create it with the sandbox provider, then issue a short-lived credential restricted to exactly one InferCrane endpoint:Bring a trained artifact into the release path
Keep training data and execution in MLflow, Kubeflow, SkyPilot, or your existing pipeline. Hand InferCrane a signed content-free artifact identity bound to one immutable candidate revision:Why the boundary is useful
- Change inference providers without rebuilding the agent sandbox.
- Change sandbox vendors without changing the logical model endpoint.
- Keep model rollout evidence separate from tool-execution security policy.
- Move externally trained artifacts into a guarded release without moving training data or keys.
- Avoid turning InferCrane into a gateway fork, workflow engine, or sandbox isolation runtime.