Useful in minutes, without a migration

scripts/demo-connect.sh in this repository against
GPU-free local fixtures. It proves the product workflow, not real-runtime performance.
Connect an existing endpoint in observe-only mode. InferCrane performs bounded discovery and records
the result without publishing a route or taking provider ownership.
connect is intentionally conservative. Runtime detection is reported only when the endpoint
provides grounded signals. Unknown capability remains unknown rather than being inferred from a URL.
Qualify before routing traffic
Observe-only discovery does not prove optional runtime behavior. Before changing application code or promoting routing ownership:unknown or unqualified, retain observe-only ownership.
Compare the existing and managed request paths
Use the same non-sensitive fixture first against the current upstream, then through InferCrane. The application remains on its existing URL during this comparison.data: [DONE]. Record direct latency and TTFT
outside InferCrane; observe-only mode does not proxy the request and cannot attribute that direct
call to a managed request record.
Approve cost and ownership before migration
Observe-only connection creates no provider capacity and sends no application traffic, but the existing workload continues to incur whatever infrastructure cost its owner already pays. InferCrane does not import that provider bill automatically. Before approving client migration, record one comparable window:
Use Cost and performance for comparable evidence. Unknown price,
currency, coverage, provider identity, or data policy remains
unavailable. The approver must be the
team accountable for provider spend and data handling; InferCrane does not turn missing cost into a
savings estimate. Keep clients on the original URL and retain observe-only ownership until those
inputs and the direct-versus-managed compatibility check pass.
[DONE], then inspect the managed
X-Request-Id. To test a delayed provider, inject a bounded delay only in a staging upstream you
control and repeat both paths; compare direct elapsed time with InferCrane queue_ms, ttft_ms, and
latency_ms. A provider-hidden delay remains combined upstream time, not a fabricated allocation or
model-load boundary.
Stop and keep application clients on the original URL if the managed path changes a required
protocol, model identity, cancellation behavior, or error contract; if evidence is missing; or if
the delayed request exceeds the approved timeout. Traffic-managed ownership is not lifecycle
ownership, so provider recovery still occurs in the existing owning system.
Turn evidence into a daily workflow
Evidence → Rule → Finding logic. Request Inspector reconstructs the
logical endpoint, resolved target, revision or upstream where known, queue and response timing,
tokens, retry count, and fallback reason. Prompt and output content are not recorded by default.
After this comparison passes, applications may move to the stable InferCrane endpoint. The workload
remains externally managed: InferCrane does not create, scale, update, or delete it.
Use a user-managed LiteLLM gateway
LiteLLM can sit behind InferCrane as an optional external OpenAI-compatible gateway:Why this pattern matters
- Platform teams can observe one workload before adopting a fleet.
- vLLM operators can retain their existing compute and container setup.
- LiteLLM users can retain broad managed-provider translation while InferCrane owns durable operational evidence and logical endpoint identity.
- Migration-sensitive teams can progress from observe-only to traffic-managed ownership without transferring provider lifecycle.