Skip to main content
InferCrane can add control-plane evidence around an existing OpenAI-compatible workload without replacing it. Adoption is incremental and does not silently transfer lifecycle ownership.

Adopt an existing workload

observe-only records health and evidence but never publishes a route. traffic-managed may route through the existing target after its qualified runtime health check succeeds. Neither mode creates, scales, updates, or deletes the imported workload. lifecycle-managed is available only through an InferCrane Deployment binding. The URL, source, logical model and physical upstream model form an immutable adoption identity. A safe retry returns the same adoption; changing immutable identity fails with a conflict.

Diagnose an adapter qualification gap

Read the evidence as an intersection, not a checklist of independently implemented parts:
  • registered means an adapter exists; it is not qualification.
  • local-qualified proves the maintained hermetic contract only.
  • experimental or in qualification requires the named real environment evidence.
  • a model recipe proves immutable configuration and license metadata, not GPU fit or protocol behavior.
  • unknown capability or missing exact-combination evidence fails closed.
Use compatibility to identify the missing provider/runtime/mode/model/accelerator evidence. If lifecycle qualification is absent but the existing endpoint is healthy, keep observe-only; move to traffic-managed only after direct protocol qualification. Do not choose lifecycle-managed for an externally owned workload. For a candidate revision that already exists, follow the retry-safe provisioning and validation procedure. It reuses the persisted revision and provider identities, checks orphan inventory, and preserves the active route. Adoption is the fallback only when the workload is genuinely externally owned; it is not a way to hide a failed InferCrane lifecycle operation. Promote ownership explicitly after reviewing health evidence:
The next command publishes a route through InferCrane. Verify upstream model identity, every client protocol, authentication behavior, streaming termination, and the ownership boundary first. It does not transfer create, scale, update, or delete ownership and does not change the existing upstream URL; keep clients on that URL until the managed path passes qualification.

Inspect one request

Use the X-Request-Id response header as the lookup key. The JSON result reconstructs one persisted request without reading prompt or output content: For standalone managed replicas, target identifies the selected registered replica target. Some provider-native Serverless and external APIs do not expose a worker identity; InferCrane leaves that boundary unavailable instead of inventing a replica ID. Missing upstream timing or token fields are also reported as unavailable. content_recorded is always false: prompts, outputs, tool arguments, embedding vectors, and authorization values are not part of request evidence. Request lookup is tenant-scoped; an ID from another tenant is returned as not found rather than disclosed.

When the request ID is missing

InferCrane does not currently expose a broad request search endpoint. Without X-Request-Id, it cannot prove that one request record belongs to a particular client call. Do not infer an individual route from aggregate metrics. Preserve the incident window and inspect only evidence that can still be attributed safely:
Record endpoint, active/candidate plan, revision, event timestamps, aggregate sample window, status, and explicitly unavailable request-level fields. Keep the current active plan and policy unchanged: do not run endpoint promote, rollout promote, admission set, or another mutation whose decision depends on the missing correlation. If policy requires request-level auditability, the correct outcome is an unauditable incident and a stopped change, not a reconstructed request ID. Capture X-Request-Id at the application or ingress boundary before retrying the qualification. For slow or unreachable providers, uncertain buffered/streaming outcomes, duplicate-submission prevention, and provider ownership reconciliation, use the Provider outage runbook.

Deterministic Doctor

Doctor persists reproducible Evidence → Rule → Finding results. Current rules cover unavailable endpoint routes, statistically bounded elevated error rate, and queue-dominant latency. A healthy result explicitly reports that no deterministic issue is active; it does not guess a cause.

Signed alerts

Deliveries include InferCrane-Delivery, InferCrane-Timestamp, and InferCrane-Signature: v1=<HMAC-SHA256>. The signed input is <timestamp>.<exact request body>. Delivery is idempotent per policy/finding, retries are bounded, and public-network validation rejects private, loopback, link-local and unspecified destinations by default. Resolved signing secrets are never persisted or returned.
Adoption does not import provider credentials or guarantee ownership of the workload. Keep the existing workload available until traffic-managed health and request evidence have been verified.