Adopt an existing workload
observe-only records health and evidence but never publishes a route. traffic-managed may route
through the existing target after its qualified runtime health check succeeds. Neither mode creates,
scales, updates, or deletes the imported workload. lifecycle-managed is available only through an
InferCrane Deployment binding.
The URL, source, logical model and physical upstream model form an immutable adoption identity. A
safe retry returns the same adoption; changing immutable identity fails with a conflict.
Diagnose an adapter qualification gap
registeredmeans an adapter exists; it is not qualification.local-qualifiedproves the maintained hermetic contract only.experimentalorin qualificationrequires the named real environment evidence.- a model recipe proves immutable configuration and license metadata, not GPU fit or protocol behavior.
unknowncapability or missing exact-combination evidence fails closed.
observe-only; move to traffic-managed only after direct
protocol qualification. Do not choose lifecycle-managed for an externally owned workload.
For a candidate revision that already exists, follow the
retry-safe provisioning and validation procedure.
It reuses the persisted revision and provider identities, checks orphan inventory, and preserves the
active route. Adoption is the fallback only when the workload is genuinely externally owned; it is
not a way to hide a failed InferCrane lifecycle operation.
Promote ownership explicitly after reviewing health evidence:
Inspect one request
X-Request-Id response header as the lookup key. The JSON result reconstructs one persisted
request without reading prompt or output content:
For standalone managed replicas,
target identifies the selected registered replica target. Some
provider-native Serverless and external APIs do not expose a worker identity; InferCrane leaves that
boundary unavailable instead of inventing a replica ID. Missing upstream timing or token fields are
also reported as unavailable.
content_recorded is always false: prompts, outputs, tool arguments, embedding vectors, and
authorization values are not part of request evidence. Request lookup is tenant-scoped; an ID from
another tenant is returned as not found rather than disclosed.
When the request ID is missing
InferCrane does not currently expose a broad request search endpoint. WithoutX-Request-Id, it
cannot prove that one request record belongs to a particular client call. Do not infer an individual
route from aggregate metrics.
Preserve the incident window and inspect only evidence that can still be attributed safely:
endpoint promote, rollout promote, admission set, or another mutation whose decision
depends on the missing correlation. If policy requires request-level auditability, the correct
outcome is an unauditable incident and a stopped change, not a reconstructed request ID. Capture
X-Request-Id at the application or ingress boundary before retrying the qualification.
For slow or unreachable providers, uncertain buffered/streaming outcomes, duplicate-submission
prevention, and provider ownership reconciliation, use the
Provider outage runbook.
Deterministic Doctor
Evidence → Rule → Finding results. Current rules cover unavailable
endpoint routes, statistically bounded elevated error rate, and queue-dominant latency. A healthy
result explicitly reports that no deterministic issue is active; it does not guess a cause.
Signed alerts
InferCrane-Delivery, InferCrane-Timestamp, and
InferCrane-Signature: v1=<HMAC-SHA256>. The signed input is
<timestamp>.<exact request body>. Delivery is idempotent per policy/finding, retries are bounded,
and public-network validation rejects private, loopback, link-local and unspecified destinations by
default. Resolved signing secrets are never persisted or returned.