Start with one command
observe resolves either a stable endpoint or a deployment and presents the highest-value persisted
evidence together. Endpoint output includes its logical model, environment, active/candidate serving
plans, Release Guard, admission policy, and alerts. Deployment output includes desired and observed
lifecycle, traffic, the active operation, Guard decision, and recent events.
--watch redraws only when state changes and can be closed safely;
the control-plane operation continues.
Reconnect to durable work
Ask for a deterministic diagnosis
X-Request-Id header. Request Inspector ties it
to the logical endpoint, serving plan, binding, deployment, revision, selected target, latency,
tokens, retry, and fallback evidence where the upstream exposes those boundaries. See the complete
Request Inspector field contract; prompt and
output content are not persisted.
--diagnose is the one intentionally mutating form: it persists a fresh Doctor evaluation derived
from metrics, events, provider observations, and revision state. InferCrane does not ask an LLM to
invent the cause.
Trace a slow async request through the control plane
InferCrane uses distinct durable handles; it does not pretend one trace ID spans boundaries a provider does not expose:
Start with the symptom’s handle and walk outward:
- client/gateway rejection: an OpenAI-compatible status appears before upstream work;
- async deadline or cancellation: the durable job is terminal even if an already transmitted provider request does not stop immediately;
- provider delay: the operation remains waiting with a stable provider identity and ordered stage;
- runtime failure: provider capacity exists, but health/model identity or request evidence fails;
- scaling delay: a persisted decision exists, while desired and ready replica counts differ.
Choose a reversible recovery action
Diagnose first, then act on the persisted finding rather than treating every latency increase as a deployment failure:
Example recovery loop:
rollout inspect identifies a retained prior revision. Otherwise keep
the diagnosis and use a new immutable candidate; do not restart or delete capacity merely because a
single request was slow. See Release Guard recovery.