Skip to main content

Start with one command

observe resolves either a stable endpoint or a deployment and presents the highest-value persisted evidence together. Endpoint output includes its logical model, environment, active/candidate serving plans, Release Guard, admission policy, and alerts. Deployment output includes desired and observed lifecycle, traffic, the active operation, Guard decision, and recent events.
The default command is read-only. --watch redraws only when state changes and can be closed safely; the control-plane operation continues.

Reconnect to durable work

Closing a laptop or losing a network connection stops the client watcher, not the deployment. The operation ID is the resume handle. Retrying a mutation with the original idempotency key is safe when the request result was uncertain.

Ask for a deterministic diagnosis

The request ID comes from the inference response’s X-Request-Id header. Request Inspector ties it to the logical endpoint, serving plan, binding, deployment, revision, selected target, latency, tokens, retry, and fallback evidence where the upstream exposes those boundaries. See the complete Request Inspector field contract; prompt and output content are not persisted. --diagnose is the one intentionally mutating form: it persists a fresh Doctor evaluation derived from metrics, events, provider observations, and revision state. InferCrane does not ask an LLM to invent the cause.

Trace a slow async request through the control plane

InferCrane uses distinct durable handles; it does not pretend one trace ID spans boundaries a provider does not expose: Start with the symptom’s handle and walk outward:
An async job can remain queued before it has an inference request ID. A provider allocation can be pending without a physical replica, and some Serverless APIs never reveal a worker identity. Keep those fields unavailable. Distinguish outcomes this way:
  • client/gateway rejection: an OpenAI-compatible status appears before upstream work;
  • async deadline or cancellation: the durable job is terminal even if an already transmitted provider request does not stop immediately;
  • provider delay: the operation remains waiting with a stable provider identity and ordered stage;
  • runtime failure: provider capacity exists, but health/model identity or request evidence fails;
  • scaling delay: a persisted decision exists, while desired and ready replica counts differ.
If a required handle is missing, preserve the remaining time window and use the missing-request-ID procedure. Do not join unrelated records by timestamp alone or report a provider-hidden boundary as measured. For automation, use the fail-closed release evidence CI runbook. It captures stable JSON request, operation, deployment/revision, Guard, orphan, and direct provider outcomes and rejects a release when any required boundary is unavailable.

Choose a reversible recovery action

Diagnose first, then act on the persisted finding rather than treating every latency increase as a deployment failure: Example recovery loop:
Rollback is valid only when rollout inspect identifies a retained prior revision. Otherwise keep the diagnosis and use a new immutable candidate; do not restart or delete capacity merely because a single request was slow. See Release Guard recovery.
Use infercrane ui for a keyboard-driven operations workspace and the separately released operations console for the browser view. Both reconstruct state from the same control-plane API; neither is the source of truth.