Skip to main content

CLI reference

The public CLI uses Cobra for grouped help, typo suggestions, aliases, and shell completion. It talks to the authenticated control-plane API and never opens PostgreSQL directly. Run infercrane help for the current command tree and infercrane version for the build version.

Global conventions

  • Commands that return data accept --output human|json.
  • status --watch --output json emits one JSON document per state refresh.
  • Mutation commands accept --idempotency-key; generated keys are printed in human and JSON output and attached to uncertain-request errors. Reuse the same key after an uncertain result.
  • --wait polls persisted operation state and prints only changed progress. Closing the client does not cancel server-side work; the operation ID and exact infercrane operation watch ID resume command are printed before waiting begins.
  • --wait-timeout 20m bounds the local watcher without cancelling the durable operation. Resume with infercrane operation watch ID; cancel only with infercrane operation cancel ID.
  • JSON failures contain code, category, message, retryable, remediation, and provider HTTP status when available.
  • Command help is available before authentication: infercrane COMMAND --help never contacts the control plane.
  • infercrane --context NAME COMMAND selects a context for one invocation.
  • infercrane --no-color COMMAND and the standard NO_COLOR environment variable disable ANSI styling. Redirected output and JSON are never decorated.

Configure and diagnose

init

Validate and store a control-plane URL and an already-issued credential in a named private client context.
init verifies authentication through the read-only identity endpoint before writing configuration. Use --skip-check only when intentionally configuring an offline control plane. It does not create or enroll a principal.

Contexts and authentication

JSON output uses stable lowercase fields (id, tenant_id, name, role, kind, and scopes) so authentication checks can be consumed directly by scripts without depending on Go field names. Legacy single-context configuration is migrated on the next init. Context listing and display never print stored credentials.

Shell completion

Homebrew installs generated completions automatically. Completion is read-only and may suggest deployment names from the configured control plane; an unavailable control plane produces no completion error or mutation.

doctor

Ask the control plane to check its own dependencies.
--cloud adds SkyPilot credentials and RunPod advisory accelerator availability; --serverless adds RunPod Serverless credentials and template validation; --aws performs a read-only role assumption and dependency probe for the configured AWS BYOC adapter; --gcp checks Application Default Credentials, every configured Compute dependency, immutable images, and Private Google Access; --kubernetes checks the explicit context, optional KServe CRD, and required namespaced permissions. Diagnostics are read-only and do not create provider resources. The capability table distinguishes supported, unsupported, and unknown. InferCrane never silently changes hardware, and never claims cache or fast-resume behavior that an adapter cannot observe.

Local accelerator discovery

Inspect NVIDIA GPU inventory without installing a runtime or changing the host:
Discovery reports its evidence source and limitations. It does not infer runtime compatibility, model fit, or performance from a device name alone.

Run the control plane

serve starts the control plane and inference gateway from validated environment configuration. Production deployments should use the container and service guidance in Production operations, including PostgreSQL, TLS, bootstrap authentication, stable instance identity, and graceful shutdown.

Operations workspaces

inbox

The inbox reads tenant-scoped endpoint and deployment summaries, ranks non-serving state and staged candidates deterministically, and prints the exact persisted field behind each item. It fails the whole read if either fleet API is unavailable, so a partial inventory cannot look healthy. It does not run Doctor implicitly and never reads prompt or response content.

observe

The default snapshot is read-only. Endpoint snapshots combine logical identity, serving plans, Guard, admission, and alerts; deployment snapshots combine lifecycle, traffic, active operation, Guard, and events. --diagnose explicitly persists a fresh deterministic Doctor evaluation.

telemetry collect dcgm

The command parses a bounded, documented subset of DCGM Prometheus metrics locally and sends normalized content-free evidence through the authenticated control API. Use --file SNAPSHOT for offline input. The snapshot is immutable, active-revision-bound, and becomes stale at its declared TTL; arbitrary Prometheus metrics are not accepted.

finops collect opencost

The command selects exact OpenCost allocation keys, records a currency-explicit average hourly rate against the active immutable revision, and creates a persisted FinOps report. Multiple exact keys may be comma-separated. --allocation and --currency are mandatory; cluster-wide attribution and currency conversion are never inferred. Use --file allocation.json for offline evidence.

artifact

Prefetch records durable intent; only a fresh provider observation is cache-hit evidence.

evaluation

ingest accepts the strict evaluator-result v1 interchange contract, rejects unknown content fields, binds it to one immutable revision, signs it, and optionally attaches it. Evidence is signature-verified and contains aggregate values only.

mcp

Starts a stdio MCP server with read-only deployment, endpoint, request, operation, and curated-recipe tools. No mutation tool is registered. See MCP for coding agents.

Admission and async inference

async submit persists protocol-native request content only after explicit encrypted-storage consent, returns a durable job ID, and does not depend on the client remaining connected. See Admission and async inference for retry, retention and webhook rules.

Terminal

The terminal workspace is reconnectable, supports state-valid guarded actions, and can be forced read-only with infercrane ui --read-only. It requires an interactive terminal and never uses tmux for persistence. See Terminal operations workspace.

Plan and deploy

optimize propose

Generate inspectable candidate DeploymentSpecs without contacting a control plane or provider. Check the optional modeled-candidate source first:
The default --source auto uses the pinned AIConfigurator adapter only when its exact supported Python tuple is available and visibly falls back to the reviewed catalog otherwise. Use --source catalog for a deterministic offline proposal, or --source aiconfigurator to require the estimator and fail rather than falling back.
Catalog output is unmeasured; AIConfigurator output is modeled. Neither is qualified. Use --output json to retain the input digest, source version, modeled-result digest, compatibility evidence, missing evidence, and limitations. --workload-fingerprint can bind a content-free observed workload identity to the proposal. Simulated compatibility is excluded unless --include-simulated is explicit. The command refuses to overwrite a non-empty output directory. The reviewed catalog spans instruction, reasoning, coding, embeddings, multimodal, and enterprise-text workloads. A profile for one model is never evidence for another model, precision, GPU, runtime, or topology. See Evidence-gated optimization.

optimize create

Persist the proposal and candidate identities without provider mutation, then grant bounded paid authority explicitly:
Approval retries do not extend the first persisted expiry. Creation and inspection allocate no infrastructure. The default new_endpoint intent stops at human activation after qualification. --target-deployment selects evolve_endpoint, which must stop at Release Guard and human promotion. optimize activate is the explicit boundary in both cases: it marks the qualified new deployment active, or delegates an evolution to the existing guarded rollout operation. Retrying the command adopts the same durable operation; InferCrane never promotes from the optimizer loop.

artifact optimize

Plan and attest an externally built optimized checkpoint, speculator, or TensorRT engine:
Presets are fp8, awq, gptq, nvfp4, eagle3, mtp, dflash, and tensorrt. Builders are never executed inside the InferCrane API process, and no optimized artifact qualifies without exact candidate, revision, digest, and semantic-quality evidence.

workload

Create a repository-local inference project, validate it without allocating a GPU, and use the same DeploymentSpec for build, plan, and deployment:
workload init --recipe NAME --profile PROFILE pins a reviewed model commit and writes the selected candidate’s runtime version, immutable workload when present, GPU type and count, provider hint, replica bounds, and routing policy. The default profile is the recipe’s general starting point; profiles are not benchmark claims. workload build --tag IMAGE performs a local Docker Buildx build for custom OCI projects; add --push to record a registry-confirmed immutable digest. workload dev binds a custom OCI workload only to loopback. See Inference projects.

plan

Preview semantic changes without provider or database mutation.
Provisioned and serverless plans include an explicit readiness-evidence boundary:
This is intentionally not a time estimate. If the tenant has at least three successful historical observations for the exact provider adapter, runtime, region, and GPU, plan instead shows an observed p50 from durable replica intent through runtime readiness. It shows p95 only after twenty successes and labels the sample count and boundary. That is historical evidence, not a capacity guarantee. For a DeploymentSpec pinned to an immutable model revision, plan can additionally show per-stage observed p50 after three successes and p95 after twenty only when model commit, runtime version and arguments, provider, region, GPU type/count, and compute mode all match. A stage with no provider marker or insufficient samples remains absent/unknown. Once an operation starts, durable events replace unknown stages with fresh provider and runtime observations where those boundaries are actually exposed.

deploy and apply

deploy creates a cloud or existing-target deployment. apply declaratively converges a deployment using the same input shape.
A YAML path cannot be combined with deployment flags. See DeploymentSpec. Custom OCI uses the YAML form because its runtime contract is part of the immutable revision. Advanced Dynamo topology also uses YAML. See SGLang, NVIDIA Dynamo, and Custom OCI workloads. After submission, human output prints the exact commands to follow durable progress. A successful wait prints the command for the first inference request.

integrations

Displays registered provider, runtime, and external composition capabilities, evidence state, and the exact qualified runtime/provider/compute-mode combinations. Registration alone never implies production support.

sandbox

Compose an externally operated sandbox with one stable inference endpoint:
connect and rotate reveal a credential once. It expires within 24 hours, can invoke only the selected endpoint alias, and cannot use control-plane routes. revoke disables InferCrane access without mutating the external sandbox. See External agent sandboxes.

training

Verify a signed, content-free artifact handoff from an external training system:
The private signing key remains local. Attachment binds immutable identity to one revision but does not promote it or claim that InferCrane executed training. See Training artifact handoffs.

Send a request

Use the logical deployment name without assembling an HTTP request by hand:
--output json returns the OpenAI-compatible choices and usage object for non-streaming requests. The command uses the configured endpoint and credential; applications can continue to use any OpenAI-compatible SDK directly. --protocol accepts chat, responses, embeddings, completions, or batch; the endpoint must explicitly qualify the selected surface. Inspect the persisted routing and timing evidence for a returned X-Request-Id without retrieving prompt or output content:

Stable endpoints

Separate the name applications call from the deployment that currently serves it:
endpoint plan makes the first plan active and stages later plans as candidates. Inspect and promote explicitly:
--bindings is ordered and accepts optional bounded weights, for example primary:80,overflow:20. See stable endpoints and serving plans. Authenticated external APIs add provider-neutral policy flags to endpoint bind:
The control plane rejects raw credentials, missing consent, absent hard limits, cross-tenant secret references, and adapter/target mismatches before creating the binding. infercrane provider list shows reusable connection metadata without resolving credentials. infercrane provider delete NAME removes only the reusable configuration; immutable endpoint bindings and their release history are deliberately unchanged. --from-env creates or reuses a secret reference by variable name, making a disconnected retry idempotent without accepting the raw key. Use --secret-reference ID instead when secret references are provisioned separately. The two flags are mutually exclusive. Stage one environment’s active immutable plan as another endpoint’s candidate:
The first command is a non-mutating preview. Staging never switches production traffic; destination Release Guard and explicit endpoint promotion remain required.

Connect an existing workload

The simple path verifies the upstream model list through the control-plane API and starts in observe-only mode:
Use --model when the endpoint exposes more than one physical model. Use --manage-traffic only when the endpoint has been qualified and should immediately enter the InferCrane logical route. The advanced ownership-compatible command remains available:
Use traffic-managed only when InferCrane should publish the healthy existing target. Neither mode transfers provider lifecycle ownership. See Adopt and diagnose.

Existing targets and routing

Supported strategies are round-robin, consistent-hash, power-of-two, and cache-aware, delegated to the pinned vLLM Router.

Observe and explain

Long operations belong to the control plane, not the terminal. deploy, apply, rollout, scale, and deletion continue after the CLI disconnects. Pressing Ctrl-C stops only the local watcher. Run infercrane operation watch ID later to resume from persisted progress; add --wait-timeout 20m to bound only the new local watch. Explicit infercrane operation cancel ID is the separate action that requests cooperative server-side cancellation and provider cleanup. status deliberately separates two concerns. Serving answers whether the current route can accept inference traffic. Convergence answers whether desired capacity, a rollout, or deletion is still progressing. A deployment can therefore be serving · converging while a provider allocates another replica; this is not reported as an outage. JSON output exposes the same stable fields under lifecycle_status, including ready and desired replica counts, provisioning and draining counts, candidate state, and the blocking durable operation. Deterministic explanations use only persisted state and measurements:

Models and recipes

Explore the built-in reviewed catalog without contacting a control plane:
models inspect prints the immutable revision, license URL, protocol capabilities, serving profiles, accelerator topology, immutable runtime workload when present, and the exact qualification boundary. Create a project with infercrane workload init --recipe NAME. Catalog profiles are configuration-verified, not benchmark, price, capacity, or universal GPU-fit claims. See Verified Models.

Benchmark

Run AIPerf and persist the exact reproduction metadata.
--profile accepts balanced, buffered, interactive, long-context, long-generation, overload, or throughput. Profiles are immutable workload shapes; explicit request, concurrency, token, and streaming flags override their defaults. --revision accepts active, candidate, or an explicit revision ID. Without a profile, defaults remain 100 requests, concurrency 10, and seed 17. Declared TTFT, TPOT, or total-latency SLOs produce measured goodput. Capture and search immutable recipes only after a matching benchmark exists:
Compare persisted measured configurations without provisioning new capacity:
The objective can be interactive, latency, throughput, or cost-efficiency. InferCrane recommends only across an exact common workload digest; mixed shapes remain UNRANKED. See Recipes and Inference Lab for the provenance and missing-evidence contract.

Replay and capacity

Replay captures content-free production shape by default. Running an AIPerf approximation requires both --execute and --acknowledge-cost. See Replay and capacity intelligence. Create or inspect a best-effort logical session hint when an application benefits from affinity:
The returned Context Passport is an opaque, expiring routing hint. It never promises durable KV state, and health always overrides a stale preferred target or binding.

Revisions and Release Guard

Candidate creation also accepts --model-revision, --runtime, --runtime-version, --runtime-args, --routing, and --region. --file preserves a complete validated DeploymentSpec, including advanced serving topology, and cannot be mixed with candidate override flags. Promotion remains policy-gated; no LLM decides the outcome. rollout validate is explicit synthetic traffic, not shadowing. It runs the existing AIPerf adapter against active and candidate revisions using the same workload and persisted hard bounds. The acknowledgement is mandatory because both runs may incur provider cost.

Signed release evidence

passport keygen and passport verify work without a configured control-plane context. Issue and list use the authenticated API. See Inference Passports.

Delete

Preview cleanup, then confirm it explicitly.
--plan is side-effect-free. --yes is required for mutation. After a paid test, also verify the provider inventory reaches zero.

Tenant administration

These commands require an admin/bootstrap credential.
New and rotated tokens are returned once. Service-account scopes can only restrict the role ceiling; they can never grant an action unavailable to viewer, operator, or admin. Omitting --scopes uses every action allowed by the selected role for compatibility.

Secret references

Register metadata that points to an injected environment value. InferCrane never accepts the raw secret on the command line or through its API.

Signed alerts

Alerts contain deterministic Doctor findings and HMAC headers; delivery is idempotent and bounded.

Governed external fallback

External targets require manage_external. They are selected only when ordinary targets are unhealthy and require an explicit privacy acknowledgement plus hard budgets.
See Governed external capacity for transmission and reservation semantics. Evaluate a bounded overflow decision without allowing a second controller to mutate replicas:
Burst Guard requires fresh health, repeated pressure, and sourced cost within the hard limit. Its result is a policy decision; the configured mutation owner remains responsible for execution.

SLO policy and recommendations

Define explicit fail-closed thresholds, then evaluate persisted benchmark evidence. Recommendations are advisory and never mutate the deployment.
recommended identifies a qualified candidate satisfying the policy. no_match means measured candidates violate it. unknown means required evidence is missing. See Inference decisions. Advisory Autopilot turns that evidence into a reviewable plan without mutating production:
Approval records human intent. It does not bypass Release Guard, evidence requirements, hard budgets, or the single mutation-owner boundary.