Skip to main content

A model is enough to start

InferCrane turns a model reference and a serving plan into a durable deployment. The CLI submits the intent to the control plane and can disconnect safely while provider allocation, artifact transfer, runtime startup, readiness, and route publication continue.
deploy can create billable provider resources. Start with plan, complete the selected provider setup and qualification review. InferCrane does not fabricate a price when the provider has not supplied trustworthy cost evidence.
For the first real deployment, the selected adapter’s setup path must include credential mounting, production control-plane startup, read-only preflight, durable reconnection, teardown, and direct provider-inventory verification. RunPod is only one independently qualified adapter.
If you want a reviewed immutable starting point first:
Catalog entries make repository, commit, license, protocol, and configuration evidence visible. They are not performance or provider-availability claims. See Verified Models. Qwen3-8B is intentionally used as the release acceptance baseline: it is small enough for frequent single-GPU qualification and exercises the production lifecycle without making every test unusually expensive. It is not embedded in the InferCrane domain model.

Use the model your workload needs

The same model-to-endpoint path accepts other Hugging Face repository identities. These examples show portable input syntax; they do not claim that every model fits every GPU or that every adapter has been production-qualified for it.
plan validates InferCrane’s serving intent without creating resources. Before deploy, confirm the model card license or access grant, total and active parameter memory, quantization, tensor-parallel topology, runtime version, multimodal requirements, and provider capacity. Frontier MoE models such as Kimi K3, Qwen3.8 A95B, and DeepSeek V4 require materially different topology from a compact dense model. Omitting a GPU flag is deliberate: use Inference Lab or an explicit qualified serving plan rather than presenting one accelerator as a sizing recommendation.
Model popularity changes faster than InferCrane releases. The source of truth is the immutable model artifact and the evidence attached to its serving plan, not a marketing list of “supported models.”
The short path uses configured defaults. Make infrastructure choices explicit when you need them:
The second command below creates provider capacity and can incur cost. The first command is the required side-effect-free review. For a deliberately small first qualification, keep exactly one replica with --min 1 --max 1, confirm the provider’s own price/stock view, and record the returned operation ID before running deploy. A replica ceiling is not a monetary budget.
plan is side-effect free. deploy returns a durable operation unless --wait is supplied. Closing the terminal stops only the local watch; use the returned operation ID to reconnect:

Diagnose readiness before changing the plan

Do not submit a second deployment when readiness fails. Preserve the operation and provider identity, then collect:
Retry the same operation or desired state only with the original idempotency key. After a terminal failure, choose between an immutable corrected candidate and plan-first deletion. Never erase PostgreSQL to clear readiness state. Cleanup is complete only when InferCrane shows no deployment or orphan and direct provider inventory proves zero run-owned resources. An uncertain create or delete must keep the same persisted replica identity and operation; never submit a replacement by name.

Clean up the exact deployment

Deletion can terminate billable provider capacity. First render the semantic plan, verify the deployment name and every persisted provider identity, then submit one durable delete operation:
Do not guess a resource ID, retry with a different idempotency key, or delete by provider display name. delete --plan is side-effect free; --yes authorizes termination. InferCrane inventory is necessary but not sufficient: use the selected provider guide to query its authoritative inventory read-only and prove that every run-owned pod, instance, endpoint, service, and attached billable resource reached zero. If provider deletion is eventually consistent, keep watching the same durable operation and inventory; do not create replacement capacity during cleanup. Unknown ownership is a stop condition: preserve durable state, retry only read-only observation, and never create, adopt, or delete until one authoritative identity or absence is proven. Capture canonical provider inventory before deployment and after deletion with the command for the selected adapter. In a shared account, require equality with the baseline rather than global zero; never delete another deployment’s resource:
Any authentication, authorization, parsing, pagination, or API failure makes cleanup unproven. An empty successful result is distinct from a command that could not see inventory.

Before opening a shared endpoint

A ready GPU is not the complete production boundary. Use a non-bootstrap, tenant-scoped credential, persist endpoint admission, set the distributed tenant request ceiling, and qualify only the protocols your exact runtime/model pair claims:
  1. Authentication and secret boundaries
  2. Admission and tenant quota procedure
  3. Protocol and cancellation qualification
  4. Autoscaling and active-stream qualification
  5. Request Inspector evidence
If authentication, quota exhaustion, request correlation, or a required protocol cannot be proven in the target environment, keep the endpoint private and the candidate unpromoted. Local fixture success does not upgrade an unqualified runtime/provider combination.

Send the first real request

Wait until status reports the deployment ready before sending application traffic. The application uses the control-plane URL as its stable OpenAI-compatible base URL and the deployment name as model; provider worker addresses remain infrastructure details:
Copy the X-Request-Id response header from cURL to prove which revision and target served it:
The real provider/runtime/model/GPU combination remains unqualified until its separate acceptance evidence passes; a successful single request is readiness evidence, not a production performance or reliability claim.
RunPod is the first real-infrastructure acceptance target, not the InferCrane domain model. Its elastic SkyPilot adapter remains experimental until the credentialed lifecycle and zero-leak gate is complete. AWS BYOC, Kubernetes, existing targets, and future providers implement the same provider contract with independent evidence. Check compatibility and qualification before selecting a backend.

Bring an immutable runtime when a model flag is not enough

For a qualified custom runtime, describe the OCI image, digest, startup contract, health contract, and protocol capabilities in a DeploymentSpec. InferCrane persists that configuration as an immutable revision and manages its lifecycle; it does not build images or execute arbitrary build scripts.
See Custom OCI workloads for the exact schema and safety requirements.

What “build” means here

This boundary keeps the product extensible: vLLM, SGLang, a custom OCI runtime, or another qualified runtime can change without changing the application-facing endpoint.

Separate development, staging, and production

Environments are first-class, tenant-scoped resources. Bind separate stable endpoints to the same logical model so application configuration and policy do not leak across stages:
Each endpoint has independent active and candidate serving plans. Request records preserve logical model, endpoint, and environment identity. Stage the exact serving plan from staging as a production candidate without changing production traffic:
The first command is a side-effect-free plan. The second atomically clones destination-scoped bindings and stages an immutable candidate. Production Release Guard remains the authority for activation. See Environment promotion.

Make the next revision safely

Use the same declarative path for updates:
InferCrane creates an immutable candidate and preserves the active revision until promotion policy is satisfied. Continue with the safe rollout showcase.