A model is enough to start
InferCrane turns a model reference and a serving plan into a durable deployment. The CLI submits the intent to the control plane and can disconnect safely while provider allocation, artifact transfer, runtime startup, readiness, and route publication continue. For the first real deployment, the selected adapter’s setup path must include credential mounting, production control-plane startup, read-only preflight, durable reconnection, teardown, and direct provider-inventory verification. RunPod is only one independently qualified adapter.Use the model your workload needs
The same model-to-endpoint path accepts other Hugging Face repository identities. These examples show portable input syntax; they do not claim that every model fits every GPU or that every adapter has been production-qualified for it.plan validates InferCrane’s serving intent without creating resources. Before deploy, confirm the
model card license or access grant, total and active parameter memory, quantization, tensor-parallel
topology, runtime version, multimodal requirements, and provider capacity. Frontier MoE models such
as Kimi K3, Qwen3.8 A95B, and DeepSeek V4 require materially different topology from a compact dense
model. Omitting a GPU flag is deliberate: use Inference Lab or an explicit qualified serving plan
rather than presenting one accelerator as a sizing recommendation.
Model popularity changes faster than InferCrane releases. The source of truth is the immutable model
artifact and the evidence attached to its serving plan, not a marketing list of “supported models.”
plan is side-effect free. deploy returns a durable operation unless --wait is supplied. Closing
the terminal stops only the local watch; use the returned operation ID to reconnect:
Diagnose readiness before changing the plan
Do not submit a second deployment when readiness fails. Preserve the operation and provider identity, then collect:
Retry the same operation or desired state only with the original idempotency key. After a terminal
failure, choose between an immutable corrected candidate and plan-first deletion. Never erase
PostgreSQL to clear readiness state. Cleanup is complete only when InferCrane shows no deployment or
orphan and direct provider inventory proves zero run-owned resources. An uncertain create or delete
must keep the same persisted replica identity and operation; never submit a replacement by name.
Clean up the exact deployment
Deletion can terminate billable provider capacity. First render the semantic plan, verify the deployment name and every persisted provider identity, then submit one durable delete operation:delete --plan is side-effect free; --yes authorizes termination. InferCrane inventory is
necessary but not sufficient: use the selected provider guide to query its
authoritative inventory read-only and prove that every run-owned pod, instance, endpoint, service,
and attached billable resource reached zero. If provider deletion is eventually consistent, keep
watching the same durable operation and inventory; do not create replacement capacity during
cleanup. Unknown ownership is a stop condition: preserve durable state, retry only read-only
observation, and never create, adopt, or delete until one authoritative identity or absence is
proven.
Capture canonical provider inventory before deployment and after deletion with the command for the
selected adapter. In a shared account, require equality with the baseline rather than global zero;
never delete another deployment’s resource:
- RunPod
- AWS EC2
- GCP Compute
- Kubernetes or KServe
Before opening a shared endpoint
A ready GPU is not the complete production boundary. Use a non-bootstrap, tenant-scoped credential, persist endpoint admission, set the distributed tenant request ceiling, and qualify only the protocols your exact runtime/model pair claims:- Authentication and secret boundaries
- Admission and tenant quota procedure
- Protocol and cancellation qualification
- Autoscaling and active-stream qualification
- Request Inspector evidence
Send the first real request
Wait untilstatus reports the deployment ready before sending application traffic. The application
uses the control-plane URL as its stable OpenAI-compatible base URL and the deployment name as
model; provider worker addresses remain infrastructure details:
X-Request-Id response header from cURL to prove which revision and target served it:
RunPod is the first real-infrastructure acceptance target, not the InferCrane domain model. Its
elastic SkyPilot adapter remains experimental until the credentialed lifecycle and zero-leak
gate is complete. AWS BYOC, Kubernetes, existing targets, and future providers implement the same
provider contract with independent evidence. Check compatibility and qualification before
selecting a backend.
Bring an immutable runtime when a model flag is not enough
For a qualified custom runtime, describe the OCI image, digest, startup contract, health contract, and protocol capabilities in a DeploymentSpec. InferCrane persists that configuration as an immutable revision and manages its lifecycle; it does not build images or execute arbitrary build scripts.What “build” means here
This boundary keeps the product extensible: vLLM, SGLang, a custom OCI runtime, or another qualified
runtime can change without changing the application-facing endpoint.