One directory, one serving intent
An inference project keeps the deployment specification beside the code and configuration that produce it. InferCrane discoversinfercrane.yaml from the current directory or a parent, validates
it locally, and uses the same file for planning and deployment.
workload init requires one serving source: either --model MODEL or --recipe NAME. If you have
not selected one yet, run infercrane recipes curated to inspect the reviewed starting points.
--profile must name a serving profile belonging to the selected recipe. When omitted, the first
balanced/general profile is used. A profile is a reviewed candidate configuration: it writes its
runtime and version, immutable workload image/argv when present, accelerator type and count, provider
hint, replica bounds, and routing strategy into the project. A profile can restrict accelerator
types when its runtime configuration is hardware-specific. It does not claim that those choices are
fastest or currently available on the selected provider.
workload init writes a valid DeploymentSpec with a JSON Schema directive understood by YAML
language servers in VS Code, JetBrains editors, and other schema-aware tools. Fields, enums, and
unknown keys are checked while you type. A curated recipe also pins the model commit. Recipes are
reviewed configuration starting points, not performance claims; use benchmark or replay evidence
for serving decisions.
The command prints one provider-specific path from validation to first request. For AWS that path is:
deploy --wait prints and persists the durable operation identity. Closing the terminal does not
cancel the operation; reattach with infercrane operation watch OPERATION_ID. The application keeps
using model="fraud-explainer" while future revisions, runtimes, or infrastructure change behind the
stable endpoint.
Local validation before GPU allocation
plan is also non-mutating. This keeps an invalid container command or unsupported
runtime combination from becoming a paid provisioning attempt.
The versioned schema is published from
schemas/deployment-v1.schema.json in this repository.
The Go loader remains authoritative; the schema improves editing and never weakens server-side
validation.
Build a custom OCI workload
Custom OCI projects can build through Docker Buildx:Run a custom workload locally
127.0.0.1. The command is available for custom-oci
projects; vLLM and SGLang model projects use their qualified runtime adapters.
InferCrane owns the reproducible serving intent, not a new container build system. Docker/Buildx,
your registry, vLLM, SGLang, and custom OCI runtimes retain their established responsibilities.