Your first InferCrane request
In this guide you will start the complete control and request path, call a stable model endpoint, and inspect the durable state behind it. No GPU or cloud account is required.The local stack uses GPU-free fake vLLM workers. It proves application and lifecycle behavior, not
model quality, GPU performance, or real-provider compatibility.
Fast proof: one command
From a repository checkout, one command runs the local product proof and cleans up its isolated Docker project:Choose either the one-command proof above or the manual tour below.
make demo already starts an
isolated stack, exercises the workflow, and tears that stack down; you do not need to run both paths.Manual tour: inspect each step
Prerequisites
- Git, Python 3, cURL, and jq
- Docker with Compose v2 and a reachable engine
- Port
18000available
1
Clone InferCrane
Clone the public repository and enter the checkout:
2
Start the local stack
18000 is already allocated, identify the owner before changing
anything:docker compose down stops it without
deleting the named PostgreSQL volume, then retry up. If another application owns the port,
leave it alone and choose a different host port:18000 with the selected port in the remaining manual-tour URLs. Do not use
docker compose down --volumes unless you intentionally want to erase local InferCrane state.If the port opens but /readyz does not become healthy, inspect the existing containers before
retrying setup:/livez proves the process is running; /readyz additionally requires PostgreSQL. Fix the first
concrete container, configuration, or migration error, then rerun docker compose up --build -d.
That retry reuses the same named volume and does not create provider resources. Never delete the
PostgreSQL volume, edit migration-ledger rows, or start a paid deployment to repair a local
readiness failure. If logs show a migration checksum, gap, or newer-schema error, stop and follow
Upgrade and compatibility rather than retrying destructively.The stack includes PostgreSQL, InferCrane, two fake workers, and a development router. The
development API key is infercrane; production mode rejects that default.3
Send a request
The CLI and cURL examples require no language package. For the Python example, install the
official OpenAI client in your active virtual environment with Add
python -m pip install openai.--stream to the CLI request to print response text as chunks arrive.4
Inspect what happened
5
Open the terminal operations view
6
Stop the stack
--volumes only when you intend to delete the local state.Choose your next path
Deploy a real model
Plan provider and runtime choices before creating billable capacity.
Connect an endpoint
Observe an existing vLLM, LiteLLM, or OpenAI-compatible endpoint first.
Share with a teammate
Use a separate TLS-protected environment and scoped credentials; never expose the local fixture.
Connect a model API
Register OpenRouter or another OpenAI-compatible API without enabling traffic or spend.
Understand the model
Learn endpoints, environments, serving plans, deployments, and revisions.
Before creating real infrastructure
Do not expose the local fixture as a team service
The Compose quickstart binds a development control plane and uses the known development credentialinfercrane. It is not a supported shared-access mode. Do not port-forward it, publish it through a
tunnel, send the credential to a teammate, or treat loopback fake workers as a shared model.
To collaborate without creating paid infrastructure, keep each developer’s fixture local and share
the same version-controlled DeploymentSpec and qualification evidence. If the team needs one shared
endpoint, use a customer-owned, TLS-protected environment with tenant-scoped credentials and network
policy. Follow Share a development endpoint safely for the complete
deployment, credential, qualification, request, and revocation path. An existing in-region
vLLM/SGLang service can be connected observe-only without provisioning capacity; see
Connect existing inference. Otherwise stop after plan until a
provider budget and exact real-environment qualification are approved.
The local fixture cannot be promoted into a shared environment: its workers do not contain the
model, runtime, accelerator, or provider settings of a real deployment. Recreate the intended
combination explicitly and stop before mutation if its evidence is incomplete:
deploy.
Install the public-beta CLI with Homebrew or download a matching archive from the
v1.0.0-rc.1 prerelease:
plan is side-effect free. A real deploy can create billable resources, so check
provider setup and compatibility and qualification first. Adapter
registration does not prove that an exact model, runtime, GPU, and provider combination has passed
real-infrastructure qualification.