Skip to main content

Your first InferCrane request

In this guide you will start the complete control and request path, call a stable model endpoint, and inspect the durable state behind it. No GPU or cloud account is required.
The local stack uses GPU-free fake vLLM workers. It proves application and lifecycle behavior, not model quality, GPU performance, or real-provider compatibility.
Local proof and teammate access are two separate workflows. The quickstart binds a loopback fixture with the public development credential infercrane; sharing, tunneling, or port-forwarding that endpoint is unsupported. To serve a teammate without provisioning new GPU capacity, connect an already-operated upstream in a customer-owned TLS environment and issue a tenant-scoped credential. Follow Share a development endpoint safely; do not reuse the local fixture or its credential.

Fast proof: one command

From a repository checkout, one command runs the local product proof and cleans up its isolated Docker project:
The GPU-free proof connects a fixture endpoint because it cannot provision a real model. It sends and inspects a request, creates an intentionally unready candidate, records a deterministic Release Guard rejection, verifies the active revision did not change, rejects the candidate, and removes all temporary state. This proves the application and lifecycle loop locally; it does not replace the model-to-endpoint deployment path or claim GPU/provider qualification.
Choose either the one-command proof above or the manual tour below. make demo already starts an isolated stack, exercises the workflow, and tears that stack down; you do not need to run both paths.

Manual tour: inspect each step

Prerequisites

  • Git, Python 3, cURL, and jq
  • Docker with Compose v2 and a reachable engine
  • Port 18000 available
1

Clone InferCrane

Clone the public repository and enter the checkout:
2

Start the local stack

If Docker reports that port 18000 is already allocated, identify the owner before changing anything:
If it is an earlier InferCrane stack from this checkout, docker compose down stops it without deleting the named PostgreSQL volume, then retry up. If another application owns the port, leave it alone and choose a different host port:
Replace 18000 with the selected port in the remaining manual-tour URLs. Do not use docker compose down --volumes unless you intentionally want to erase local InferCrane state.If the port opens but /readyz does not become healthy, inspect the existing containers before retrying setup:
/livez proves the process is running; /readyz additionally requires PostgreSQL. Fix the first concrete container, configuration, or migration error, then rerun docker compose up --build -d. That retry reuses the same named volume and does not create provider resources. Never delete the PostgreSQL volume, edit migration-ledger rows, or start a paid deployment to repair a local readiness failure. If logs show a migration checksum, gap, or newer-schema error, stop and follow Upgrade and compatibility rather than retrying destructively.The stack includes PostgreSQL, InferCrane, two fake workers, and a development router. The development API key is infercrane; production mode rejects that default.
3

Send a request

The CLI and cURL examples require no language package. For the Python example, install the official OpenAI client in your active virtual environment with python -m pip install openai.
Add --stream to the CLI request to print response text as chunks arrive.
4

Inspect what happened

These commands use the authenticated control-plane API. Public CLI workflows never connect directly to PostgreSQL.
5

Open the terminal operations view

The terminal workspace reads the same durable control API and can be closed without cancelling operations. InferCrane also operates the browser console at console.infercrane.com; see Operations console for its source and deployment boundary.
6

Stop the stack

This preserves PostgreSQL data. Add --volumes only when you intend to delete the local state.

Choose your next path

Deploy a real model

Plan provider and runtime choices before creating billable capacity.

Connect an endpoint

Observe an existing vLLM, LiteLLM, or OpenAI-compatible endpoint first.

Share with a teammate

Use a separate TLS-protected environment and scoped credentials; never expose the local fixture.

Connect a model API

Register OpenRouter or another OpenAI-compatible API without enabling traffic or spend.

Understand the model

Learn endpoints, environments, serving plans, deployments, and revisions.

Before creating real infrastructure

Do not expose the local fixture as a team service

The Compose quickstart binds a development control plane and uses the known development credential infercrane. It is not a supported shared-access mode. Do not port-forward it, publish it through a tunnel, send the credential to a teammate, or treat loopback fake workers as a shared model. To collaborate without creating paid infrastructure, keep each developer’s fixture local and share the same version-controlled DeploymentSpec and qualification evidence. If the team needs one shared endpoint, use a customer-owned, TLS-protected environment with tenant-scoped credentials and network policy. Follow Share a development endpoint safely for the complete deployment, credential, qualification, request, and revocation path. An existing in-region vLLM/SGLang service can be connected observe-only without provisioning capacity; see Connect existing inference. Otherwise stop after plan until a provider budget and exact real-environment qualification are approved. The local fixture cannot be promoted into a shared environment: its workers do not contain the model, runtime, accelerator, or provider settings of a real deployment. Recreate the intended combination explicitly and stop before mutation if its evidence is incomplete:
The three results must agree on the adapter, runtime/version, compute mode, immutable model, accelerator, environment, and required real-system qualification. The local smoke result is only control-flow evidence and does not satisfy any missing row. Follow the selected provider guide and exact-combination policy before deploy. Install the public-beta CLI with Homebrew or download a matching archive from the v1.0.0-rc.1 prerelease:
plan is side-effect free. A real deploy can create billable resources, so check provider setup and compatibility and qualification first. Adapter registration does not prove that an exact model, runtime, GPU, and provider combination has passed real-infrastructure qualification.