Production operations
InferCrane is designed to run as multiple stateless gateway replicas backed by PostgreSQL. Each replica has a stableINFERCRANE_INSTANCE_ID and supervises its own loopback vLLM Router
processes. Router generations are scoped by instance, so one replica never publishes another
replica’s loopback endpoint.
Required configuration:
INFERCRANE_ENV=production: enables production security validation.INFERCRANE_DATABASE_URL: PostgreSQL URL with TLS enabled outside a trusted private network.INFERCRANE_API_KEY: at least 32 characters, supplied by the workload secret manager; no production default exists.INFERCRANE_URL: absolute HTTP(S) control-plane base URL used by lifecycle CLI commands.INFERCRANE_TLS_CERT_FILEandINFERCRANE_TLS_KEY_FILE: optional native server identity. Both are required together. AddINFERCRANE_TLS_CLIENT_CA_FILEto require and verify client certificates.INFERCRANE_PASSPORT_SIGNING_KEY_FILE: optional mounted Ed25519 private-key file for issuing Inference Passports. It must be readable only by its owner (0600); the private key is never persisted in PostgreSQL. Back it up and rotate it through the workload secret manager.
Hosted control plane on Fly.io
deploy/fly/control-plane.toml is the maintained single-machine hosted-control-plane profile. It
runs the Go API, gateway, reconciler, and durable-operation worker on a small CPU machine. It does
not run model inference and does not make Fly.io a GPU supplier. Managed model APIs, RunPod
Serverless, bounded Modal experiments, BYOC deployments, and connected endpoints remain behind
InferCrane’s provider contracts.
Use a Supabase session pooler connection string on port 5432 for this persistent IPv4 backend.
Keep sslmode=require in the URL. Direct database URLs may require IPv6; transaction pooling is
intended for transient or serverless clients and can change prepared-statement/session behavior.
The profile deliberately keeps one machine running. Auto-stop is unsafe for the background
reconciler and would interrupt durable work until another request arrived. Start with one machine
to minimize fixed cost; move to two machines only after completing the database-failover and
rolling-upgrade qualifications described below.
Configure secrets without printing or committing them:
RUNPOD_API_KEY and INFERCRANE_MANAGED_DEPLOYMENTS_ENABLED=true as Fly secrets. The RunPod
key is the platform’s supply credential, not a tenant BYOC credential. Leave the flag absent or
false until current price ingestion, Stripe funding, auto-stop, provider deletion, and settlement
are all operational.
The Fly Machine hostname is the default replica identity. Do not set INFERCRANE_INSTANCE_ID to
one shared value: two live replicas with the same identity would violate lease ownership and route
generation isolation. Set the variable only when the host supplies a different stable, unique
value for every replica.
The provider-neutral runtime requirements, scale-out gates, and Fly-to-ECS/Kubernetes exit path are
defined in Control-plane hosting. The maintained Fly
profile remains deliberately single-machine until those real failure and rolling-upgrade gates are
recorded as passing.
Use test-mode Stripe keys until the complete checkout, signed webhook, idempotent ledger-credit,
refund, and replay journey passes. The hosted verifier accepts either
INFERCRANE_HOSTED_AUTH_JWT_KEY from a secret manager or the file variant used by mounted-secret
platforms, never both. A static Clerk PEM avoids a network dependency during request verification;
rotate it when the Clerk signing key changes.
Public hosted signup also requires INFERCRANE_HOSTED_AUTH_AUTO_PROVISION=true. After Clerk verifies
the session and active organization, InferCrane atomically creates a dedicated tenant for a new
organization, grants its first member administrator access, and maps later verified members without
changing existing roles or restoring revoked access. Leave this setting disabled for invite-only
or manually provisioned installations.
For a test-mode Stripe account authenticated with the Stripe CLI, create or verify the fixed-price
prepaid catalog with scripts/bootstrap-stripe-prepaid.sh. The command is idempotent by product
name and denomination and prints the non-secret INFERCRANE_STRIPE_PRICE_IDS_JSON value. It never
creates recurring prices and never runs in live mode. Keep secret keys and webhook signing secrets
in the deployment secret manager.
The default Model API catalog is a deliberately small discovery shelf. It publishes model identity
and explicit unknown states, not managed availability or price claims. A hosted operator can load a
strict model-api-catalog/v1 file with INFERCRANE_MODEL_API_CATALOG_FILE to attach internal supply
contracts and InferCrane rate cards. Treat that file as an operational secret: it may contain
procurement routing, but the customer API collapses it to one infercrane-standard service offer,
standardizes price provenance, and omits internal offer IDs, regions, adapters, URLs, credentials,
and supplier identities. Invalid or incomplete priced entries fail control-plane startup.
Customer-wallet bindings additionally require an internal input/output cost basis, provenance, an
RFC3339 rate-card expiry, and a minimum gross-margin floor of at least 15%. The retail rate must
satisfy that floor for each charged token dimension. Authorization fails before supplier
transmission when the immutable rate card has expired. Public catalog responses also stop emitting
expired pricing and downgrade availability to unknown until a fresh contract is installed. See
Managed Model API architecture for supply lanes, prepaid loss
limits, and the evidence boundary for capacity-derived pricing.
For a single-host first installation, copy .env.production.example to a private path, replace
every example secret and URL, then render and start the maintained production stack:
The example pins the published v1.0.0-rc.1 image. Hosted documentation follows main, which may
describe work scheduled for the next prerelease. Do not combine a release image with a newer schema
or configuration contract without following upgrade and compatibility.
compose.yaml, this stack contains no fake workers or development router. Unlike
compose.runpod-acceptance.yaml, it contains no fault proxy or acceptance credential. PostgreSQL
is private to the Compose network; only the InferCrane API port is published. The bundled database
generates a persistent, self-signed server certificate and requires an encrypted
sslmode=require connection across that bridge. This encrypts the single-host transport but does
not provide CA-backed server identity; use managed PostgreSQL with verify-full, external secret
management, and multiple control-plane instances for a production service that must survive a host
failure.
The base production stack is provider-neutral: it does not require a RunPod credential and does
not start SkyPilot. Provider adapters remain dormant until a DeploymentSpec selects them. The
image contains pinned AWS CLI v2, Google Cloud CLI, and kubectl clients because those are the
explicit process boundaries used by the AWS, GCP, and Kubernetes adapters; credentials and
kubeconfig are always supplied by the operator at runtime.
For RunPod, add the explicit production overlay and the variables from .env.runpod.example:
INFERCRANE_SKYPILOT_API=disabled; SkyPilot starts only for an explicit provider
manifest in an operator-owned overlay.
InferCrane-paid GPU deployments remain disabled unless
INFERCRANE_MANAGED_DEPLOYMENTS_ENABLED=true is set explicitly in that private environment file.
When enabled, RunPod becomes internal InferCrane Cloud supply: the API requires a current exact
price, atomically holds sufficient prepaid wallet credit before queueing provider work, limits the
first self-serve product to one elastic GPU for 1–24 hours, and deletes resources before settling
actual runtime and releasing unused credit. BYOC requests cannot use this platform-owned provider
connection. See Managed deployments.
AWS and Kubernetes follow the same explicit composition pattern:
INFERCRANE_SKYPILOT_PROVIDERS_JSON; the entrypoint
then supervises its foreground API server and exits if that required subprocess stops.
INFERCRANE_INSTANCE_ID: stable and unique per control-plane replica.INFERCRANE_DATABASE_MAX_OPENandINFERCRANE_DATABASE_MAX_IDLE: size these with the total replica count below PostgreSQL’s connection budget or place PgBouncer in transaction mode.INFERCRANE_REQUEST_RETENTION_HOURS: bounds high-volume request-accounting storage; the default is 24 hours and cleanup runs in small batches.
/livez: process liveness; it does not depend on downstream services./readyz: PostgreSQL connectivity with a bounded timeout./metrics: Prometheus-format gateway request, failure, active request, byte, duration histogram, and operation claim, completion, failure, retry, and cancellation counters.
deploy/prometheus-rules.yaml from the repository, then tune thresholds
from real model latency and traffic. Follow the compatibility policy and perform
the backup/restore drill for every release containing migrations.
Use a disruption budget, topology spread constraints, anti-affinity across failure domains, and
at least two replicas. Termination grace must exceed INFERCRANE_SHUTDOWN_TIMEOUT_SECONDS so
streaming requests and buffered request accounting can drain. Do not configure an HTTP server
write timeout: it would terminate legitimate long-running streaming inference responses.
Each replica publishes its live binary and protocol interval. Use
infercrane system instances --output json before and during a rolling upgrade. Membership does not
elect a leader: durable operation claims are independently fenced in PostgreSQL. Configure CLI mTLS
with INFERCRANE_CLIENT_TLS_CA_FILE, INFERCRANE_CLIENT_TLS_CERT_FILE, and
INFERCRANE_CLIENT_TLS_KEY_FILE; Python SDK callers can pass ca_file, cert_file, and key_file.
Put private DNS, firewalls, and workload identity at the deployment boundary; TLS does not replace
network policy.
The production image includes InferCrane, the pinned upstream vLLM Router, the pinned SkyPilot
RunPod client, AWS CLI v2, and kubectl. Including provider clients does not select or configure a
provider. Development
workers and the simple development router exist only in the development image target and are
not performance or reliability substitutes for vLLM and vLLM Router.
Before rollout, qualify the exact PostgreSQL, vLLM, vLLM Router, model, GPU, and provider versions
with sustained load, streaming cancellation, worker loss, database failover, pod termination,
and soak tests. Capacity limits must be based on those measurements rather than defaults.
Release maintainers can validate packaging metadata without publishing with make release-check.
With syft installed, make release-artifacts RELEASE_TAG=v1.0.0 creates and
verifies four exact-version archives, checksums, archive SBOMs, and a generated Homebrew formula
under dist/. It pushes no tag, image, package, or release.
The control-plane API accepts the bootstrap bearer secret or hashed tenant-scoped credentials:
GET /api/v1/operations/{id}POST /api/v1/operations/{id}/cancelPOST /api/v1/deployments/applyPOST /api/v1/deploymentsDELETE /api/v1/deployments/{name}GET /api/v1/deploymentsGET|POST /api/v1/targetsGET /api/v1/orphansGET /api/v1/audit-eventsPUT /api/v1/tenant/quota- principal creation, rotation, and revocation endpoints