The hosted console is a company-operated service over the public InferCrane control contracts.
The Apache-2.0 core remains CLI-, API-, Terraform-, and terminal-first and does not require an
InferCrane account.
What the console covers
- fleet attention, stable endpoints, and concrete deployments;
- durable operations with reconnectable timelines and cooperative cancellation;
- Request Inspector and deterministic Doctor findings;
- endpoint monitoring for request rate, errors, fallback, TTFT, queueing, latency, and reported token throughput;
- typed freshness, source, sample count, and availability for every monitoring measurement;
- Release Guard, benchmark, replay, revision, and replica evidence where available;
- lifecycle overlays, signed alert policies, sourced FinOps reports, and signed output-quality evidence;
- provider/runtime capability inventory and secret references;
- private/gated model-access references and authenticated existing-endpoint discovery;
- content-free telemetry imports and workload replay history;
- scoped API-key lifecycle, organization roles, hosted SSO entry, and approval boundaries;
- a curated Model API shelf with InferCrane rates, capabilities, limits, and availability.
First-run paths
Create endpoint starts with the application identity the user wants, then asks how it should be served. The primary paths are ordered by intent:- Deploy an open model starts from a reviewed model configuration, exposes only the required serving choices, and keeps runtime, accelerator, region, and scaling under Advanced.
- Connect existing inference discovers vLLM, SGLang, LiteLLM, or another compatible endpoint and starts observe-only. The endpoint page then presents a separate, explicit ownership action before InferCrane can route traffic. Promotion changes routing ownership only; it does not grant lifecycle control over the external workload. Saved Baseten, Fireworks, OpenRouter, and generic OpenAI-compatible connections can authenticate the bounded discovery call without putting their bearer credential in the browser or endpoint URL.
- Use a Model API selects a hosted InferCrane product, creates one API key, and uses the same OpenAI-compatible endpoint across the catalog. The customer surface shows only the InferCrane rate, capability, limit, and availability contract. Procurement routing, internal offer IDs, credentials, and upstream rate provenance remain inside the control plane. Prepaid reservation happens before a managed request can leave InferCrane.
Self-hosted interface boundary
The public repository ships the control API, OpenAI-compatible gateway, CLI, Python and TypeScript SDKs, Terraform provider, and terminal workspace. Those surfaces are sufficient to self-operate InferCrane and do not depend on the hosted console. The browser console source is maintained as a separate company-operated service and is not part of the Apache-2.0 repository.Hosted identity and workspace provisioning
Hosted mode uses Clerk for sign-in, sessions, and organization selection. The Go control plane then independently requires a mapped InferCrane user, organization membership, role/scopes, and activeweb_console_access entitlement.
Configure the Go verifier with the Clerk instance’s HTTPS issuer and public JWT verification key.
Use the inline value when the deployment platform stores environment secrets, or the file value
when it mounts secrets; never configure both:
Durable work
Closing the browser does not cancel deployment work. Reopen the console or resume from any terminal:Terminal equivalent
For headless, low-bandwidth, or incident use, the terminal surfaces remain first-class:npm run test:e2e:empty-control-plane journey starts with a new PostgreSQL
volume and verifies this first-run flow through the real control API. It is the regression boundary
for empty states and browser-to-control-plane wiring; the hosted identity provider and real GPU/cloud
boundaries remain separate release qualifications.