Skip to main content
Open-source inference release system
Any compatible modelSafe releasesExplainable operationsNo Kubernetes required
Prove the release loop locally
InferCrane turns a compatible model or immutable OCI workload into one OpenAI-compatible logical endpoint. Durable operations provision, reconcile, scale, and safely replace workers behind it. Already running inference? Connect it in observe-only mode without migrating or transferring ownership.
requestOpenAI client
stable identityendpoint
atomic policyroute snapshot
readyruntime binding

Choose your starting point

Call a qualified Model API

Use the OpenAI-compatible contract when the catalog publishes a qualified managed offer.

Deploy a model

Plan, review, and deploy a Hugging Face model or explicit serving plan into infrastructure you control.

Connect existing inference

Discover vLLM, SGLang, LiteLLM, or another OpenAI-compatible endpoint before transferring ownership.

Optimize a workload

Search bounded candidates, compare exact workload evidence, and promote only a qualified winner.

Run an agent sandbox

Give an agent a persistent workspace with commands, files, previews, sleep, resume, and scoped model access.

Prove the loop locally

Run the complete GPU-free control and request path in under five minutes.

Measured evidence

On the recorded Qwen/Qwen3.8-27B-FP8, one-H200, 4K-input/512-output, concurrency-12 workload, the selected configuration moved aggregate output throughput from 810.4 to 1,393.6 tokens/s and median TTFT from 1,591 ms to 739 ms. That is exact-tuple evidence, not a universal model or provider claim. Read the benchmark index and the linked raw receipt before reusing the result.

The production operating loop

1

Deploy and converge

Persist the desired model, runtime, infrastructure, and scale policy before provider work begins.
2

Observe the real request path

Correlate endpoint, revision, replica, queueing, TTFT, generation, errors, and cost evidence.
3

Validate a candidate

Compare active and candidate revisions with benchmark, replay, runtime, and optional signed task-quality evidence.
4

Promote, reject, or explain

Apply deterministic Release Guard policy and preserve the decision for audit and recovery.
Provider and runtime capabilities are declared and conformance-tested, but each adapter has an independent qualification state. Review compatibility and qualification before relying on an exact provider, runtime, model, and accelerator combination.

See production patterns

Follow real-life deployment, adoption, rollout, cost, and composition journeys.

Understand the architecture

See the durable control plane, database-free request path, and replaceable adapter boundaries.

Inspect the CLI

Browse current commands, durable operation behavior, and machine-readable output.

Choose an operating model

Understand Community, the hosted Cloud direction, and the Enterprise boundary.

Release and upgrade safely

Qualify client, protocol, runtime, and control-plane changes with explicit restoration paths.