Skip to main content

One product model, three operating choices

InferCrane’s domain model does not change depending on who operates the control plane or supplies compute. Applications keep the same logical endpoint, deployments remain durable, and release evidence keeps the same meaning.

InferCrane Community

The public repository contains the complete self-operated inference control and request path and is licensed under Apache-2.0. It includes the control plane, gateway, provider and runtime contracts, CLI, SDKs, Terraform provider, terminal workspace, monitoring, durable operations, Release Guard, Replay, benchmarking, and optimization evidence. The company-operated browser console is a separate hosted product; it is not required to run the open-source system. Community users can:
  • deploy a compatible model through vLLM, SGLang, or an immutable custom OCI workload;
  • operate capacity in AWS, GCP, Kubernetes, RunPod, or another implemented adapter;
  • connect existing OpenAI-compatible inference without migrating it first;
  • keep one stable application endpoint while serving plans change;
  • inspect requests, diagnose production behavior, and gate releases with evidence.
The catalog is not an allowlist. A model accepted by the selected runtime can enter the same planning and lifecycle flow. Hardware fit, model terms, runtime capabilities, and performance remain properties of the exact serving plan.

InferCrane hosted console

The hosted console is the company-operated workspace for Model APIs, BYOC planning, optimization campaigns, usage, access, and sandboxes. It consumes the same public control contracts while adding hosted identity, tenancy, and service billing. A console account does not imply that InferCrane owns the user’s GPU capacity: BYOC operations use the customer’s provider account, and managed Model API offers appear only after their exact supplier, price, protocol, and capacity boundaries are qualified. The portability rule is:
Hosted supplier credentials, payment systems, abuse controls, rate cards, and warm-pool operations belong outside the open-source process. Cloud code may consume public InferCrane APIs and extension contracts; the open-source core must never depend on a proprietary package.

InferCrane Enterprise direction

Enterprise packaging is for contractual operation and organization-wide governance, not a second inference engine. Expected commercial capabilities include SSO and SCIM, fine-grained access control, audit export, private connectivity, residency policy, customer-managed encryption, multi-region recovery, approval workflows, support, and SLAs. Basic deployment, monitoring, autoscaling, release safety, and provider/runtime extensibility stay in Community. Commercial value comes from operating InferCrane, providing governed capacity, reducing organizational risk, and accepting contractual responsibility.

Repository and license boundary

Apache-2.0 is deliberately used for the core so individuals and organizations can inspect, modify, and self-operate the product under a permissive license with explicit contributor patent terms. A proprietary service does not change the license of the public core, and mixing proprietary source into this repository is not required for Cloud or Enterprise operation.

Existing orchestration systems

An existing OpenAI-compatible serving system can be connected through InferCrane’s generic adoption path without transferring lifecycle ownership. The existing system remains responsible for its own GPU scheduling, engine configuration, workers, and replica mutation. InferCrane can own stable endpoint identity, higher-level revisions, release policy, and evidence only after the operator chooses the corresponding ownership mode. This boundary prevents two controllers from scaling the same replicas. InferCrane will not add a native lifecycle adapter for another orchestration system unless repeated user demand justifies the extra controller and qualification surface.

Current availability

Product direction is not current availability. Review compatibility and qualification before relying on a provider, runtime, hosted surface, or exact model and accelerator combination.
Today, InferCrane is distributed as an open-source public beta with a company-operated console. The core control plane, gateway, CLI, SDKs, terminal workspace, and provider/runtime contracts are self-hostable. The browser console, hosted tenancy, supplier credentials, and service billing are operated surfaces. Managed Model APIs fail closed when no qualified offer is installed; enterprise SLAs and a blanket InferCrane-owned GPU cloud are not implied by console access.