Skip to main content

Admission and async inference

InferCrane applies endpoint admission before an inference request reaches a runtime. Policies are loaded from PostgreSQL into an in-memory snapshot, so the request path never waits on a database.
The policy bounds concurrency, queue depth, queue wait, encoded request size, requested output tokens and accepted priority classes. Rejections use OpenAI-compatible errors and happen before upstream transmission. Tenant request/token limits and external-capacity budgets remain separate hard controls.
Endpoint concurrency and queue depth are shared by every tenant authorized to call that endpoint; this release has no per-tenant concurrency or per-tenant queue partition. A tenant request-per-minute quota limits admitted volume over a UTC window but cannot reserve an instantaneous capacity share. If strict noisy-neighbor isolation is mandatory, use separate tenant boundaries, stable endpoints, and dedicated backend capacity pools - or separate control-plane environments - then qualify each pool. Separate endpoint names backed by the same saturated deployment are not capacity isolation. Stop before onboarding another tenant if that boundary cannot be provided.
For external capacity, keep a distinct binding with its own hard request and worst-case cost reservations for each isolated endpoint. For self-hosted capacity, GPU billing continues while the dedicated replicas exist, including after one request is cancelled. InferCrane has no provider-neutral per-tenant GPU invoice or automatic refund claim. Inspect the effective endpoint policy after mutation:

Set a distributed tenant request ceiling

Tenant quotas are currently a control-API operation rather than a dedicated CLI command. Use an admin-scoped credential and set all three tenant ceilings explicitly; omitted JSON fields decode as zero and therefore block that resource:
Quota updates affect the entire active tenant and the API has no quota-read endpoint. Before this mutation, record the reviewed existing JSON outside InferCrane as previous-quota.json. Test only in a dedicated staging tenant; never guess deployment or replica ceilings, and restore exactly the recorded values afterward.
The update is tenant-scoped and audited as quota.update. Request reservations are distributed per UTC minute and refresh into gateway memory within approximately one second. A configured zero request limit denies immediately after refresh; a lower nonzero limit is fully authoritative at the next UTC-minute boundary because already issued leases cannot be revoked. There is no token-per-minute tenant quota in this release: bound requested output tokens with endpoint admission and stop if a separate distributed token budget is mandatory. To qualify the request ceiling in staging, set a small nonzero limit immediately after a UTC-minute boundary, wait for policy refresh, issue exactly that many successful requests, then verify the next request returns OpenAI-compatible 429. Inspect response request IDs and the quota.update audit event. Missing 429, negative/duplicate accounting, or unavailable audit evidence is a stop condition; do not open the endpoint to more tenants. Restore the externally recorded values rather than constructing a new guess:

Distinguish quota, admission, and backend failure

Backend failure is supported by request status/error type, attempt count, route identity, runtime health, and durable provider events, not by weakening quota. If the request ID or provider outcome is unknown, keep the safety control in place and use the provider outage runbook.

Qualify synchronous queue outcomes

Admission overflow and queue timeout both return HTTP 429 with OpenAI-compatible JSON:
The message distinguishes endpoint concurrency and queue capacity are exhausted from endpoint admission queue timed out. Both overload outcomes include Retry-After: 1 and preserve a caller-supplied request ID or create one before admission. Request-size rejection is 413 request_too_large; output-token or priority rejection is 422 admission_policy_violation. An endpoint’s request_timeout_ms is one absolute deadline measured from gateway arrival through admission, all buffered retry attempts, and response streaming. Expiry before response headers returns 504 request_deadline_exceeded; expiry after a stream begins terminates the stream and records the same error type without replaying any partial output. Rejected and timed-out requests persist content-free Request Inspector evidence. Successful request latency starts at gateway ingress, queue_ms records the bounded admission wait, and buffered retry attempts are counted. For streaming responses, generation_ms is the measured interval from first response bytes to completion; it remains unavailable when that boundary cannot be observed. The endpoint monitoring response can also include a live admission snapshot with accepting, queueing, or saturated capacity state. Its scope is explicitly gateway_instance: it is useful for distinguishing a busy gateway from a dead runtime, but it is not a cluster-global concurrency claim. /livez, /readyz, runtime health, and served-model identity remain separate signals. Use a controllable staging runtime that can hold one request longer than the queue timeout. Do not try to create this state by overloading production:
slow-request.json
request.json
After the running request finishes, a new request must complete normally. A synchronous queued request has no pollable intermediate object: it either acquires capacity and completes, its client cancels, or it returns the queue-timeout 429. This is distinct from durable async execution, where deadline_seconds produces a persisted terminal job state retrievable with async get. Record policy JSON, client correlation IDs, response status/body, elapsed queue time, completed request evidence, and the exact staging runtime behavior. If the runtime cannot deterministically hold the first request, treat this procedure as inconclusive and use the maintained local fault test instead of making a production guarantee.

Protect an existing workload before taking traffic ownership

Admission does not require InferCrane to own provider lifecycle. Connect an existing OpenAI-compatible workload without changing its current request path, inspect it, and configure the stable endpoint before any client migration:
The connection remains observe-only; no traffic or provider lifecycle has transferred. Complete the direct-versus-managed compatibility comparison before adopt promote. That promotion grants routing ownership only. If qualification fails, keep applications on the original upstream URL; InferCrane does not scale, update, or delete it. There is no automatic ownership rollback command in this release, so do not switch application clients until the managed path has passed. Delete the InferCrane endpoint only through its explicit dependency and plan-first lifecycle if you later choose to remove the observation record.

Qualify external spend exhaustion

External budgets are binding-scoped, worst-case authorization reservations. They are separate from tenant request quota and do not replace provider billing.
This staging-only test transmits non-sensitive input to an external provider and may incur cost. Get data-processing approval, verify the provider/model credential and current worst-case price, and use a dedicated endpoint and request limit of one. Stop if price, currency, residency, or provider billing evidence is unavailable.
Send one approved request and inspect its request ID. A second request must return 429 with external_budget_exhausted evidence before upstream transmission. The persisted reservation must not exceed one request or USD 0.01. Compare actual provider usage/billing separately; the reservation is not proof that the first request cost USD 0.01. After collecting evidence, remove only the dedicated staging objects:
Do not reuse this tiny budget as a production price or mutate an existing binding to reset consumed reservations. Missing 429, more than one upstream request, absent provider billing evidence, or unknown data handling fails the qualification.

Retry semantics

Retries are intentionally narrow. InferCrane retries only buffered requests to InferCrane-managed capacity, up to the persisted endpoint budget, for connection failures and HTTP 502, 503 or 504. It never automatically retries streaming requests or paid external fallback. Attempts retain one request ID and expose X-InferCrane-Attempt to the runtime.

Durable async requests

Durable async inference is local-qualified for encrypted persistence, idempotent submission, fenced lease recovery, bounded retry, cancellation races, and webhook transport controls. Real provider execution, upstream cancellation behavior, secret-manager key rotation, and external webhook delivery remain separate environment-specific qualification. It is not a hosted queue or a general workflow service.
Before async submit, inspect the endpoint binding and provider/resource identity, obtain approval for the data boundary and worst-case spend, and use non-sensitive staging input. Submission may transmit work immediately. Cancelling after transmission may not stop upstream compute or billing; an external budget reservation is a worst-case authorization, not a refundable invoice, and a self-hosted GPU remains billable while its replica exists. If provider identity, maximum exposure, or cancellation behavior is unknown, do not submit. See Cancellation and billing.
Create a protocol-native request file:
request.json
Submit it and close the terminal safely:
The command returns a job ID immediately. Reading job state is non-mutating:
Before cancellation, inspect the JSON state. If the job is already running, assume the request may have reached the provider and that billing may continue. Cancellation makes InferCrane’s job state terminal; it does not prove upstream work stopped or refund a prior external reservation.
Async execution is one bounded inference request, not a workflow engine. Claims use PostgreSQL leases and fencing tokens. Another worker adopts an expired lease; a stale worker cannot commit the result. The default execution deadline is 900 seconds and default encrypted retention is 86,400 seconds from submission. Retention must be longer than the deadline and cannot exceed 604,800 seconds (seven days). Jobs have at most three execution attempts. Cancellation is terminal only while a job is queued or running. It prevents a leased worker from committing a later result because the fenced status update no longer matches. It does not promise that an already transmitted provider request stops instantly. If completion or cancellation won the state transition first, the later request returns the current terminal outcome rather than rewriting history. Inspect status, completed_at, error_code, and expires_at with infercrane async get JOB_ID --output json.

Cancellation and billing

InferCrane does not claim that cancelling a job cancels provider billing: For an external binding, the hard worst-case request/cost reservation is consumed before transmission and is not a settled invoice or refundable billing record. For self-hosted capacity, GPU cost continues while the replica exists regardless of one job’s cancellation. Verify actual spend with the provider’s authoritative billing/export for the same time window and resource/model identity. InferCrane currently has no provider-neutral per-job billing importer. Collect the job JSON, request ID when present, Request Inspector record, endpoint binding, provider operation/resource identity, cancellation timestamp, and provider usage/billing record. If transmission or billing cannot be proven, report it as unavailable; do not claim saved cost.

Content and keys

Async mode is disabled until the control plane receives an encryption key of at least 32 bytes:
The API requires store_encrypted_content: true. Payloads and results use AES-256-GCM with the tenant and job identity as associated data. Plaintext is never written to PostgreSQL. Losing or rotating a key without retaining the previous key makes existing results unreadable; production key rotation therefore requires an operator-managed overlap procedure.

Signed completion webhook

Webhooks are HTTPS-only and signed with InferCrane-Timestamp and InferCrane-Signature: v1=<hmac-sha256>. Delivery has at most three attempts. The outbound transport rejects redirects and private, loopback, link-local and multicast addresses after DNS resolution.
Completion webhooks contain the inference result. Configure them only when the destination is authorized to receive that content.

Current limits

  • Queue order is priority then creation time; it is not a general scheduler.
  • Results are limited to 32 MiB.
  • Async streaming is not exposed; use synchronous streaming when incremental tokens matter.
  • Cross-key decryption and automatic key re-encryption are not currently implemented.
  • Tenant requests-per-minute is supported through the control API; tenant tokens-per-minute is not.