Admission and async inference
InferCrane applies endpoint admission before an inference request reaches a runtime. Policies are loaded from PostgreSQL into an in-memory snapshot, so the request path never waits on a database.Set a distributed tenant request ceiling
Tenant quotas are currently a control-API operation rather than a dedicated CLI command. Use an admin-scoped credential and set all three tenant ceilings explicitly; omitted JSON fields decode as zero and therefore block that resource:quota.update. Request reservations are distributed per
UTC minute and refresh into gateway memory within approximately one second. A configured zero
request limit denies immediately after refresh; a lower nonzero limit is fully authoritative at the
next UTC-minute boundary because already issued leases cannot be revoked. There is no token-per-minute
tenant quota in this release: bound requested output tokens with endpoint admission and stop if a
separate distributed token budget is mandatory.
To qualify the request ceiling in staging, set a small nonzero limit immediately after a UTC-minute
boundary, wait for policy refresh, issue exactly that many successful requests, then verify the next
request returns OpenAI-compatible 429. Inspect response request IDs and the quota.update audit
event. Missing 429, negative/duplicate accounting, or unavailable audit evidence is a stop condition;
do not open the endpoint to more tenants.
Restore the externally recorded values rather than constructing a new guess:
Distinguish quota, admission, and backend failure
Qualify synchronous queue outcomes
Admission overflow and queue timeout both return HTTP429 with OpenAI-compatible JSON:
endpoint concurrency and queue capacity are exhausted from
endpoint admission queue timed out. Both overload outcomes include Retry-After: 1 and preserve a
caller-supplied request ID or create one before admission. Request-size rejection is
413 request_too_large; output-token or priority rejection is 422 admission_policy_violation.
An endpoint’s request_timeout_ms is one absolute deadline measured from gateway arrival through
admission, all buffered retry attempts, and response streaming. Expiry before response headers
returns 504 request_deadline_exceeded; expiry after a stream begins terminates the stream and
records the same error type without replaying any partial output.
Rejected and timed-out requests persist content-free Request Inspector evidence. Successful request
latency starts at gateway ingress, queue_ms records the bounded admission wait, and buffered retry
attempts are counted. For streaming responses, generation_ms is the measured interval from first
response bytes to completion; it remains unavailable when that boundary cannot be observed.
The endpoint monitoring response can also include a live admission snapshot with accepting,
queueing, or saturated capacity state. Its scope is explicitly gateway_instance: it is useful
for distinguishing a busy gateway from a dead runtime, but it is not a cluster-global concurrency
claim. /livez, /readyz, runtime health, and served-model identity remain separate signals.
Use a controllable staging runtime that can hold one request longer than the queue timeout. Do not
try to create this state by overloading production:
slow-request.json
request.json
429. This is distinct from durable async execution, where
deadline_seconds produces a persisted terminal job state retrievable with async get.
Record policy JSON, client correlation IDs, response status/body, elapsed queue time, completed
request evidence, and the exact staging runtime behavior. If the runtime cannot deterministically
hold the first request, treat this procedure as inconclusive and use the maintained local fault test
instead of making a production guarantee.
Protect an existing workload before taking traffic ownership
Admission does not require InferCrane to own provider lifecycle. Connect an existing OpenAI-compatible workload without changing its current request path, inspect it, and configure the stable endpoint before any client migration:observe-only; no traffic or provider lifecycle has transferred. Complete
the direct-versus-managed compatibility comparison
before adopt promote. That promotion grants routing ownership only. If qualification fails, keep
applications on the original upstream URL; InferCrane does not scale, update, or delete it. There is
no automatic ownership rollback command in this release, so do not switch application clients until
the managed path has passed. Delete the InferCrane endpoint only through its explicit dependency and
plan-first lifecycle if you later choose to remove the observation record.
Qualify external spend exhaustion
External budgets are binding-scoped, worst-case authorization reservations. They are separate from tenant request quota and do not replace provider billing.429 with
external_budget_exhausted evidence before upstream transmission. The persisted reservation must
not exceed one request or USD 0.01. Compare actual provider usage/billing separately; the reservation
is not proof that the first request cost USD 0.01.
After collecting evidence, remove only the dedicated staging objects:
Retry semantics
Retries are intentionally narrow. InferCrane retries only buffered requests to InferCrane-managed capacity, up to the persisted endpoint budget, for connection failures and HTTP502, 503 or
504. It never automatically retries streaming requests or paid external fallback. Attempts retain
one request ID and expose X-InferCrane-Attempt to the runtime.
Durable async requests
Durable async inference is local-qualified for encrypted persistence, idempotent submission,
fenced lease recovery, bounded retry, cancellation races, and webhook transport controls. Real
provider execution, upstream cancellation behavior, secret-manager key rotation, and external
webhook delivery remain separate environment-specific qualification. It is not a hosted queue or a
general workflow service.
request.json
queued or running. It prevents a leased worker
from committing a later result because the fenced status update no longer matches. It does not
promise that an already transmitted provider request stops instantly. If completion or cancellation
won the state transition first, the later request returns the current terminal outcome rather than
rewriting history. Inspect status, completed_at, error_code, and expires_at with
infercrane async get JOB_ID --output json.
Cancellation and billing
InferCrane does not claim that cancelling a job cancels provider billing:
For an external binding, the hard worst-case request/cost reservation is consumed before
transmission and is not a settled invoice or refundable billing record. For self-hosted capacity,
GPU cost continues while the replica exists regardless of one job’s cancellation. Verify actual
spend with the provider’s authoritative billing/export for the same time window and resource/model
identity. InferCrane currently has no provider-neutral per-job billing importer.
Collect the job JSON, request ID when present, Request Inspector record, endpoint binding, provider
operation/resource identity, cancellation timestamp, and provider usage/billing record. If
transmission or billing cannot be proven, report it as unavailable; do not claim saved cost.
Content and keys
Async mode is disabled until the control plane receives an encryption key of at least 32 bytes:store_encrypted_content: true. Payloads and results use AES-256-GCM with the
tenant and job identity as associated data. Plaintext is never written to PostgreSQL. Losing or
rotating a key without retaining the previous key makes existing results unreadable; production key
rotation therefore requires an operator-managed overlap procedure.
Signed completion webhook
InferCrane-Timestamp and
InferCrane-Signature: v1=<hmac-sha256>. Delivery has at most three attempts. The outbound transport
rejects redirects and private, loopback, link-local and multicast addresses after DNS resolution.
Current limits
- Queue order is priority then creation time; it is not a general scheduler.
- Results are limited to 32 MiB.
- Async streaming is not exposed; use synchronous streaming when incremental tokens matter.
- Cross-key decryption and automatic key re-encryption are not currently implemented.
- Tenant requests-per-minute is supported through the control API; tenant tokens-per-minute is not.