Inference telemetry
InferCrane records operational measurements for each request without storing prompts or generated content. Request bodies are used only to forward the OpenAI-compatible request, and response chunks are observed transiently to extract timing and usage metadata when the runtime supplies it. Persisted request dimensions are:- deployment and active revision
- provider, runtime, and compute mode
- GenAI operation name (
chat) - requested logical model and runtime-reported response model
- OpenTelemetry GenAI schema identity (
https://opentelemetry.io/schemas/gen-ai/1.42.0) - HTTP status, error type, and whether the response streamed
- response model, when returned by the runtime
gen_ai.response.time_to_first_chunk; InferCrane does not claim it is
the model server’s internal gen_ai.server.time_to_first_token. Missing token usage remains unknown;
InferCrane does not estimate it. Replica identity remains unset when the standalone router does not
provide a trustworthy selected-worker identity.
The corresponding OpenTelemetry GenAI concepts are gen_ai.operation.name,
gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model,
gen_ai.request.stream, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens,
gen_ai.server.request.duration, and gen_ai.response.time_to_first_chunk. InferCrane-specific
deployment, revision, provider/runtime, compute-mode, and operation dimensions are retained with
the durable request record so decisions can be reproduced later. runpod is a grounded custom
provider value because the convention permits provider-specific names outside its well-known list.
Aggregated deployment statistics expose request rate, error rate, latency p50/p95, TTFT p50/p95, and observed input/output tokens per second over the selected window. Token throughput is emitted only from runtime-reported usage.
The endpoint monitoring API returns the same content-free evidence as a bounded time series, plus
binding/deployment/revision attribution and persisted lifecycle overlays. The console uses this API;
it never queries PostgreSQL or an OpenTelemetry collector from the browser. Missing buckets and
unsupported runtime metrics stay null or explicitly unavailable rather than becoming zero.
Each operator-visible metric also carries a normalized evidence envelope: value, unit,
availability, evidence class, source, observation time, freshness boundary, sample count, and a
machine-readable absence reason. Availability distinguishes available, stale, not_observed,
and unsupported; evidence class distinguishes measured, provider-reported, modeled, and estimated
values. This lets CLIs, SDKs, and consoles preserve the same truth boundary instead of inferring one
from a nullable number.
Import existing production telemetry
Customers can keep Baseten, Fireworks, Datadog, or an OpenTelemetry pipeline as their monitoring source and post a strict content-free observation batch:source_request_id is deterministically scoped to the tenant,
deployment, and source, so retrying an import updates the same request record.
After import, POST /api/v1/deployments/{name}/replays captures an immutable workload-shape trace;
GET /api/v1/deployments/{name}/replays lists its history. Replay evidence contains arrival,
duration, token-count, session-hash, streaming, and concurrency shapes only. It does not duplicate
live traffic. Real prompt shadowing remains a separate explicit data-policy and approval boundary.
The Prometheus endpoint also exposes accounting queue depth and capacity, persisted and dropped
request-record counters, and persistence failures. These distinguish inference-path health from
telemetry backpressure without putting PostgreSQL on the request path. vLLM running/waiting and
cache signals are persisted with autoscaling decisions; streaming cancellations and upstream
disconnects are persisted as client_cancelled and upstream_disconnect error types. GPU metrics
remain unavailable unless a qualified provider/runtime collector supplies a trustworthy fresh
measurement.
NVIDIA GPU evidence
When an environment already runs NVIDIA DCGM Exporter, an operator can import a bounded snapshot without exposing the exporter to the browser:--file dcgm.prom instead of --url for an offline snapshot. --selector accepts exact,
comma-separated Prometheus label matches. Shared exporters should always be scoped to the intended
workload. DCGM_FI_DEV_GPU_UTIL is percent by default; pass --utilization-unit ratio only when an
upstream compatibility layer deliberately rewrites it to 0..1.
InferCrane normalizes only:
The snapshot is attached to the active revision at ingestion time. It cannot silently move to a
later rollout. Once its TTL expires, the monitoring API returns a stale envelope with no value.
This integration proves the control-plane semantics locally; actual DCGM field availability varies
by GPU and DCGM version and still requires real-hardware qualification. Collector snapshots follow
the configured high-volume request-evidence retention period (
INFERCRANE_REQUEST_RETENTION_HOURS).