ADR 0040: Durable evidence-gated optimization campaigns
Status: accepted.Context
InferCrane could propose conservative or modeled serving candidates, but a proposal directory did not provide a durable authority boundary for paid experiments. Quantized checkpoints, speculative draft models, and TensorRT engines also need stronger identity than a runtime flag. Their builders can succeed after a client disconnect, fail with partial output, or produce an artifact that is valid for a different model, runtime, GPU architecture, or quality evaluation. The upstream systems evolve independently. LLM Compressor and ModelOpt transform checkpoints, vLLM Speculators creates draft artifacts, TensorRT-LLM builds and executes hardware-specific engines, LMCache owns KV storage, and llm-d or AIBrix may own Kubernetes routing and scheduling. Importing their objects into the core domain would create version coupling and overlapping mutation owners.Decision
- Persist an immutable
OptimizationCampaignand ordered candidate runs before any provider mutation. Creation is idempotent and free of infrastructure side effects. - Require an actor, hard maximum spend, and expiry before an execution worker may claim paid work. Retrying approval never extends the first persisted authority.
- Keep predicted evidence separate from actual evidence. Candidate transitions are fenced by the state observed by the worker; stale workers cannot overwrite newer evidence.
- Derive campaign state from candidate truth. API handlers, CLI clients, and future workers do not implement independent campaign state machines.
- Represent transformed checkpoints, speculators, and TensorRT engines as immutable optimized artifacts with the base artifact, exact tool version, digest-pinned builder image, algorithm, configuration, hardware constraints, SPDX expression, and output digest.
- Execute builders outside the control-plane process. InferCrane records an attestation but never treats a builder process exit as semantic or performance proof.
- Require the exact candidate run, immutable revision, passing signed quality evidence, and optimized artifact identity to agree before the artifact becomes qualified.
- Compile runtime arguments only from exact, locally or really qualified capability descriptors. Unknown versions, model families, artifact precision, accelerators, and conflicting mechanisms fail closed.
- Register external builders, caches, and orchestrators through the replaceable composition contract. One component owns scaling and routing mutations for a serving graph.
- PostgreSQL remains outside the inference request path. Gateways consume immutable endpoint route generations; campaign and artifact state is control-plane evidence only.
Consequences
Operators can safely create and inspect an optimization campaign without allocating a GPU, and a later worker can resume from durable state after a process or client restart. Advanced users can provide a complete builder plan, while the common CLI offers small presets that remain exact, digest-pinned, externally executed, and quality-gated. This foundation does not make an optimization executable or fast by itself. Real quantization, speculative decoding, LMCache, NIXL, TensorRT-LLM, llm-d, and AIBrix claims remain deferred until the exact software/model/artifact/hardware/workload tuple passes its qualification gate.Alternatives rejected
- Run builders inside the API process: mixes untrusted, GPU-heavy Python/CUDA execution with the durable control plane and makes recovery unsafe.
- Store only generated flags: loses artifact provenance and cannot prove verifier, precision, or hardware compatibility.
- Allow estimates to qualify candidates: confuses modeled evidence with measured performance.
- Let InferCrane and an upstream orchestrator both scale or route: creates controller oscillation, double deletion, and unsafe drains.
- Automatically promote the highest-ranked candidate: bypasses Release Guard and human authority.
Verification
go test ./internal/optimizationcampaign ./internal/optimizationcapability ./internal/optimizedartifact- PostgreSQL integration tests cover duplicate creation, bounded approval, cancellation, stale transitions, partial candidate outcomes, exact artifact/candidate/revision/quality binding, and deterministic cleanup.
- The integration catalog test verifies that every supported composition capability references an existing regression test and that unqualified execution remains unsupported.
- The OpenAPI route-coverage test includes campaign and optimized-artifact endpoints.
- The Docker verifier rebuilds source before running race, migration, PostgreSQL, and CLI tests.