Skip to main content

Healthy infrastructure is not enough

A candidate may be ready, fast, and error-free while producing worse answers. InferCrane accepts aggregate evidence from a customer-owned evaluation suite, verifies its signature and revision identity, and lets Release Guard enforce explicit score thresholds. InferCrane does not store prompts or generated outputs in this evidence format and does not choose an LLM judge for you.

Create an evaluator key

Keep the private Ed25519 key outside the control plane. The operator authorized to attach evidence is the trust boundary.

Ingest a result from any evaluator

An artifact digest proves which private result bundle was evaluated, not that the evaluator was well-designed. Before generating or attaching evidence, approve the representative dataset, privacy boundary, evaluator and version, calibration procedure, sampling method, and required human review for high-risk changes. The commands below persist signed release evidence and --attach mutates the candidate record; they do not upload prompt/output bodies or authorize promotion.
Have Ragas, DeepEval, a custom CI evaluator, or another customer-controlled system emit the strict content-free evaluator result contract at schemas/evaluator-result-v2.schema.json:
evaluator-result.json
The abbreviated score array above is illustrative; in a valid file scores contains exactly sample_count finite values from 0 through 1 in the evaluator’s canonical pair order. Active and candidate files must use the same pairing digest. The digest commits the private sample identities without storing those identities, prompts, or responses in InferCrane. Bind that aggregate result to the exact immutable candidate revision, sign it, and explicitly attach it through the control-plane API:
Unknown fields are rejected, so prompt and generated-output bodies cannot accidentally enter the evidence envelope. If API attachment fails, the signed local file remains available for a safe evaluation attach retry.

Sign aggregate evidence directly

The signed payload records deployment, immutable revision, suite and evaluator versions, normalized score, pass/fail result, sample count, result-artifact digest, and evaluation time. Attaching the same evidence twice is idempotent.

Fail closed in Release Guard

Active and candidate evidence must use the same suite version and evaluator version to be comparable. Missing or incomparable required evidence leaves the decision waiting. A failed suite, score below the minimum, or regression beyond policy rejects the candidate. The evaluator cannot activate a revision; promotion remains a separate, authorized control-plane decision. threshold remains the default comparison mode. In bootstrap mode InferCrane performs 10,000 deterministic paired resamples. A confidence interval wholly below the allowed regression passes; one wholly beyond it rejects; an overlapping interval or too few pairs waits. The persisted decision records the seed, alpha, active/candidate counts, observed regression, and interval so a retry is reproducible. The three-score example demonstrates the contract, not a recommended production sample size. Choose and review a representative minimum for the evaluator and risk class; insufficient pairs must remain waiting.