Healthy infrastructure is not enough
A candidate may be ready, fast, and error-free while producing worse answers. InferCrane accepts aggregate evidence from a customer-owned evaluation suite, verifies its signature and revision identity, and lets Release Guard enforce explicit score thresholds. InferCrane does not store prompts or generated outputs in this evidence format and does not choose an LLM judge for you.Create an evaluator key
Ingest a result from any evaluator
Have Ragas, DeepEval, a custom CI evaluator, or another customer-controlled system emit the strict content-free evaluator result contract atschemas/evaluator-result-v2.schema.json:
evaluator-result.json
scores contains exactly
sample_count finite values from 0 through 1 in the evaluator’s canonical pair order. Active and
candidate files must use the same pairing digest. The digest commits the private sample identities
without storing those identities, prompts, or responses in InferCrane.
Bind that aggregate result to the exact immutable candidate revision, sign it, and explicitly attach
it through the control-plane API:
evaluation attach retry.
Sign aggregate evidence directly
Fail closed in Release Guard
threshold remains the default comparison mode. In bootstrap mode InferCrane performs 10,000
deterministic paired resamples. A confidence interval wholly below the allowed regression passes;
one wholly beyond it rejects; an overlapping interval or too few pairs waits. The persisted decision
records the seed, alpha, active/candidate counts, observed regression, and interval so a retry is
reproducible.
The three-score example demonstrates the contract, not a recommended production sample size.
Choose and review a representative minimum for the evaluator and risk class; insufficient pairs
must remain waiting.