PostgreSQL backup and restore drill
Set a TLS-protectedINFERCRANE_DATABASE_URL, then create and validate a custom-format backup:
Reconcile restored state with real infrastructure
Keep the restored control plane stopped. Read the restored database only with a DB-operator credential; this is an offline recovery procedure, not a public CLI workflow:orphans response does not prove the provider account is empty. If identity is ambiguous, keep the
restored control plane stopped and resolve ownership manually; do not issue another create or
delete.
Create an evidence directory and move the three CSV files into it before starting a worker. Keep
database URLs and provider credentials out of the bundle:
- RunPod
- AWS EC2
- GCP Compute
- Kubernetes or KServe
provider_resource_id, provider, immutable revision, and replica ordinal;
compare external_key to the provider ownership tag/label or its provider-specific hash. Also prove
that no provider object with an InferCrane ownership marker is absent from restored state. A missing,
duplicate, empty, permission-denied, truncated, or unparseable result is not an empty inventory
and blocks reconciliation. In shared accounts, compare against a timestamped pre-drill baseline;
never require global zero or adopt another tenant’s resource.
After the identities agree, start exactly one restored control-plane worker. Only now check
readiness, query the API, watch durable operations/events, and confirm it adopts existing resources
rather than creating replacements:
Recover the active control plane after a failed migration
This recovery changes the database used by the control plane, not the serving resources. Keep the active runtime and provider capacity intact while all InferCrane application replicas are stopped.- Retain the failed/migrated database read-only and export the available post-upgrade evidence using the evidence-preserving rollback procedure.
-
Have the database operator create a separate empty rollback database. Set the TLS-protected
INFERCRANE_DATABASE_URLto that database, not to the failed database - and verify its name before restore. -
Restore the pre-upgrade dump:
- Keep every InferCrane replica stopped. Use the offline PostgreSQL exports and provider inventory procedure above to compare the restored active revision, candidate, nonterminal operations, provider identities, external keys, replica ordinals, and desired replica count with the failed-database exports and direct provider inventory. The pre-upgrade backup is authoritative through its timestamp; later provider state is an observation to reconcile, not permission to overwrite history.
-
Keep reconciliation and provider cleanup stopped while any identity is missing, duplicated, or
mismatched. Never run
delete, create a replacement deployment, or clear an orphan merely to make the restored database resemble current inventory. -
Only after every owned resource maps unambiguously to one restored replica intent, start exactly
one control-plane replica at the pre-upgrade application version. Do not start a mixed old/new
set. Wait for
/readyz, then collect API state: -
Confirm adoption without a provider create, send one bounded request through the unchanged active
endpoint, inspect its
X-Request-Id, and only then restore the normal same-version control-plane replica count.