Skip to main content

PostgreSQL backup and restore drill

Set a TLS-protected INFERCRANE_DATABASE_URL, then create and validate a custom-format backup:
Restore only into a verified empty or disposable target first:
Do not restart InferCrane or send a request after the restore yet. Keep every application replica stopped and the restored database unreachable from the active control plane until the offline state and authoritative provider inventory reconcile exactly. The restore refuses a target-name mismatch or a database with live control-plane heartbeats, but that check does not prove external ownership. Record recovery-point and recovery-time measurements in the release evidence. Never test destructive restore behavior against the only production copy. Set an operational backup schedule from the required recovery point objective (RPO); InferCrane does not silently choose one. Measure restore plus reconciliation against the recovery time objective (RTO). A database restore is incomplete until owned external resources have been inventoried and reconciled without duplicate mutation.

Reconcile restored state with real infrastructure

Keep the restored control plane stopped. Read the restored database only with a DB-operator credential; this is an offline recovery procedure, not a public CLI workflow:
For every lifecycle-managed deployment, compare its persisted provider identity and ownership tags with the provider’s read-only inventory. A database backup can be older than the provider: an empty orphans response does not prove the provider account is empty. If identity is ambiguous, keep the restored control plane stopped and resolve ownership manually; do not issue another create or delete. Create an evidence directory and move the three CSV files into it before starting a worker. Keep database URLs and provider credentials out of the bundle:
Then capture authoritative provider inventory using the same scoped identity, account/project, region, and namespace as the deployment:
Before starting the single restored worker, require every restored replica row to map to exactly one provider object. Compare provider_resource_id, provider, immutable revision, and replica ordinal; compare external_key to the provider ownership tag/label or its provider-specific hash. Also prove that no provider object with an InferCrane ownership marker is absent from restored state. A missing, duplicate, empty, permission-denied, truncated, or unparseable result is not an empty inventory and blocks reconciliation. In shared accounts, compare against a timestamped pre-drill baseline; never require global zero or adopt another tenant’s resource. After the identities agree, start exactly one restored control-plane worker. Only now check readiness, query the API, watch durable operations/events, and confirm it adopts existing resources rather than creating replacements:
Wait for adoption and routing evidence to converge, then send one bounded inference request through the unchanged endpoint and inspect its request ID. Only after that passes may you restore the normal control-plane replica count. Provider-console or provider-API inventory remains a required external check because InferCrane cannot prove resources hidden by stale, incomplete, or mis-scoped provider credentials.

Recover the active control plane after a failed migration

This recovery changes the database used by the control plane, not the serving resources. Keep the active runtime and provider capacity intact while all InferCrane application replicas are stopped.
  1. Retain the failed/migrated database read-only and export the available post-upgrade evidence using the evidence-preserving rollback procedure.
  2. Have the database operator create a separate empty rollback database. Set the TLS-protected INFERCRANE_DATABASE_URL to that database, not to the failed database - and verify its name before restore.
  3. Restore the pre-upgrade dump:
  4. Keep every InferCrane replica stopped. Use the offline PostgreSQL exports and provider inventory procedure above to compare the restored active revision, candidate, nonterminal operations, provider identities, external keys, replica ordinals, and desired replica count with the failed-database exports and direct provider inventory. The pre-upgrade backup is authoritative through its timestamp; later provider state is an observation to reconcile, not permission to overwrite history.
  5. Keep reconciliation and provider cleanup stopped while any identity is missing, duplicated, or mismatched. Never run delete, create a replacement deployment, or clear an orphan merely to make the restored database resemble current inventory.
  6. Only after every owned resource maps unambiguously to one restored replica intent, start exactly one control-plane replica at the pre-upgrade application version. Do not start a mixed old/new set. Wait for /readyz, then collect API state:
  7. Confirm adoption without a provider create, send one bounded request through the unchanged active endpoint, inspect its X-Request-Id, and only then restore the normal same-version control-plane replica count.
If the active revision cannot be matched or post-backup provider mutations cannot be explained, leave the data plane in its current safe state and escalate manual recovery. A successful database restore alone does not prove lifecycle convergence. Backup/restore does not transfer tenant, organization, endpoint, or provider ownership. For credential rotation, same-tenant operator handoff, and the unsupported cross-tenant reassignment boundary, follow Ownership and credential transfer.