Start with an API, add your own GPUs later
You do not need GPU infrastructure to get value from InferCrane. Start with a model API, add budgets and evidence, then introduce customer-operated capacity when measurements justify it. The application continues using one model name throughout.1. Connect the API
--from-env stores only the variable name. The variable and its value must exist in the control-plane
environment. InferCrane does not put the provider key in endpoint configuration, browser state, or
CLI output. This is non-billable configuration; it does not route a request.
2. Create the stable application endpoint
3. Call one endpoint
5. Add self-hosted capacity without changing the app
model="support-production". What changed is the governed serving
plan, not the application contract.
What this proves
- API-first onboarding without building GPU infrastructure first.
- One stable model alias across external and customer-operated compute.
- Provider keys remain server-side references.
- External data transmission and spend are explicit and bounded.
- Migration is an immutable candidate plus evidence, not a risky cutover script.