Choose a model without accepting a black box
InferCrane’s Verified Models catalog is a small, reviewed set of deployment starting points. It is not a hosted-model marketplace and it does not imply that one model runs on every GPU, runtime, or provider.- the Hugging Face repository and immutable commit;
- publisher, task, protocol, capabilities, and modalities;
- license URL and gated-access requirement;
- reviewed runtime configuration and serving profiles;
- review date and an explicit evidence boundary.
Pick an operating objective, then measure it
Generation recipes may expose three reviewed vLLM candidate profiles:
These are configuration-verified candidates, not performance claims. The selected profile is written
into
infercrane.yaml, including the runtime version, immutable workload image and executable argv
when required, accelerator type and count, provider hint, replica bounds, and cache-aware routing.
Review or edit that ordinary DeploymentSpec before plan; editing a reviewed tuple creates an
explicitly unreviewed configuration that must earn its own evidence.
The candidates follow upstream vLLM optimization guidance,
which explicitly describes latency/throughput trade-offs for chunked prefill and requires workload-
specific measurement. Benchmark the exact model, GPU, runtime version, and traffic shape before
promotion; a value that helps one workload can hurt another or increase memory pressure.
Evidence levels
The current built-in catalog contains only configuration-verified entries. Measured results stay
tenant-scoped in benchmark history and immutable model recipes; InferCrane does not silently promote
them into universal claims.
General profiles for fast-moving runtimes
Serving profiles are not limited to InferCrane’s built-in vLLM image. A profile can carry a custom OCI image pinned by digest, a runtime version label, complete executable argv and probes, an exact GPU count, compatible GPU types, and a provider hint. The same revision contract flows through planning, provider compilation, pricing, capacity evidence, benchmarking, recommendations, and signed inference passports. Lifecycle code contains no model-specific branch.glm-5.3-flash and qwen3.8-flash-next are experimental catalog entries exercising this general
path. Each reviewed profile pins the model commit, a dedicated upstream vLLM image, exact GPU
topology, and complete executable argv. Both remain configuration-verified, not real-GPU or
production-qualified.
The primary GLM profile requests one native RunPod Pod with eight H200 GPUs and tensor parallelism,
matching the current upstream H200 recipe. A separate four-B200 Blackwell profile follows upstream’s
TP4 Blackwell guidance but remains an exact-tuple qualification candidate. The runpod-pods adapter
launches the pinned upstream image directly, without an SSH-derived wrapper.
The reviewed command first materializes the complete commit-pinned Hugging Face snapshot at a local
container path, then replaces the bootstrap process with vLLM. This provider-neutral argv pattern
works around a day-zero runtime ordering issue where multimodal processor metadata is opened before
the runtime’s ordinary Hub resolver runs; it is reusable by other custom OCI profiles that require a
filesystem model path. It deliberately uses a bounded 131K context starting point rather than
converting the model card’s declared 1M
maximum into an unmeasured production promise. See the official model card
and upstream vLLM recipe.
plan. Deployment still requires current RunPod
capacity, enough configured storage, paid-resource acknowledgement, and fresh qualification
evidence for the exact generated tuple. The FP8 snapshot is about 306 GiB before runtime overhead.
Use an appropriately sized identity-bound network volume or container disk; storage and topology
remain plan inputs, not inferred defaults.
The Qwen profile uses the separate FP8 checkpoint and the official eight-H200 TEP8 topology. Plain
TP8 is incompatible with that checkpoint’s quantization blocks, so the reviewed argv enables expert
parallelism and the Triton MoE backend. InferCrane bounds the starting context at 131K; native 262K
and one-million-token YaRN remain separate qualification targets. This is the experimental
Qwen3.8-Flash-Next open-weight preview, not the forthcoming production Qwen3.8-Flash QwenCloud
service. See the official model card and
upstream vLLM recipe.
plan. Real deployment requires current
RunPod capacity for an exact eight-H200 Pod and fresh protocol, memory, performance, and cleanup
evidence. The catalog also exposes custom-oci-blackwell-tp2 and
custom-oci-blackwell-tep4 as explicitly unqualified two- and four-B200 candidates for environments
where the H200 node shape is unavailable; selecting either does not create a support claim.
Current market watch
The public model story is intentionally broader than InferCrane’s small regression fixtures. As of 2026-08-23, three different signals point to the families teams are evaluating now:- Artificial Analysis ranks Kimi K3, Qwen3.8 A95B, and DeepSeek V4 Pro among the leading open-weight models by its Intelligence Index;
- OpenRouter’s usage ranking shows substantial current token traffic for Kimi, DeepSeek, and MiniMax families;
- official model cards and releases show active ecosystems around GLM-5.3, gpt-oss, Gemma 4, and Meta Llama.
Popularity is a discovery signal, not a market-share percentage or compatibility proof. OpenRouter
ranks token volume rather than users or spend, and model-hub downloads can include automation. The
leaderboard date and exact model revision must remain visible whenever InferCrane publishes a
comparison.
Frontier families also expose the central reason InferCrane exists: Kimi K3, Qwen3.8 A95B, DeepSeek
V4, and GLM-5.3 do not share one license, topology, runtime requirement, provider surface, or cost
profile. The application can retain one stable model identity while each backend binding keeps its
own qualification evidence.
Reviewed deployment collections
The built-in catalog follows application tasks rather than chasing a weekly popularity ranking. Selection considers current ecosystem adoption, an official upstream model card, license clarity, and compatibility with InferCrane’s qualified runtime line. Older compact models are useful here because they make frequent single-GPU lifecycle qualification affordable; they do not define the marketing story.
This is not a claim that every entry is best for its category. Evaluate quality on the application’s
own data, then benchmark the exact serving plan before promotion.
Catalog review uses the upstream Hugging Face model metadata, the
vLLM supported-model matrix, and each
publisher’s model card and license. Task-oriented discovery is informed by model-library patterns
used by Replicate and
Baseten, but InferCrane entries remain provider-neutral and do not
copy their performance or availability claims.
Popular models that require a newer runtime, a multi-node topology, unpublished weights, or custom
license handling remain available through an external binding or explicit-model path. Catalog
starting points provide defaults; they are not an allowlist and they do not imply measured
performance or provider availability.
Any Hugging Face model remains available
The catalog is convenience, not an allowlist:ModelArtifact identity during the durable deployment
lifecycle. A catalog entry simply starts with a reviewed commit and clearer license/protocol
metadata.
API
Authenticated consoles and automation can use:performance_claims: false. Trustworthy performance comes from
AIPerf benchmarking, Inference Lab, and the exact
evidence attached to a revision.
InferCrane does not mirror a third-party discovery catalog. Customer-visible models are reviewed
InferCrane products, while private deployment workflows resolve each artifact to an immutable
repository revision before any capacity can be created.