Skip to main content

Choose a model without accepting a black box

InferCrane’s Verified Models catalog is a small, reviewed set of deployment starting points. It is not a hosted-model marketplace and it does not imply that one model runs on every GPU, runtime, or provider.
Each entry exposes:
  • the Hugging Face repository and immutable commit;
  • publisher, task, protocol, capabilities, and modalities;
  • license URL and gated-access requirement;
  • reviewed runtime configuration and serving profiles;
  • review date and an explicit evidence boundary.
Create a project from one reviewed entry:
The browser console exposes the same control-plane catalog under Models. A model detail page shows the immutable identity and qualification scope before the deployment form is populated.

Pick an operating objective, then measure it

Generation recipes may expose three reviewed vLLM candidate profiles: These are configuration-verified candidates, not performance claims. The selected profile is written into infercrane.yaml, including the runtime version, immutable workload image and executable argv when required, accelerator type and count, provider hint, replica bounds, and cache-aware routing. Review or edit that ordinary DeploymentSpec before plan; editing a reviewed tuple creates an explicitly unreviewed configuration that must earn its own evidence. The candidates follow upstream vLLM optimization guidance, which explicitly describes latency/throughput trade-offs for chunked prefill and requires workload- specific measurement. Benchmark the exact model, GPU, runtime version, and traffic shape before promotion; a value that helps one workload can hurt another or increase memory pressure.

Evidence levels

The current built-in catalog contains only configuration-verified entries. Measured results stay tenant-scoped in benchmark history and immutable model recipes; InferCrane does not silently promote them into universal claims.

General profiles for fast-moving runtimes

Serving profiles are not limited to InferCrane’s built-in vLLM image. A profile can carry a custom OCI image pinned by digest, a runtime version label, complete executable argv and probes, an exact GPU count, compatible GPU types, and a provider hint. The same revision contract flows through planning, provider compilation, pricing, capacity evidence, benchmarking, recommendations, and signed inference passports. Lifecycle code contains no model-specific branch. glm-5.3-flash and qwen3.8-flash-next are experimental catalog entries exercising this general path. Each reviewed profile pins the model commit, a dedicated upstream vLLM image, exact GPU topology, and complete executable argv. Both remain configuration-verified, not real-GPU or production-qualified. The primary GLM profile requests one native RunPod Pod with eight H200 GPUs and tensor parallelism, matching the current upstream H200 recipe. A separate four-B200 Blackwell profile follows upstream’s TP4 Blackwell guidance but remains an exact-tuple qualification candidate. The runpod-pods adapter launches the pinned upstream image directly, without an SSH-derived wrapper. The reviewed command first materializes the complete commit-pinned Hugging Face snapshot at a local container path, then replaces the bootstrap process with vLLM. This provider-neutral argv pattern works around a day-zero runtime ordering issue where multimodal processor metadata is opened before the runtime’s ordinary Hub resolver runs; it is reusable by other custom OCI profiles that require a filesystem model path. It deliberately uses a bounded 131K context starting point rather than converting the model card’s declared 1M maximum into an unmeasured production promise. See the official model card and upstream vLLM recipe.
Those commands are local and non-mutating through plan. Deployment still requires current RunPod capacity, enough configured storage, paid-resource acknowledgement, and fresh qualification evidence for the exact generated tuple. The FP8 snapshot is about 306 GiB before runtime overhead. Use an appropriately sized identity-bound network volume or container disk; storage and topology remain plan inputs, not inferred defaults. The Qwen profile uses the separate FP8 checkpoint and the official eight-H200 TEP8 topology. Plain TP8 is incompatible with that checkpoint’s quantization blocks, so the reviewed argv enables expert parallelism and the Triton MoE backend. InferCrane bounds the starting context at 131K; native 262K and one-million-token YaRN remain separate qualification targets. This is the experimental Qwen3.8-Flash-Next open-weight preview, not the forthcoming production Qwen3.8-Flash QwenCloud service. See the official model card and upstream vLLM recipe.
Qwen3.8-Flash-Next uses the Qwen Community License 1.0, not an OSI open-source license. Its terms require a separate Qwen license for specified commercial Model-as-a-Service and AI work-assistant uses. Review the exact pinned license before deployment.
Those commands are also local and non-mutating through plan. Real deployment requires current RunPod capacity for an exact eight-H200 Pod and fresh protocol, memory, performance, and cleanup evidence. The catalog also exposes custom-oci-blackwell-tp2 and custom-oci-blackwell-tep4 as explicitly unqualified two- and four-B200 candidates for environments where the H200 node shape is unavailable; selecting either does not create a support claim.

Current market watch

The public model story is intentionally broader than InferCrane’s small regression fixtures. As of 2026-08-23, three different signals point to the families teams are evaluating now:
  • Artificial Analysis ranks Kimi K3, Qwen3.8 A95B, and DeepSeek V4 Pro among the leading open-weight models by its Intelligence Index;
  • OpenRouter’s usage ranking shows substantial current token traffic for Kimi, DeepSeek, and MiniMax families;
  • official model cards and releases show active ecosystems around GLM-5.3, gpt-oss, Gemma 4, and Meta Llama.
Popularity is a discovery signal, not a market-share percentage or compatibility proof. OpenRouter ranks token volume rather than users or spend, and model-hub downloads can include automation. The leaderboard date and exact model revision must remain visible whenever InferCrane publishes a comparison. Frontier families also expose the central reason InferCrane exists: Kimi K3, Qwen3.8 A95B, DeepSeek V4, and GLM-5.3 do not share one license, topology, runtime requirement, provider surface, or cost profile. The application can retain one stable model identity while each backend binding keeps its own qualification evidence.

Reviewed deployment collections

The built-in catalog follows application tasks rather than chasing a weekly popularity ranking. Selection considers current ecosystem adoption, an official upstream model card, license clarity, and compatibility with InferCrane’s qualified runtime line. Older compact models are useful here because they make frequent single-GPU lifecycle qualification affordable; they do not define the marketing story. This is not a claim that every entry is best for its category. Evaluate quality on the application’s own data, then benchmark the exact serving plan before promotion. Catalog review uses the upstream Hugging Face model metadata, the vLLM supported-model matrix, and each publisher’s model card and license. Task-oriented discovery is informed by model-library patterns used by Replicate and Baseten, but InferCrane entries remain provider-neutral and do not copy their performance or availability claims.
Popular models that require a newer runtime, a multi-node topology, unpublished weights, or custom license handling remain available through an external binding or explicit-model path. Catalog starting points provide defaults; they are not an allowlist and they do not imply measured performance or provider availability.
gpu_hint supplies the default accelerator. Profiles with compatible_gpus reject a different GPU during scaffolding because their runtime argv is hardware-specific. You can edit the generated YAML to experiment, but that is a new unreviewed configuration. Run plan, confirm memory/topology requirements, and qualify the exact provider/runtime combination before production.

Any Hugging Face model remains available

The catalog is convenience, not an allowlist:
Mutable references are resolved into ModelArtifact identity during the durable deployment lifecycle. A catalog entry simply starts with a reviewed commit and clearer license/protocol metadata.

API

Authenticated consoles and automation can use:
The response includes performance_claims: false. Trustworthy performance comes from AIPerf benchmarking, Inference Lab, and the exact evidence attached to a revision. InferCrane does not mirror a third-party discovery catalog. Customer-visible models are reviewed InferCrane products, while private deployment workflows resolve each artifact to an immutable repository revision before any capacity can be created.