Guide on LLM Inference Resources

Last updated: 2026-03-03 17:44:43
This document aims to introduce the configuration resources required to ensure the normal running of LLMs during LLM inference on Tencent Cloud TI-ONE Platform (TI-ONE). It is provided for your reference only.

Guide on Inference Resources for Built-in LLMs

Note:
1. For the inventory and pricing of each instance type, go to the Cloud Virtual Machine (CVM) console and see Instance Creation Guide. For the PNV6/HCCPNV6 instance types, you need to contact your designated Tencent Cloud sales representative for activation and purchase.
2. When deploying the DeepSeek V3 or R1 model, if you only need a low concurrency experience, you can opt for a single-node deployment. If you have high requirements for inference performance and context length, and you have sufficient computing resources, it is recommended to use a deployment of at least 2 nodes.
3. The inference resource configuration in the table below is slightly smaller than the CVM instance configuration because TI-ONE occupies a small amount of resources when managing CVM instances. For example, a CVM instance specification with 128 cores will have 125.6 cores available after the CVM instance is added to the resource group.
Built-in LLM
Model List
Recommended Inference Resource
CVM instance source: Select from CVM Instances (yearly/monthly subscription)
CVM instance source: Purchase from the TI-ONE Platform (pay-as-you-go)
Hunyuan-Large
Hunyuan-Large-chat
Deployment mode: standard deployment (with nf4 quantization enabled)
[Recommended configuration 1]
CVM instance specifications: PNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards

[Recommended configuration 2]
CVM instance specifications: HCCPNV4h.48XLARGE1024
CVM instance configuration: 192C, 1024 GB, and 8 A100 GPU cards
Inference resource configuration: 189C, 980 GB, and 8 A100 GPU cards
Hunyuan-A13B
Hunyuan-A13B-Instruct
Deployment mode: standard deployment
CVM instance specifications: PNV6.16XLARGE640
CVM instance configuration: 64C, 640 GB, and 4 GPU cards
Inference resource configuration: 60C, 600 GB, and 4 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640)
DeepSeek series models
DeepSeek-V3.1-Terminus
[Minimum configuration]
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

[Recommended configuration]
Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)

DeepSeek-V3.1
[Minimum configuration]
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

[Recommended configuration]
Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)

DeepSeek-R1-0528-AngelACC
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-R1-0528-AngelACC-PD
[Minimum configuration] Deployment mode: multi-machine distributed deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Note: This model is deployed using a PD disaggregation architecture, with 2 instances configured as 1p1d, 3 instances as 2p1d, 4 instances as 2p2d, 5 instances as 3p2d, and so on.
-
DeepSeek-R1-0528
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-R1-AngelACC
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-R1-AngelACC-PD
[Minimum configuration] Deployment mode: multi-machine distributed deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Note: This model is deployed using a PD disaggregation architecture, with 2 instances configured as 1p1d, 3 instances as 2p1d, 4 instances as 2p2d, 5 instances as 3p2d, and so on.
-
DeepSeek-R1
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Deployment mode: Multi-machine Distributed Deployment
Number of nodes: 2
Inference resource configuration (per node): 380C, 2214 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)

DeepSeek-V3-0324-AngelACC
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-V3-0324-AngelACC-PD
[Minimum configuration] Deployment mode: multi-machine distributed deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Note: This model is deployed using a PD disaggregation architecture, with 2 instances configured as 1p1d, 3 instances as 2p1d, 4 instances as 2p2d, 5 instances as 3p2d, and so on.
-
DeepSeek-V3-0324
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

[Recommended configuration] Deployment mode: Multi-machine distributed deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards * (2 instances)
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-V3-AngelACC
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-V3
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Number of nodes: 2
Inference resource configuration (per node): 380C, 2214 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-Prover-V2-7B
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

[Recommended configuration] Deployment mode: Multi-machine distributed deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)

DeepSeek-Prover-V2-671B
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards

[Recommended configuration] Deployment mode: Multi-machine distributed deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
DeepSeek-R1-Distill-Qwen-1.5B
Deployment mode: standard deployment
CVM instance specification: GNV4.3XLARGE44
CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card
Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card
Deployment mode: standard deployment
CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card
DeepSeek-R1-Distill-Qwen-7B
Deployment mode: standard deployment
CVM instance specification: GNV4.3XLARGE44
CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card
Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card
Deployment mode: standard deployment
CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card
DeepSeek-R1-Distill-Qwen-14B
[Recommended configuration 1] Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C,160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card

[Recommended configuration 2] Deployment mode: Standard deployment
CVM instance specifications: PNV5b.8XLARGE96
CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card
Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card
[Recommended configuration 1]
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
[Recommended configuration 2]
Deployment mode: standard deployment
CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card
DeepSeek-R1-Distill-Qwen-32B-AngelACC
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
DeepSeek-R1-Distill-Qwen-32B
[Recommended configuration 1] Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card

[Recommended configuration 2] Deployment mode: Standard deployment
CVM instance specifications: PNV5b.16XLARGE192
CVM instance configuration: 64C, 192 GB, and 2 PNV5b GPU cards
Inference resource configuration: 62C, 172 GB, and 2 PNV5b GPU cards
[Recommended configuration 1]
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
[Recommended configuration 2]
Deployment mode: standard deployment
CVM instance specifications: 64C, 192 GB, and 2 PNV5b GPU cards
gpt-oss
gpt-oss-120b
Deployment mode: standard deployment
CVM instance specification: PNV6.8XLARGE320
CVM instance configuration: 32C, 320 GB, and 2 GPU cards
Inference resource configuration: 30C, 288 GB, and 2 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 32C, 320 GB, and 2 GPUs (PNV6.8XLARGE320)
gpt-oss-20b
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Kimi-K2
Kimi-K2-Instruct
Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
-
Kimi-K2-Instruct-0905
Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
-
Cosmos-Reason1
Cosmos-Reason1
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3 series models
Qwen3-0.6B
Deployment mode: standard deployment
CVM instance specification: GNV4.3XLARGE44
CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card
Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card
Deployment mode: standard deployment
CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card
Qwen3-1.7B
Deployment mode: standard deployment
CVM instance specification: GNV4.3XLARGE44
CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card
Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card
Deployment mode: standard deployment
CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card
Qwen3-4B
Deployment mode: standard deployment
CVM instance specification: GNV4.3XLARGE44
CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card
Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card
Deployment mode: standard deployment
CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card
Qwen3-8B
Deployment mode: standard deployment
CVM instance specifications: PNV6/HCCPNV6 series
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-14B
Deployment mode: standard deployment
CVM instance specifications: PNV6/HCCPNV6 series
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-30B-A3B
Deployment mode: standard deployment
CVM instance specifications: PNV6/HCCPNV6 series
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-32B
Deployment mode: standard deployment
CVM instance specifications: PNV6/HCCPNV6 series
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-235B-A22B
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
Qwen3-235B-A22B-Instruct-2507
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
Qwen3-235B-A22B-Thinking-2507
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
Qwen3-0.6B-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-1.7B-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-4B-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-8B-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-14B-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-30B-A3B-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-32B-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-235B-A22B-FP8
Deployment mode: standard deployment
CVM instance specifications: PNV6.16XLARGE640
CVM instance configuration: 64C, 640 GB, and 4 GPU cards
Inference resource configuration: 60C, 600 GB, and 4 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640)
Qwen3-235B-A22B-Instruct-2507-FP8
Deployment mode: standard deployment
CVM instance specifications: PNV6.16XLARGE640
CVM instance configuration: 64C, 640 GB, and 4 GPU cards
Inference resource configuration: 60C, 600 GB, and 4 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640)
Qwen3-235B-A22B-Thinking-2507-FP8
Deployment mode: standard deployment
CVM instance specifications: PNV6.16XLARGE640
CVM instance configuration: 64C, 640 GB, and 4 GPU cards
Inference resource configuration: 60C, 600 GB, and 4 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640)
Qwen3-Coder series models
Qwen3-Coder-480B-A35B-Instruct
Deployment mode: Multi-machine Distributed Deployment
CVM instance specification: HCCPNV6.96XLARGE2304
CVM instance configuration: 384C, 2304 GB, and 8 GPU cards
Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances)
-
Qwen3-Coder-480B-A35B-Instruct-FP8
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
QwQ series models
qwq_32b
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen series models
qwen-14b-base
CVM instance specifications: PNV5b.8XLARGE96
CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card
Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card
Deployment mode: standard deployment
Deployment mode: standard deployment
CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card
qwen-14b-chat
CVM instance specifications: PNV5b.8XLARGE96
CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card
Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card
Deployment mode: standard deployment
Deployment mode: standard deployment
CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card
Qwen3-VL series models
Qwen3-VL-235B-A22B-Instruct
Deployment mode: standard deployment
CVM instance specification: PNV6.32XLARGE1280
CVM instance configuration: 128C, 1280 GB, and 8 GPU cards
Inference resource configuration: 125C, 1207 GB, and 8 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280)
Qwen2.5-VL series models
Qwen2.5-VL-32B-Instruct
Deployment mode: standard deployment
CVM instance specification: PNV6.8XLARGE320
CVM instance configuration: 32C, 320 GB, and 2 GPU cards
Inference resource configuration: 30C, 288 GB, and 2 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 32C, 320 GB, and 2 GPUs (PNV6.8XLARGE320)
Qwen2.5-VL-72B-Instruct
Deployment mode: standard deployment
CVM instance specifications: PNV6.16XLARGE640
CVM instance configuration: 64C, 640 GB, and 4 GPU cards
Inference resource configuration: 60C, 600 GB, and 4 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640)
Qwen2.5-VL-72B-Instruct-AWQ
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Qwen3-Embedding series models
Qwen3-Embedding-8B
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
Baichuan2 series models
baichuan2-7b-base
CVM instance specifications: PNV4.7XLARGE116
CVM instance configuration: 28C, 116 GB, and 1 A10 GPU card
Deployment mode: standard deployment
Inference resource configuration: 24C, 96 GB, and 1 A10 GPU card
Deployment mode: standard deployment
CVM instance specifications: 28C, 116 GB, and 1 A10 GPU card
baichuan2-7b-chat
CVM instance specifications: PNV4.7XLARGE116
CVM instance configuration: 28C, 116 GB, and 1 A10 GPU card
Deployment mode: standard deployment
Inference resource configuration: 24C, 96 GB, and 1 A10 GPU card
Deployment mode: standard deployment
CVM instance specifications: 28C, 116 GB, and 1 A10 GPU card
baichuan2-13b-base
CVM instance specifications: PNV5b.8XLARGE96
CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card
Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card
Deployment mode: standard deployment
Deployment mode: standard deployment
CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card
baichuan2-13b-chat
CVM instance specifications: PNV5b.8XLARGE96
CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card
Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card
Deployment mode: standard deployment
Deployment mode: standard deployment
CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card
Gemma 3 series models
gemma-3-27b-it
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
gemma-3-12b-it
Deployment mode: standard deployment
CVM instance specification: PNV6.4XLARGE160
CVM instance configuration: 16C, 160 GB, and 1 GPU card
Inference resource configuration: 15C, 144 GB, and 1 GPU card
Deployment mode: standard deployment
CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160)
GLM-4.5V series models
GLM-4.5V
Deployment mode: standard deployment
CVM instance specifications: PNV6.16XLARGE640
CVM instance configuration: 64C, 640 GB, and 4 GPU cards
Inference resource configuration: 60C, 600 GB, and 4 GPU cards
Deployment mode: standard deployment
CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640)