This document aims to introduce the configuration resources required to ensure the normal running of LLMs during LLM inference on Tencent Cloud TI-ONE Platform (TI-ONE). It is provided for your reference only.
Guide on Inference Resources for Built-in LLMs
Note:
1. For the inventory and pricing of each instance type, go to the Cloud Virtual Machine (CVM) console and see Instance Creation Guide. For the PNV6/HCCPNV6 instance types, you need to contact your designated Tencent Cloud sales representative for activation and purchase.
2. When deploying the DeepSeek V3 or R1 model, if you only need a low concurrency experience, you can opt for a single-node deployment. If you have high requirements for inference performance and context length, and you have sufficient computing resources, it is recommended to use a deployment of at least 2 nodes.
3. The inference resource configuration in the table below is slightly smaller than the CVM instance configuration because TI-ONE occupies a small amount of resources when managing CVM instances. For example, a CVM instance specification with 128 cores will have 125.6 cores available after the CVM instance is added to the resource group.
Built-in LLM | Model List | Recommended Inference Resource | |
| | CVM instance source: Select from CVM Instances (yearly/monthly subscription) | CVM instance source: Purchase from the TI-ONE Platform (pay-as-you-go) |
Hunyuan-Large | Hunyuan-Large-chat | Deployment mode: standard deployment (with nf4 quantization enabled) [Recommended configuration 1] CVM instance specifications: PNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards [Recommended configuration 2] CVM instance specifications: HCCPNV4h.48XLARGE1024 CVM instance configuration: 192C, 1024 GB, and 8 A100 GPU cards Inference resource configuration: 189C, 980 GB, and 8 A100 GPU cards | – |
Hunyuan-A13B | Hunyuan-A13B-Instruct | Deployment mode: standard deployment CVM instance specifications: PNV6.16XLARGE640 CVM instance configuration: 64C, 640 GB, and 4 GPU cards Inference resource configuration: 60C, 600 GB, and 4 GPU cards | Deployment mode: standard deployment CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640) |
DeepSeek series models | DeepSeek-V3.1-Terminus | [Minimum configuration] Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards [Recommended configuration] Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-V3.1 | [Minimum configuration] Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards [Recommended configuration] Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-R1-0528-AngelACC | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-R1-0528-AngelACC-PD | [Minimum configuration] Deployment mode: multi-machine distributed deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) Note: This model is deployed using a PD disaggregation architecture, with 2 instances configured as 1p1d, 3 instances as 2p1d, 4 instances as 2p2d, 5 instances as 3p2d, and so on. | - |
| DeepSeek-R1-0528 | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-R1-AngelACC | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-R1-AngelACC-PD | [Minimum configuration] Deployment mode: multi-machine distributed deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) Note: This model is deployed using a PD disaggregation architecture, with 2 instances configured as 1p1d, 3 instances as 2p1d, 4 instances as 2p2d, 5 instances as 3p2d, and so on. | - |
| DeepSeek-R1 | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Deployment mode: Multi-machine Distributed Deployment Number of nodes: 2 Inference resource configuration (per node): 380C, 2214 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-V3-0324-AngelACC | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-V3-0324-AngelACC-PD | [Minimum configuration] Deployment mode: multi-machine distributed deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) Note: This model is deployed using a PD disaggregation architecture, with 2 instances configured as 1p1d, 3 instances as 2p1d, 4 instances as 2p2d, 5 instances as 3p2d, and so on. | - |
| DeepSeek-V3-0324 | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards [Recommended configuration] Deployment mode: Multi-machine distributed deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards * (2 instances) | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-V3-AngelACC | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-V3 | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Number of nodes: 2 Inference resource configuration (per node): 380C, 2214 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-Prover-V2-7B | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards [Recommended configuration] Deployment mode: Multi-machine distributed deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-Prover-V2-671B | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards [Recommended configuration] Deployment mode: Multi-machine distributed deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| DeepSeek-R1-Distill-Qwen-1.5B | Deployment mode: standard deployment CVM instance specification: GNV4.3XLARGE44 CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card | Deployment mode: standard deployment CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card |
| DeepSeek-R1-Distill-Qwen-7B | Deployment mode: standard deployment CVM instance specification: GNV4.3XLARGE44 CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card | Deployment mode: standard deployment CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card |
| DeepSeek-R1-Distill-Qwen-14B | [Recommended configuration 1] Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C,160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card [Recommended configuration 2] Deployment mode: Standard deployment CVM instance specifications: PNV5b.8XLARGE96 CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card | [Recommended configuration 1] Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) [Recommended configuration 2] Deployment mode: standard deployment CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card |
| DeepSeek-R1-Distill-Qwen-32B-AngelACC | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| DeepSeek-R1-Distill-Qwen-32B | [Recommended configuration 1] Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card [Recommended configuration 2] Deployment mode: Standard deployment CVM instance specifications: PNV5b.16XLARGE192 CVM instance configuration: 64C, 192 GB, and 2 PNV5b GPU cards Inference resource configuration: 62C, 172 GB, and 2 PNV5b GPU cards | [Recommended configuration 1] Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) [Recommended configuration 2] Deployment mode: standard deployment CVM instance specifications: 64C, 192 GB, and 2 PNV5b GPU cards |
gpt-oss | gpt-oss-120b | Deployment mode: standard deployment CVM instance specification: PNV6.8XLARGE320 CVM instance configuration: 32C, 320 GB, and 2 GPU cards Inference resource configuration: 30C, 288 GB, and 2 GPU cards | Deployment mode: standard deployment CVM instance specifications: 32C, 320 GB, and 2 GPUs (PNV6.8XLARGE320) |
| gpt-oss-20b | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
Kimi-K2 | Kimi-K2-Instruct | Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | - |
| Kimi-K2-Instruct-0905 | Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | - |
Cosmos-Reason1 | Cosmos-Reason1 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
Qwen3 series models | Qwen3-0.6B | Deployment mode: standard deployment CVM instance specification: GNV4.3XLARGE44 CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card | Deployment mode: standard deployment CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card |
| Qwen3-1.7B | Deployment mode: standard deployment CVM instance specification: GNV4.3XLARGE44 CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card | Deployment mode: standard deployment CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card |
| Qwen3-4B | Deployment mode: standard deployment CVM instance specification: GNV4.3XLARGE44 CVM instance configuration: 12C, 44 GB, and 1 A10 GPU card Inference resource configuration: 11C, 35 GB, and 1 A10 GPU card | Deployment mode: standard deployment CVM instance specifications: 12C, 44 GB, and 1 A10 GPU card |
| Qwen3-8B | Deployment mode: standard deployment CVM instance specifications: PNV6/HCCPNV6 series Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-14B | Deployment mode: standard deployment CVM instance specifications: PNV6/HCCPNV6 series Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-30B-A3B | Deployment mode: standard deployment CVM instance specifications: PNV6/HCCPNV6 series Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-32B | Deployment mode: standard deployment CVM instance specifications: PNV6/HCCPNV6 series Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-235B-A22B | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| Qwen3-235B-A22B-Instruct-2507 | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| Qwen3-235B-A22B-Thinking-2507 | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
| Qwen3-0.6B-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-1.7B-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-4B-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-8B-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-14B-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-30B-A3B-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-32B-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| Qwen3-235B-A22B-FP8 | Deployment mode: standard deployment CVM instance specifications: PNV6.16XLARGE640 CVM instance configuration: 64C, 640 GB, and 4 GPU cards Inference resource configuration: 60C, 600 GB, and 4 GPU cards | Deployment mode: standard deployment CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640) |
| Qwen3-235B-A22B-Instruct-2507-FP8 | Deployment mode: standard deployment CVM instance specifications: PNV6.16XLARGE640 CVM instance configuration: 64C, 640 GB, and 4 GPU cards Inference resource configuration: 60C, 600 GB, and 4 GPU cards | Deployment mode: standard deployment CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640) |
| Qwen3-235B-A22B-Thinking-2507-FP8 | Deployment mode: standard deployment CVM instance specifications: PNV6.16XLARGE640 CVM instance configuration: 64C, 640 GB, and 4 GPU cards Inference resource configuration: 60C, 600 GB, and 4 GPU cards | Deployment mode: standard deployment CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640) |
Qwen3-Coder series models | Qwen3-Coder-480B-A35B-Instruct | Deployment mode: Multi-machine Distributed Deployment CVM instance specification: HCCPNV6.96XLARGE2304 CVM instance configuration: 384C, 2304 GB, and 8 GPU cards Inference resource configuration: 380C, 2214 GB, and 8 GPU cards (2 instances) | - |
| Qwen3-Coder-480B-A35B-Instruct-FP8 | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
QwQ series models | qwq_32b | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
Qwen series models | qwen-14b-base | CVM instance specifications: PNV5b.8XLARGE96 CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card Deployment mode: standard deployment | Deployment mode: standard deployment CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card |
| qwen-14b-chat | CVM instance specifications: PNV5b.8XLARGE96 CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card Deployment mode: standard deployment | Deployment mode: standard deployment CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card |
Qwen3-VL series models | Qwen3-VL-235B-A22B-Instruct | Deployment mode: standard deployment CVM instance specification: PNV6.32XLARGE1280 CVM instance configuration: 128C, 1280 GB, and 8 GPU cards Inference resource configuration: 125C, 1207 GB, and 8 GPU cards | Deployment mode: standard deployment CVM instance specifications: 128C, 1280 GB, and 8 GPUs (PNV6.32XLARGE1280) |
Qwen2.5-VL series models | Qwen2.5-VL-32B-Instruct | Deployment mode: standard deployment CVM instance specification: PNV6.8XLARGE320 CVM instance configuration: 32C, 320 GB, and 2 GPU cards Inference resource configuration: 30C, 288 GB, and 2 GPU cards | Deployment mode: standard deployment CVM instance specifications: 32C, 320 GB, and 2 GPUs (PNV6.8XLARGE320) |
| Qwen2.5-VL-72B-Instruct | Deployment mode: standard deployment CVM instance specifications: PNV6.16XLARGE640 CVM instance configuration: 64C, 640 GB, and 4 GPU cards Inference resource configuration: 60C, 600 GB, and 4 GPU cards | Deployment mode: standard deployment CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640) |
| Qwen2.5-VL-72B-Instruct-AWQ | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
Qwen3-Embedding series models | Qwen3-Embedding-8B | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
Baichuan2 series models | baichuan2-7b-base | CVM instance specifications: PNV4.7XLARGE116 CVM instance configuration: 28C, 116 GB, and 1 A10 GPU card Deployment mode: standard deployment Inference resource configuration: 24C, 96 GB, and 1 A10 GPU card | Deployment mode: standard deployment CVM instance specifications: 28C, 116 GB, and 1 A10 GPU card |
| baichuan2-7b-chat | CVM instance specifications: PNV4.7XLARGE116 CVM instance configuration: 28C, 116 GB, and 1 A10 GPU card Deployment mode: standard deployment Inference resource configuration: 24C, 96 GB, and 1 A10 GPU card | Deployment mode: standard deployment CVM instance specifications: 28C, 116 GB, and 1 A10 GPU card |
| baichuan2-13b-base | CVM instance specifications: PNV5b.8XLARGE96 CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card Deployment mode: standard deployment | Deployment mode: standard deployment CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card |
| baichuan2-13b-chat | CVM instance specifications: PNV5b.8XLARGE96 CVM instance configuration: 32C, 96 GB, and 1 PNV5b GPU card Inference resource configuration: 30C, 80 GB, and 1 PNV5b GPU card Deployment mode: standard deployment | Deployment mode: standard deployment CVM instance specifications: 32C, 96 GB, and 1 PNV5b GPU card |
Gemma 3 series models | gemma-3-27b-it | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
| gemma-3-12b-it | Deployment mode: standard deployment CVM instance specification: PNV6.4XLARGE160 CVM instance configuration: 16C, 160 GB, and 1 GPU card Inference resource configuration: 15C, 144 GB, and 1 GPU card | Deployment mode: standard deployment CVM instance specifications: 16C, 160 GB, and 1 GPU (PNV6.4XLARGE160) |
GLM-4.5V series models | GLM-4.5V | Deployment mode: standard deployment CVM instance specifications: PNV6.16XLARGE640 CVM instance configuration: 64C, 640 GB, and 4 GPU cards Inference resource configuration: 60C, 600 GB, and 4 GPU cards | Deployment mode: standard deployment CVM instance specifications: 64C, 640 GB, and 4 GPUs (PNV6.16XLARGE640) |