AI 推理性能基准测试

这是 NVIDIA 的数据中心深度学习产品性能中心,汇集可复现的 AI 性能基准数据,覆盖 NVIDIA 最新一代数据中心 GPU。

仅以原始运算速度作为唯一核心指标的时代已经过去。当下,行业更关注规模化场景下的吞吐量、能效与综合成本效益。随着 AI 从单次直接作答,逐步演进为支持多步推理,市场对推理能力及其底层经济性的需求持续上涨。这种转变会大幅提升算力需求,因为单次请求需要生成远更多 token。除吞吐量外,每瓦 token 数、每百万 token 成本、单用户每秒 token 数等指标同样至关重要。

对于受功耗约束的 AI 算力工厂,NVIDIA 持续迭代的软件优化能够长期提升 token 产出收益,凸显了我们技术创新的价值。

帕累托曲线 (Pareto curves) 直观展现,NVIDIA Blackwell 架构可在各类生产核心目标间实现最优均衡,包含:

  • 成本

  • 能效

  • 吞吐量

  • 响应速度

若仅针对单一场景优化系统,会限制部署灵活性,导致在曲线其他工况下效率不足。NVIDIA 的全栈设计思路,可在多种真实生产场景中兼顾能效与收益。Blackwell 的领先性源自极致的软硬件协同设计,这套全栈架构专为高性能、高能效与高可扩展性打造。

了解用于获得这些结果的方法论,并通过亲自执行 Benchmarking Recipes 学习如何复现这些测试。

MLPerf Inference v6.0 性能基准测试

Offline Scenario, Closed Division

Network Throughput GPU Server GPU Version QSL Size Target Accuracy Dataset
DeepSeek R1 2,494,310 tokens/sec 288x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
486,141 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
70,326 tokens/sec 8x B300 NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
58,582 tokens/sec 8x B200 Nebius B200 n1 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
gpt-oss 120B 1,046,150 tokens/sec 72x GB300 Nebius GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
879,542 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
111,496 tokens/sec 8x B300 Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) NVIDIA B300 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
93,071 tokens/sec 8x B200 LLM-D v0.5.0,Openshift 4.20.12,NVIDIA 8xB200-SXM-180GB NVIDIA B200 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
Qwen3-VL 235B 61 tokens/sec 4x GB300 NVIDIA GB300 NVL72 (4x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
44 tokens/sec 4x GB200 NVIDIA GB200 NVL72 (4x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
78 tokens/sec 8x B300 Nebius B300 n1 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
79 tokens/sec 8x B200 Dell B200,8xB200-SXM-180GB,RHEL 10.1,vLLM CentML:mlperf-inf-mm-q3vl-v6.0 NVIDIA B200 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
Llama3.1 405B 19,512 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) Subset of LongBench, LongDataCollections, Ruler, GovReport
15,462 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) Subset of LongBench, LongDataCollections, Ruler, GovReport
1,971 tokens/sec 8x B300 Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) NVIDIA B300 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) Subset of LongBench, LongDataCollections, Ruler, GovReport
1,350 tokens/sec 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) Subset of LongBench, LongDataCollections, Ruler, GovReport
Llama2 70B 1,126,850 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
888,054 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
112,954 tokens/sec 8x B300 NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
104,572 tokens/sec 8x B200 HPE ProLiant Compute XD685 (8x NVIDIA B200 180GB, TensorRT) NVIDIA B200 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
Llama3.1 8B 166,745 tokens/sec 8x B300 XA NB3I-E12 NVIDIA B300 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) CNN Dailymail (v3.0.0, max_seq_len=2048)
160,403 tokens/sec 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) CNN Dailymail (v3.0.0, max_seq_len=2048)
Wan2.2 0.037 samples/sec 4x GB300 NVIDIA GB300 NVL72 (4x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 248 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881) VBench prompts
0.027 samples/sec 4x GB200 NVIDIA GB200 NVL72 (4x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 248 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881) VBench prompts
0.059 samples/sec 8x B300 NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 248 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881) VBench prompts
0.046 samples/sec 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 248 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881) VBench prompts
DLRMv3 104,637 samples/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 34996 99% of FP32 and 99.9% of FP32 (AUC=80.31%) Synthetic Streaming 100B Dataset
10,737 samples/sec 8x B200 Camarero PDI200A2HG-810 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 34996 99% of FP32 and 99.9% of FP32 (WER=2.0671%) Synthetic Streaming 100B Dataset
Whisper 50,562 samples/sec 8x B300 NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 1633 99% of FP32 and 99.9% of FP32 (WER=2.0671%) LibriSpeech
49,327 samples/sec 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 1633 99% of FP32 and 99.9% of FP32 (WER=2.0671%) LibriSpeech

Server Scenario - Closed Division

Network Throughput GPU Server GPU Version QSL Size Target Accuracy MLPerf Server Latency
Constraints (ms)
Dataset
DeepSeek R1 1,555,110 tokens/sec 288x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
336,106 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
60,413 tokens/sec 8x B300 Nebius B300 n1 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
51,693 tokens/sec 8x B200 Nebius B200 n1 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
gpt-oss 120B 1,096,770 tokens/sec 72x GB300 Nebius GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
899,218 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
110,655 queries/sec 8x B300 Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) NVIDIA B300 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
87,444 tokens/sec 8x B200 Nebius B200 n1 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
Qwen3-VL 235B 43 tokens/sec 4x GB300 Nebius GB300 NVL72 (4x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
38 tokens/sec 4x GB200 NVIDIA GB200 NVL72 (4x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
45 queries/sec 8x B300 Nebius B300 n1 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
68 tokens/sec 8x B200 Dell B200,8xB200-SXM-180GB,RHEL 10.1,vLLM CentML:mlperf-inf-mm-q3vl-v6.0 NVIDIA B200 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
Llama3.1 405B 18,628 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) TTFT/TPOT: 6000 ms/175 ms Subset of LongBench, LongDataCollections, Ruler, GovReport
14,134 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) TTFT/TPOT: 6000 ms/175 ms Subset of LongBench, LongDataCollections, Ruler, GovReport
1,484 tokens/sec 8x B300 QuantaGrid D75H-10U (8x B300-SXM-270GB, TensorRT) NVIDIA B300 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) TTFT/TPOT: 6000 ms/175 ms Subset of LongBench, LongDataCollections, Ruler, GovReport
984 tokens/sec 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) TTFT/TPOT: 6000 ms/175 ms Subset of LongBench, LongDataCollections, Ruler, GovReport
Llama2 70B 868,278 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
810,104 tokens/sec 72x B200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
108,392 tokens/sec 8x B300 PowerEdge XE9780L (8x B300-SXM-270GB, TensorRT) NVIDIA B300 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
103,627 tokens/sec 8x B200 HPE ProLiant Compute XD685 (8x NVIDIA B200 180GB, TensorRT) NVIDIA B200 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
Llama3.1 8B 148,067 tokens/sec 8x B300 XA NB3I-E12 NVIDIA B300 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) TTFT/TPOT: 2000 ms/100 ms CNN Dailymail (v3.0.0, max_seq_len=2048)
131,270 queries/sec 8x B200 HPE ProLiant Compute XD685 (8x NVIDIA B200 180GB, TensorRT) NVIDIA B200 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) TTFT/TPOT: 2000 ms/100 ms CNN Dailymail (v3.0.0, max_seq_len=2048)
Wan2.2** 31 seconds 4x GB300 NVIDIA GB300 NVL72 (4x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 248 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162) N/A VBench prompts
40 seconds 4x GB200 NVIDIA GB200 NVL72 (4x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 248 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162) N/A VBench prompts
21 seconds 8x B300 G894-SD3-AAX7 NVIDIA B300 248 FID range: [23.01085758, 23.95007626] and CLIP range: [31.68631873, 31.81331801] N/A VBench prompts
25 seconds 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 248 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162) N/A VBench prompts
DLRMv3 99,997 queries/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 34996 99% of FP32 (AUC=80.31%) 80 ms Synthetic Streaming 100B Dataset
10,007 queries/sec 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 34996 FID range: [23.01085758, 23.95007626] and CLIP range: [31.68631873, 31.81331801] 80 ms Synthetic Streaming 100B Dataset

Interactive Scenario - Closed Division

Network Throughput GPU Server GPU Version QSL Size Target Accuracy MLPerf Server Latency
Constraints (ms)
Dataset
DeepSeek R1 250,634 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 1500 ms/15 ms mlperf_deepseek_r1
240,318 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 1500 ms/15 ms mlperf_deepseek_r1
4,935 tokens/sec 8x B300 G894-SD3-AAX7 NVIDIA B300 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 1500 ms/15 ms mlperf_deepseek_r1
gpt-oss 120B 677,199 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 6396 99% of 83.13% TTFT/TPOT: 2000 ms/20 ms AIME25, GPQA Diamond, LiveCodeBench v6
624,929 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 6396 99% of 83.13% TTFT/TPOT: 2000 ms/20 ms AIME25, GPQA Diamond, LiveCodeBench v6
26,006 tokens/sec 8x B300 XA NB3I-E12 NVIDIA B300 6396 99% of 83.13% TTFT/TPOT: 2000 ms/20 ms AIME25, GPQA Diamond, LiveCodeBench v6
13,155 tokens/sec 8x B200 Nebius B200 n1 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 6396 99% of 83.13% TTFT/TPOT: 2000 ms/20 ms AIME25, GPQA Diamond, LiveCodeBench v6
Llama3.1 405B 18,365 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) TTFT/TPOT: 4500 ms/80 ms Subset of LongBench, LongDataCollections, Ruler, GovReport
14,010 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) TTFT/TPOT: 4500 ms/80 ms Subset of LongBench, LongDataCollections, Ruler, GovReport
765 tokens/sec 8x B300 G894-SD3-AAX7 NVIDIA B300 8313 99% of FP16 ((GovReport + LongDataCollections + 65 Sample from LongBench)rougeL=21.6666, (Remaining samples of the dataset)exact_match=90.1335). Additionally, for both cases tokens per sample should be between than 90% and 110% of the reference (tokens_per_sample=684.68) TTFT/TPOT: 4500 ms/80 ms Subset of LongBench, LongDataCollections, Ruler, GovReport
Llama2 70B 814,128 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 450 ms/40 ms OpenOrca (max_seq_len=1024)
754,855 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 450 ms/40 ms OpenOrca (max_seq_len=1024)
70,724 tokens/sec 8x B300 PowerEdge XE9780L (8x B300-SXM-270GB, TensorRT) NVIDIA B300 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 450 ms/40 ms OpenOrca (max_seq_len=1024)
61,300 tokens/sec 8x B200 HPE ProLiant Compute XD685 (8x NVIDIA B200 180GB, TensorRT) NVIDIA B200 24576 99% of FP32 and 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, for both cases the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 450 ms/40 ms OpenOrca (max_seq_len=1024)
Llama3.1 8B 128,633 tokens/sec 8x B300 G894-SD3-AAX7 NVIDIA B300 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) TTFT/TPOT: 500 ms/30 ms CNN Dailymail (v3.0.0, max_seq_len=2048)
128,750 tokens/sec 8x B200 NVIDIA DGX B200 (8x B200-SXM-180GB, TensorRT) NVIDIA B200 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) TTFT/TPOT: 500 ms/30 ms CNN Dailymail (v3.0.0, max_seq_len=2048)

**The primary metric on Wan2.2 in Server Scenario is measured in seconds (lower the better).
MLPerf™ v6.0 Inference Closed Division. NVIDIA platform results from the following entries: 6.0-0006, 6.0-0010, 6.0-0024, 6.0-0039, 6.0-0040, 6.0-0048, 6.0-0062, 6.0-0072, 6.0-0073, 6.0-0074, 6.0-0075, 6.0-0076, 6.0-0077, 6.0-0078, 6.0-0080, 6.0-0081, 6.0-0083, 6.0-0084, 6.0-0085, 6.0-0089, 6.0-0091, 6.0-0094, 6.0-0098. MLPerf name and logo are trademarks. See https://mlcommons.org/ for more information.
For MLPerf™ various scenario data, click here
For MLPerf™ latency constraints, click here

NVIDIA 数据中心产品的推理性能

B200 推理性能

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
Stable Video Diffusion 1 7.32 videos/min - 8202.75 1x B200 DGX B200 26.02-py3 Mixed Synthetic TensorRT 10.15.1 NVIDIA B200
Stable Diffusion XL 1 2.89 images/sec - 507.41 1x B200 DGX B200 26.02-py3 FP8 Synthetic TensorRT 10.15.1 NVIDIA B200
BEVFusion Head 1 2,543 images/sec 6.03 images/sec/watt 0.39 1x B200 DGX B200 26.06-py3 INT8 Synthetic TensorRT 11.0.0.114 NVIDIA B200
Flux Image Generator 1 0.47 images/sec - 2130.4 1x B200 DGX B200 26.02-py3 FP4 Synthetic TensorRT 10.15.1 NVIDIA B200
HF Swin Base 128 5,220 samples/sec 5.69 samples/sec/watt 24.52 1x B200 DGX B200 26.04-py3 FP8 Synthetic TensorRT 10.16.1.11 NVIDIA B200
HF Swin Large 64 3,305 samples/sec 3.38 samples/sec/watt 19.36 1x B200 DGX B200 26.04-py3 FP8 Synthetic TensorRT 10.16.1.11 NVIDIA B200
HF ViT Base 512 9,879 samples/sec 10.38 samples/sec/watt 51.83 1x B200 DGX B200 26.04-py3 FP8 Synthetic TensorRT 10.16.1.11 NVIDIA B200
HF ViT Large 512 3,399 samples/sec 3.55 samples/sec/watt 301.21 1x B200 DGX B200 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 NVIDIA B200
Yolo v10 M 1 870 images/sec 1.16 images/sec/watt 1.15 1x B200 DGX B200 26.06-py3 INT8 Synthetic TensorRT 11.0.0.114 NVIDIA B200
Yolo v11 M 1 1,069 images/sec 1.38 images/sec/watt 0.94 1x B200 DGX B200 26.06-py3 INT8 Synthetic TensorRT 11.0.0.114 NVIDIA B200

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

RTX PRO 6000 Blackwell Server Edition Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
Stable Diffusion XL 1 1.05 images/sec - 954 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.01-py3 FP8 Synthetic TensorRT 10.14.1 RTX PRO 6000 BSE
Flux Image Generator 1 0.2 images/sec - 5072 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.01-py3 FP4 Synthetic TensorRT 10.14.1 RTX PRO 6000 BSE
BEVFusion Head 1 1738.51 images/sec 5 images/sec/watt 0.58 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.02-py3 FP8 Synthetic TensorRT 10.15.1 RTX PRO 6000 BSE
HF Swin Base 32 2,719 samples/sec 5 samples/sec/watt 11.77 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.02-py3 FP8 Synthetic TensorRT 10.15.1 RTX PRO 6000 BSE
HF Swin Large 32 1,517 samples/sec 3 samples/sec/watt 21.1 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.02-py3 FP8 Synthetic TensorRT 10.15.1 RTX PRO 6000 BSE
HF ViT Base 512 3,419 samples/sec 5.71 samples/sec/watt 149.732 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 RTX PRO 6000 BSE
HF ViT Large 512 1,161 samples/sec 1.93 samples/sec/watt 440.78 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 RTX PRO 6000 BSE
Yolo v11 M 1 465 images/sec 1 images/sec/watt 2.15 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.02-py3 FP8 Synthetic TensorRT 10.15.1 RTX PRO 6000 BSE

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

RTX PRO 4500 Blackwell Server Edition Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
Stable Diffusion XL 1 0.4 images/sec - 2514 1x RTX PRO 4500 Supermicro SYS-521GE-TNRT 26.01-py3 FP8 Synthetic TensorRT 10.14.1 RTX PRO 4500 BSE
Flux Image Generator 1 0.07 images/sec - 13816 1x RTX PRO 4500 Supermicro SYS-521GE-TNRT 26.01-py3 FP4 Synthetic TensorRT 10.14.1 RTX PRO 4500 BSE
HF Bert Large QAT 64 2,720 samples/sec - 24 1x RTX PRO 4500 Supermicro SYS-521GE-TNRT 26.01-py3 INT8 Synthetic TensorRT 10.14.1 RTX PRO 4500 BSE
HF Bert Large 64 1,507 samples/sec - 42 1x RTX PRO 4500 Supermicro SYS-521GE-TNRT 26.01-py3 Mixed Synthetic TensorRT 10.14.1 RTX PRO 4500 BSE
HF ViT Base 16 1,403 samples/sec - 11 1x RTX PRO 4500 Supermicro SYS-521GE-TNRT 26.01-py3 FP8 Synthetic TensorRT 10.14.1 RTX PRO 4500 BSE
HF ViT Large 4 449 samples/sec - 9 1x RTX PRO 4500 Supermicro SYS-521GE-TNRT 26.01-py3 FP8 Synthetic TensorRT 10.14.1 RTX PRO 4500 BSE

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

H100 Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
Stable Diffusion XL 1 1.54 images/sec - 780.31 1x H100 DGX H100 26.02-py3 FP8 Synthetic TensorRT 10.15.1 H100 SXM5-80GB
BEVFusion Head 1 2,016 images/sec 6.13 images/sec/watt 0.5 1x H100 DGX H100 26.06-py3 INT8 Synthetic TensorRT 11.0.0.114 H100 SXM5-80GB
HF Swin Base 128 2,967 samples/sec 4.29 samples/sec/watt 43.14 1x H100 DGX H100 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 H100 SXM5-80GB
HF Swin Large 128 1,831 samples/sec 2.65 samples/sec/watt 69.91 1x H100 DGX H100 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 H100 SXM5-80GB
HF ViT Base 2048 4,939 samples/sec 7.11 samples/sec/watt 414.68 1x H100 DGX H100 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 H100 SXM5-80GB
HF ViT Large 512 1,737 samples/sec 7.54 samples/sec/watt 294.79 1x H100 DGX H100 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 H100 SXM5-80GB
Yolo v10 M 1 405 images/sec 0.68 images/sec/watt 2.47 1x H100 DGX H100 26.04-py3 FP8 Synthetic TensorRT 10.16.1.11 H100 SXM5-80GB
Yolo v11 M 1 480 images/sec 0.76 images/sec/watt 2.08 1x H100 DGX H100 26.04-py3 FP8 Synthetic TensorRT 10.16.1.11 H100 SXM5-80GB

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

L40S Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
BEVFusion Head 1 1958 images/sec 7 images/sec/watt 0.51 1x L40S Supermicro SYS-521GE-TNRT 26.02-py3 INT8 Synthetic TensorRT 10.15.1 NVIDIA L40S
HF Swin Base 32 1,396 samples/sec 4 samples/sec/watt 22.92 1x L40S Supermicro SYS-521GE-TNRT 26.02-py3 FP8 Synthetic TensorRT 10.15.1 NVIDIA L40S
HF Swin Large 32 716 samples/sec 2 samples/sec/watt 44.72 1x L40S Supermicro SYS-521GE-TNRT 26.02-py3 FP8 Synthetic TensorRT 10.15.1 NVIDIA L40S
HF ViT Base 1024 1,629 samples/sec 4.76 samples/sec/watt 628.73 1x L40S Supermicro SYS-521GE-TNRT 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 NVIDIA L40S
HF ViT Large 2048 578 samples/sec 1.7 samples/sec/watt 3,546.06 1x L40S Supermicro SYS-521GE-TNRT 26.06-py3 FP8 Synthetic TensorRT 11.0.0.114 NVIDIA L40S
Yolo v10 M 1 275 images/sec 0.79 images/sec/watt 3.64 1x L40S Supermicro SYS-521GE-TNRT 26.02-py3 INT8 Synthetic TensorRT 10.15.1 NVIDIA L40S
Yolo v11 M 1 310 images/sec 0.9 images/sec/watt 3.23 1x L40S Supermicro SYS-521GE-TNRT 26.02-py3 INT8 Synthetic TensorRT 10.15.1 NVIDIA L40S

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

查看更多性能数据

训练至收敛

要在真实场景中部署 AI,需要将网络训练到在指定精度下收敛。
这是检验 AI 系统是否已准备好在实际环境中交付有价值结果的最有效方法。

AI Pipeline

NVIDIA Riva 是一个用于构建多模态对话式 AI 服务的应用框架,能够在 GPU 上提供高性能的实际运行体验。

NVIDIA 数据中心深度学习产品性能常见问题

只看计算单价或 FLOPs per dollar 会对推理 TCO 形成不完整的认知。对于 AI 推理 TCO,最重要的指标是每个 token 的成本,也就是实际交付的性价比。GB300 NVL72 在使用 Dynamo 和 TensorRT-LLM、单用户交互吞吐 116 TPS 的场景下,实现了每百万个 token 0.123 美元的推理成本——根据截至 2026 年 4 月的 SemiAnalysis InferenceX 基准测试,这是各大平台中每 token 成本最低的水平。

Metric
NVIDIA Hopper (HGX H200)
NVIDIA Blackwell (GB300 NVL72)
NVIDIA Blackwell 相对 Hopper 的倍数
单 GPU 每小时成本 (美元)
$1.41
$2.65
2 倍
每美元 FLOP (PFLOPS)
2.8
5.6
2 倍
单 GPU 每秒 token 数
90
6,000
65 倍
每兆瓦每秒 token 数
54K
2.8M
50 倍
每百万 token 成本 (美元)
$4.20
$0.12
降低 35 倍


GB300 NVL72 在使用 Dynamo 和 TensorRT-LLM、单用户交互吞吐 116 TPS 的场景下,实现了每百万个 token 0.123 美元的推理成本——根据截至 2026 年 4 月的 SemiAnalysis InferenceX 基准测试,这是各大平台中每 token 成本最低的水平。

NVIDIA 每百万 token 的推理成本在各代 GPU 中有了显著改善:根据 2026 年第一季度的 SemiAnalysis InferenceX 基准测试,在低延迟 agentic 工作负载上,得益于软硬件协同设计,NVIDIA Blackwell Ultra(GB300 NVL72)相较 NVIDIA Hopper 实现了每 MW 吞吐最高提升至 50 倍、每 token 成本最多降低至 35 倍。软件优化也带来持续改进——GB200 的 token 产出在三个月内提升了 4 倍,对应地每 token 成本也按比例下降。

NVIDIA 的 TensorRT-LLM 和 Dynamo 软件栈在无需更换硬件的前提下,持续带来推理成本优化。根据截至 2026 年 4 月的 SemiAnalysis InferenceX 基准测试,NVIDIA Blackwell B200 在 GPT-OSS-120B 模型上的每百万 token 成本,从发布时的 0.11 美元在两个月内降至 0.02 美元,单靠软件就实现了约 5 倍的改进。每个版本的 TensorRT-LLM 通常会通过算子/内核融合、量化改进以及调度优化等方式提升吞吐,从而进一步摊薄单位 token 的推理成本。