{"id":3114,"date":"2026-09-23T06:08:32","date_gmt":"2026-09-23T06:08:32","guid":{"rendered":"https:\/\/dmarketertayeeb.com\/blog\/nvidia-b200-confidential-inference-benchmark\/"},"modified":"2026-09-23T16:58:28","modified_gmt":"2026-09-23T16:58:28","slug":"nvidia-b200-confidential-inference-benchmark","status":"publish","type":"post","link":"https:\/\/dmarketertayeeb.com\/blog\/nvidia-b200-confidential-inference-benchmark\/","title":{"rendered":"NVIDIA B200 Confidential Inference: CC-On vs CC-Off Benchmark"},"content":{"rendered":"\n<p><strong>Short answer:<\/strong> <a href=\"https:\/\/developer.nvidia.com\/blog\/enabling-private-high-performance-production-ai-inference-with-nvidia-confidential-computing\/\">NVIDIA reports<\/a> that confidential computing (CC) on an eight-GPU DGX B200 retained 96.1%\u201398.2% of CC-off output-token throughput, with 1.2%\u20134.3% mean time-per-output-token (TPOT) overhead. The result came from one specific DeepSeek-R1-0528-NVFP4 and TensorRT-LLM configuration. Use the same model, software, hardware, token lengths and concurrency on both sides of a CC-on\/CC-off comparison before applying the percentages to your own service.<\/p>\n\n\n\n<p>NVIDIA published the benchmark on September 22, 2026, alongside details on the TensorRT-LLM changes used to reduce confidential-inference overhead. This guide records the test matrix, gives a controlled comparison plan, and explains what the published result can and cannot tell an AI platform team.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What NVIDIA measured on one DGX B200<\/h2>\n\n\n\n<p>The NVIDIA performance-engineering team compared CC off with CC on while holding the model, hardware, framework version, sequence lengths, parallelism and concurrency constant. Its published workload and stack were:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Setting<\/th><th>NVIDIA configuration<\/th><\/tr><\/thead><tbody>\n<tr><td>GPU system<\/td><td>1 DGX B200 with 8 NVIDIA B200 GPUs<\/td><\/tr>\n<tr><td>CPU confidential-computing platform<\/td><td>Intel TDX<\/td><\/tr>\n<tr><td>Model<\/td><td><code><a href=\"https:\/\/huggingface.co\/nvidia\/DeepSeek-R1-0528-NVFP4\">nvidia\/DeepSeek-R1-0528-NVFP4<\/a><\/code><\/td><\/tr>\n<tr><td>Inference stack<\/td><td>TensorRT-LLM 1.3.0rc22, PyTorch backend; NCCL 2.30<\/td><\/tr>\n<tr><td>Input and output<\/td><td>32K input tokens \/ 1K output tokens<\/td><\/tr>\n<tr><td>Concurrent requests<\/td><td>1, 2, 4, 8 and 16<\/td><\/tr>\n<tr><td>Parallelism<\/td><td>Tensor parallelism 8; expert parallelism 1; pipeline parallelism 1<\/td><\/tr>\n<tr><td>KV cache<\/td><td>FP8<\/td><\/tr>\n<tr><td>Guest and host<\/td><td>Ubuntu 24.04.4 guest, 256 vCPUs and 2 NUMA nodes; Ubuntu 25.10 host<\/td><\/tr><tr><td>Host and guest kernels<\/td><td>Host 6.17.0-20-generic; guest 6.8.0-124-generic<\/td><\/tr><tr><td>GPU VBIOS and power limit<\/td><td>FW 1.4.x [97.10.64.00.0C]; 1,000 W<\/td><\/tr>\n<tr><td>Other reported versions<\/td><td>Driver 595.71.05; CUDA 13.2; OpenSSL 3.6.0; Docker with NVIDIA Container Toolkit<\/td><\/tr>\n<\/tbody><\/table><\/figure>\n\n\n\n<p>Across those five concurrency levels, NVIDIA reported the following normalized results. Throughput retained is output tokens per second with CC on divided by the CC-off baseline. TPOT overhead compares mean TPOT in the two runs; lower TPOT is better.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Concurrent requests<\/th><th>CC-on output throughput retained<\/th><th>CC-on mean TPOT overhead<\/th><\/tr><\/thead><tbody>\n<tr><td>1<\/td><td>98.2%<\/td><td>1.2%<\/td><\/tr>\n<tr><td>2<\/td><td>98.0%<\/td><td>3.2%<\/td><\/tr>\n<tr><td>4<\/td><td>96.1%<\/td><td>4.3%<\/td><\/tr>\n<tr><td>8<\/td><td>96.8%<\/td><td>3.5%<\/td><\/tr>\n<tr><td>16<\/td><td>96.8%<\/td><td>3.1%<\/td><\/tr>\n<\/tbody><\/table><\/figure>\n\n\n\n<p>These are NVIDIA-reported results for the listed DeepSeek workload, not an independent DMT benchmark. The throughput percentages are relative to NVIDIA\u2019s CC-off baseline; the original figure also labels absolute output tokens per second, reproduced below.<\/p>\n\n\n\n<figure><img decoding=\"async\" src=\"https:\/\/dmarketertayeeb.com\/blog\/wp-content\/uploads\/2026\/09\/b200-throughput.png\" alt=\"NVIDIA B200 output throughput with confidential computing off and on at five concurrency levels; exact values follow in the table.\" width=\"1400\" height=\"1130\" style=\"max-width:100%;height:auto\" loading=\"lazy\"><figcaption>DMT chart from <a href=\"https:\/\/developer.nvidia.com\/blog\/enabling-private-high-performance-production-ai-inference-with-nvidia-confidential-computing\/\">NVIDIA\u2019s September 22 benchmark figure<\/a>. Vendor-reported results for the configuration above; no independent DMT measurement.<\/figcaption><\/figure><div style=\"overflow-x:auto\"><table><thead><tr><th>Concurrent requests<\/th><th>CC off, output tok\/s<\/th><th>CC on, output tok\/s<\/th><\/tr><\/thead><tbody><tr><td>1<\/td><td>253.9<\/td><td>249.3<\/td><\/tr><tr><td>2<\/td><td>374.9<\/td><td>367.3<\/td><\/tr><tr><td>4<\/td><td>529.5<\/td><td>508.9<\/td><\/tr><tr><td>8<\/td><td>651.9<\/td><td>631.2<\/td><\/tr><tr><td>16<\/td><td>770.8<\/td><td>746.5<\/td><\/tr><\/tbody><\/table><\/div>\n\n\n\n<h2 class=\"wp-block-heading\">How to set up a controlled CC-on vs CC-off comparison<\/h2>\n\n\n\n<p>A useful comparison changes the confidential-computing state and keeps the rest of the serving stack fixed. Start on a supported B200 system with an approved TDX or SEV-SNP platform, a confidential VM, a pinned model revision, and the same guest, driver, CUDA, container, framework and NCCL versions for both passes. Follow NVIDIA\u2019s <a href=\"https:\/\/docs.nvidia.com\/nvtrust\/index.html\">current compatibility and deployment documentation<\/a> for the exact host and firmware you have.<\/p>\n\n\n\n<p>The unexecuted manifest below is a measurement template, not an executable benchmark command:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>{\n  \"model\": \"nvidia\/DeepSeek-R1-0528-NVFP4\",\n  \"backend\": \"TensorRT-LLM \/ PyTorch\",\n  \"input_tokens\": \"32K\",\n  \"output_tokens\": \"1K\",\n  \"concurrency\": [1, 2, 4, 8, 16],\n  \"tensor_parallel\": 8,\n  \"expert_parallel\": 1,\n  \"pipeline_parallel\": 1,\n  \"kv_cache\": \"FP8\",\n  \"comparison_states\": [\"CC off\", \"CC on\"],\n  \"keep_constant\": [\n    \"GPU host and GPU count\",\n    \"CPU TEE, guest image, vCPU and NUMA layout\",\n    \"model files and tokenizer revision\",\n    \"driver, CUDA, TensorRT-LLM, NCCL and container\",\n    \"request set, input\/output lengths and concurrency\"\n  ],\n  \"record\": [\n    \"output tokens per second\",\n    \"mean and p50\/p95\/p99 TPOT\",\n    \"TTFT, request errors and GPU\/CPU utilization\",\n    \"CC state and attestation result\"\n  ]\n}<\/code><\/pre>\n\n\n\n<p>For each concurrency point, run the same requests in both states and compute the two comparisons:<\/p>\n\n\n\n<ul class=\"wp-block-list\"><li><strong>Throughput retained (%)<\/strong> = 100 \u00d7 (CC-on output tokens per second \u00f7 CC-off output tokens per second)<\/li><li><strong>Mean TPOT overhead (%)<\/strong> = 100 \u00d7 ((CC-on mean TPOT \u00f7 CC-off mean TPOT) \u2212 1)<\/li><\/ul>\n\n\n\n<p>For a reader-run test, warm both configurations before collecting measurements, repeat each point, and report the spread as well as the mean. Add your production request-rate distribution, context lengths, output lengths and serving SLA after reproducing the reference shape. Those controls help show whether a slowdown is repeatable and whether the benchmark resembles your service; NVIDIA\u2019s post does not publish its full run count, warm-up procedure or raw request traces.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Check the GPU state before each pass<\/h3>\n\n\n\n<p><a href=\"https:\/\/docs.nvidia.com\/cc-deployment-guide-tdx.pdf\">NVIDIA\u2019s deployment guide<\/a> documents privileged CC mode changes and a reset after changing modes. The following per-GPU commands show the guide\u2019s flags; replace <code>xx:00.0<\/code> with the GPU\u2019s actual PCI bus\/device address, repeat for every B200 in the test, and use the approved helper path for your host. These examples are unexecuted. CC mode changes are system-administration actions, reset the GPU and persist across reboots, so perform them on an isolated test system with a maintenance plan.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code># Set one GPU to the baseline state, then reboot\/reset as required by the platform.\nsudo python3 .\/nvidia_gpu_tools.py --set-cc-mode=off --reset-after-cc-mode-switch --gpu-bdf=xx:00.0\n\n# Set the same GPU to confidential-computing mode for the paired pass.\nsudo python3 .\/nvidia_gpu_tools.py --set-cc-mode=on --reset-after-cc-mode-switch --gpu-bdf=xx:00.0\n\n# Inside the guest, record the reported configuration before benchmarking.\nnvidia-smi conf-compute -q<\/code><\/pre>\n\n\n\n<p>Confirm the CC status and complete the required attestation flow before treating a CC-on run as a security test. A throughput benchmark alone does not verify a threat model or prove that an attestation policy is configured correctly. NVIDIA\u2019s deployment guide also notes that mode settings may need to be applied from the guest when host Secure Boot policy prevents the host-side operation; keep your existing security policy intact and follow your organization\u2019s approved deployment path.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Why TensorRT-LLM needs CC-aware changes<\/h2>\n\n\n\n<p>Confidential execution changes how data moves between the CPU and GPU and how the runtime measures work. NVIDIA describes four adaptations in the B200 test stack:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Path<\/th><th>What changes under CC<\/th><th>TensorRT-LLM response described by NVIDIA<\/th><\/tr><\/thead><tbody>\n<tr><td>Host to GPU<\/td><td>Transfers pass through an encrypted bounce buffer because the GPU cannot directly access protected confidential-VM memory. Pinned memory may lose its usual asynchronous-transfer advantage.<\/td><td>Choose pageable memory on affected paths instead of always pinning memory.<\/td><\/tr>\n<tr><td>GPU to host<\/td><td>Protected token and sampling-data copies can block a caller during decode.<\/td><td>Move repeated readback to an asynchronous worker so the main scheduler can keep preparing decode work.<\/td><\/tr>\n<tr><td>Kernel autotuning<\/td><td>CUDA-event timestamps were unstable in the tested CC configuration and could lead the autotuner to choose a slower tactic.<\/td><td>Use the GPU <code>%globaltimer<\/code> for tactic timing in CC while retaining CUDA events outside CC.<\/td><\/tr>\n<tr><td>Multi-GPU communication<\/td><td>NVIDIA says NVLS multicast is unavailable in B200 CC mode. An NCCL symmetric path can still add registration and synchronization costs before falling back to a non-multicast collective.<\/td><td>Check availability and select a communication algorithm for the message size, topology and workload.<\/td><\/tr>\n<\/tbody><\/table><\/figure>\n\n\n\n<p>The pinned-memory and asynchronous-worker changes are visible in NVIDIA\u2019s merged <a href=\"https:\/\/github.com\/NVIDIA\/TensorRT-LLM\/pull\/11573\">TensorRT-LLM PR #11573<\/a>; the autotuner change is described in <a href=\"https:\/\/github.com\/NVIDIA\/TensorRT-LLM\/pull\/11657\">PR #11657<\/a>. Framework version matters: record the exact container tag and configuration you ran, rather than assuming another TensorRT-LLM build includes identical CC behavior.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How to read the result without overgeneralizing<\/h2>\n\n\n\n<p>The 96.1%\u201398.2% throughput-retention range describes this DeepSeek-R1-0528-NVFP4 configuration, not every B200 model, framework, context length or load pattern. In NVIDIA\u2019s rows, concurrency 4 has the lowest throughput retained (96.1%) and the highest mean TPOT overhead (4.3%). That is a useful example of why teams should report the whole concurrency curve instead of carrying a single average into capacity planning.<\/p>\n\n\n\n<p><a href=\"https:\/\/arxiv.org\/abs\/2608.26575\">A separate August 27 arXiv preprint<\/a> also tests TDX plus NVIDIA CC on B200, using multiple models and frameworks; it reports roughly 1%\u20133% overhead for its well-configured inference cases. <a href=\"https:\/\/tinfoil.sh\/blog\/2026-06-23-confidential-computing-overhead\">Tinfoil\u2019s June benchmark<\/a> uses different 8\u00d7B300 and H200 configurations. Keep each result attached to its model, software and traffic matrix: the percentages are not a pooled estimate and neither study is a reproduction of NVIDIA\u2019s September 22 DeepSeek run.<\/p>\n\n\n\n<p>NVIDIA\u2019s September post is useful for identifying the stack, comparison formulas and runtime changes. It does not include an executable benchmark command, a reported repeat count or confidence intervals. The results are vendor-run and DMT did not independently measure them. A B200 performance result also says nothing by itself about model quality, deployment cost, application tail latency or whether your attestation and key-release controls meet a particular policy.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Related DMT coverage<\/h2>\n\n\n\n<p>For a separate look at how to read vendor accelerator comparisons, see DMT\u2019s <a href=\"https:\/\/dmarketertayeeb.com\/blog\/nvidia-vera-rubin-nvl72-mlperf-v6-1-results\">NVIDIA Vera Rubin MLPerf analysis<\/a>. For enterprise questions about where data, control and operational responsibility sit in a private-AI deployment, see the <a href=\"https:\/\/dmarketertayeeb.com\/blog\/cohere-aleph-alpha-transatlantic-sovereign-ai\">sovereign AI controls checklist<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>NVIDIA reports 96.1%\u201398.2% throughput retained on one B200 confidential-inference workload. See the exact stack, CC-on\/CC-off formulas, TensorRT-LLM changes and limits before benchmarking your own traffic.<\/p>\n","protected":false},"author":1,"featured_media":3113,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[180,178],"tags":[300,272,441],"class_list":["post-3114","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news","category-artificial-intelligence","tag-ai-models","tag-ai-privacy","tag-llm-inference","has-featured-image"],"_links":{"self":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3114","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/comments?post=3114"}],"version-history":[{"count":1,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3114\/revisions"}],"predecessor-version":[{"id":3147,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3114\/revisions\/3147"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media\/3113"}],"wp:attachment":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media?parent=3114"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/categories?post=3114"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/tags?post=3114"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}