{"id":3142,"date":"2026-09-23T14:11:04","date_gmt":"2026-09-23T14:11:04","guid":{"rendered":"https:\/\/dmarketertayeeb.com\/blog\/nvidia-tensorrt-multi-device-dynamo-triton-config\/"},"modified":"2026-09-23T14:11:04","modified_gmt":"2026-09-23T14:11:04","slug":"nvidia-tensorrt-multi-device-dynamo-triton-config","status":"publish","type":"post","link":"https:\/\/dmarketertayeeb.com\/blog\/nvidia-tensorrt-multi-device-dynamo-triton-config\/","title":{"rendered":"TensorRT Multi-GPU in Dynamo-Triton 26.07: Config and Benchmark Interpretation"},"content":{"rendered":"\n<p>This is a config-and-benchmark-interpretation guide for NVIDIA\u2019s September 21, 2026 Dynamo-Triton example. To serve a TensorRT network across several GPUs through one model endpoint, the model must already be compiled as a distributed TensorRT plan. NVIDIA\u2019s example uses a single <code>KIND_MODEL<\/code> instance, <code>enable_multi_device<\/code> set to <code>true<\/code>, and the participating GPU IDs in <code>multi_device_gpus<\/code>. That config activates a compatible plan; it does not turn an ordinary single-GPU engine into a multi-GPU one.<\/p>\n\n\n\n<p>For the published example, pin the <a href=\"https:\/\/developer.nvidia.com\/blog\/simplifying-model-serving-across-multiple-gpus-with-nvidia-tensorrt-multi-device-integration-in-nvidia-dynamo-triton\/\">NVIDIA Cosmos 3 Nano walkthrough<\/a> to Dynamo-Triton\/Triton 26.07 and TensorRT 11.1 as packaged in that release. The 26.07 release notes list Triton 2.71.0, TensorRT 11.1.0.106, CUDA 13.3.4.1 and NCCL 2.30.7. TensorRT multi-device inference became fully supported in TensorRT 11.0; TensorRT 10.16 treated it as a preview. The version distinction matters if you are comparing older examples or build images.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Check the version and GPU path first<\/h2>\n\n\n\n<p>NVIDIA says the Dynamo-Triton TensorRT backend added multi-device inference in release 26.07. The official release notes identify the server image contents and say the 26.07 container supports compute capability 7.5 or later. For the distributed-collective path in the Cosmos Ulysses example, the current TensorRT multi-device guide specifies Ampere-class GPUs (SM 80) or later. It lists a separate requirement for multi-device attention\u2014Blackwell (SM 100) or later\u2014but NVIDIA says its Cosmos example wraps standard attention with distributed collective layers and does not use that separate attention operator.<\/p>\n\n\n\n<p>Use the GPU, CUDA and driver combination supported by the exact container and plan you deploy. The blog says its benchmark ran on a healthy eight-GPU NVIDIA system but does not name the GPU SKU, so it is not evidence that a particular H100, B200 or other machine reproduced the numbers. Confirm the release\u2019s support matrix and the TensorRT plan\u2019s build\/runtime compatibility before allocating the model.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A Dynamo-Triton model config for a distributed plan<\/h2>\n\n\n\n<p>This is NVIDIA\u2019s CP8 <code>config.pbtxt<\/code> excerpt, with the eight GPU ordinals shown in the article. It is the model-serving configuration, not a TensorRT engine-building command:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>name: \"cosmos3_cp8\"\nbackend: \"tensorrt\"\nmax_batch_size: 0\n\ninstance_group [\n  { kind: KIND_MODEL count: 1 }\n]\nparameters [\n  { key: \"enable_multi_device\" value: { string_value: \"true\" } },\n  { key: \"multi_device_gpus\" value: { string_value: \"0,1,2,3,4,5,6,7\" } }\n]<\/code><\/pre>\n\n\n\n<ul class=\"wp-block-list\"><li><code>KIND_MODEL<\/code> with <code>count: 1<\/code> defines one TensorRT model instance that owns the selected devices. It is different from creating eight independent GPU instances.<\/li><li><code>enable_multi_device<\/code> opts that TensorRT model into the backend\u2019s multi-device path.<\/li><li><code>multi_device_gpus<\/code> lists the participating GPU IDs. NVIDIA\u2019s CP8 example lists IDs 0 through 7; use IDs visible to your serving process and keep the list length aligned with the distributed plan.<\/li><li><code>max_batch_size: 0<\/code> is part of this example\u2019s config. Do not copy it or the model name blindly if your model\u2019s I\/O and batch semantics differ.<\/li><\/ul>\n\n\n\n<p>Triton stores each model&#8217;s config beside a numeric version directory. For the example, the repository shape is:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>model_repository\/\n  cosmos3_cp8\/\n    config.pbtxt\n    1\/\n      model.plan<\/code><\/pre>\n\n\n\n<p>The default TensorRT plan filename is <code>model.plan<\/code>. The file in version <code>1<\/code> must be the matching distributed plan, not a regular single-device engine. NVIDIA&#8217;s blog does not give a universal export\/build command or a complete model I\/O config for arbitrary models; match input\/output names, shapes and types to your engine and the exact TensorRT backend.<\/p>\n\n\n\n<p>At request time, the application sends one inference request for the named model over the gRPC serving interface. Dynamo-Triton and the TensorRT backend manage the participating ranks and their execution contexts. The client does not issue a separate request to every GPU rank. See NVIDIA&#8217;s <a href=\"https:\/\/docs.nvidia.com\/deeplearning\/triton-inference-server\/user-guide\/docs\/user_guide\/model_repository.html\">Triton model-repository layout<\/a> for the version-folder convention; the multi-device fields remain the NVIDIA blog excerpt above.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What NVIDIA measured\u2014and what it did not<\/h2>\n\n\n\n<p>NVIDIA compared single-device (SD) serving with two-, four- and eight-GPU context-parallel runs for Cosmos 3 Nano. Every run generated 189 frames at 1280 \u00d7 720 and 24 frames per second, using 35 denoising steps. NVIDIA reports one warm-up followed by five measured complete generations on the same eight-GPU system. Timing includes prompt work, 70 transformer RPCs, classifier-free guidance and scheduler updates, VAE decode and frame postprocessing; model loading and MP4 encoding were excluded.<\/p>\n\n\n\n<figure class=\"wp-block-table\" style=\"overflow-x:auto\"><table style=\"min-width:700px\"><thead><tr><th>Run<\/th><th>GPUs<\/th><th>Mean end-to-end generation<\/th><th>E2E speedup<\/th><th>Mean transformer RPC time<\/th><th>RPC speedup<\/th><\/tr><\/thead><tbody><tr><td>SD<\/td><td>1<\/td><td>156.595 s<\/td><td>1.00\u00d7<\/td><td>146.192 s<\/td><td>1.00\u00d7<\/td><\/tr><tr><td>CP2<\/td><td>2<\/td><td>87.999 s<\/td><td>1.78\u00d7<\/td><td>77.548 s<\/td><td>1.89\u00d7<\/td><\/tr><tr><td>CP4<\/td><td>4<\/td><td>53.093 s<\/td><td>2.95\u00d7<\/td><td>42.661 s<\/td><td>3.43\u00d7<\/td><\/tr><tr><td>CP8<\/td><td>8<\/td><td>34.183 s<\/td><td>4.58\u00d7<\/td><td>23.993 s<\/td><td>6.09\u00d7<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>These are NVIDIA-reported latency results for one video-generation configuration, not an independent DMT benchmark. The eight-GPU run reduced the mean time for one complete generation from 156.595 to 34.183 seconds (4.58\u00d7 speedup). The measured transformer RPC component fell from 146.192 to 23.993 seconds (6.09\u00d7). The RPC share of end-to-end time also fell from 93.4% to 70.2%, leaving prompt work, scheduling, decoding and postprocessing more visible in the total.<\/p>\n\n\n\n<p>NVIDIA explicitly did not measure concurrent request throughput, cost per video or total cost of ownership. More GPUs per request can shorten latency in this workload, but these measurements do not show how many videos the system serves per second or what each video costs. For a deployment decision, benchmark your own model, resolution, sequence length, concurrency, queueing and service-level target, and measure hardware utilization and cost separately.<\/p>\n\n\n\n<p>The post also reports output validation across the same-seed SD, CP2, CP4 and CP8 runs. It sampled frames 0, 47, 94, 141 and 188 and checked temporal variation and image error; CP8 measured MAE 16.316 and PSNR 19.400 dB against the single-device result. NVIDIA does not claim pixel-identical outputs. Treat this as a model-specific validation example, not a universal quality guarantee for other graphs or models.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A short pre-deployment checklist<\/h2>\n\n\n\n<ol class=\"wp-block-list\"><li>Pin the Dynamo-Triton release and its bundled TensorRT\/CUDA\/NCCL versions; verify driver and GPU compatibility for the exact container.<\/li><li>Build or obtain a distributed TensorRT plan for the intended rank count. Confirm the graph contains the required collective operations; do not assume the serving config performs graph partitioning.<\/li><li>Match <code>multi_device_gpus<\/code> to the process-visible IDs and the rank count expected by the plan.<\/li><li>Start with the single-device and multi-device plan on the same inputs. Validate correctness and output quality before comparing timing.<\/li><li>Report the number of warm-ups and measured requests, input profile, timing boundaries, concurrency, and GPU count. Keep request latency, throughput and cost as separate metrics.<\/li><\/ol>\n\n\n\n<p>For a separate example of how to read NVIDIA performance claims without broadening a result beyond its tested workload, see DMT\u2019s <a href=\"https:\/\/dmarketertayeeb.com\/blog\/nvidia-vera-rubin-nvl72-mlperf-v6-1-results\">Vera Rubin MLPerf analysis<\/a>. That article covers a different benchmark and is not a Dynamo-Triton setup guide.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Sources and scope<\/h2>\n\n\n\n<ul class=\"wp-block-list\"><li><a href=\"https:\/\/developer.nvidia.com\/blog\/simplifying-model-serving-across-multiple-gpus-with-nvidia-tensorrt-multi-device-integration-in-nvidia-dynamo-triton\/\">NVIDIA: Simplifying Model Serving Across Multiple GPUs with TensorRT Multi-Device Integration in Dynamo-Triton<\/a> (Sep 21, 2026)<\/li><li><a href=\"https:\/\/docs.nvidia.com\/deeplearning\/triton-inference-server\/release-notes\/rel-26-07.html\">NVIDIA Triton Inference Server 26.07 release notes<\/a><\/li><li><a href=\"https:\/\/docs.nvidia.com\/deeplearning\/tensorrt\/latest\/getting-started\/release-notes-11\/11.0.0.html\">NVIDIA TensorRT 11.0 release notes<\/a><\/li><li><a href=\"https:\/\/docs.nvidia.com\/deeplearning\/tensorrt\/latest\/inference-library\/multi-device-inference.html\">NVIDIA TensorRT Multi-Device Inference documentation<\/a> (current page checked Sep 23, 2026)<\/li><\/ul>\n\n\n\n<p>All benchmark numbers and implementation details above are attributed to NVIDIA\u2019s published example. DMT did not build an engine, deploy the service or run the benchmark. The model config excerpt is source-derived and intended to show the multi-device fields; it is not a complete model repository. NVIDIA\u2019s Sep 21 post links to a backend guide under an archived 27.10 documentation path, which is later than this article\u2019s Sep 23 review date, so that page was not used to establish 26.07 behavior.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Configure a precompiled distributed TensorRT plan in Dynamo-Triton 26.07, then read NVIDIA&#8217;s latency results without mistaking them for throughput or cost.<\/p>\n","protected":false},"author":1,"featured_media":3141,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[180,178],"tags":[414,446,300,315],"class_list":["post-3142","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news","category-artificial-intelligence","tag-ai-evaluation","tag-ai-hardware","tag-ai-models","tag-developer-tools","has-featured-image"],"_links":{"self":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3142","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/comments?post=3142"}],"version-history":[{"count":0,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3142\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media\/3141"}],"wp:attachment":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media?parent=3142"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/categories?post=3142"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/tags?post=3142"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}