<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Nvidia on François Duthilleul</title><link>https://testdev.ovh/tags/nvidia/</link><description>Recent content in Nvidia on François Duthilleul</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 François Duthilleul</copyright><atom:link href="https://testdev.ovh/tags/nvidia/index.xml" rel="self" type="application/rss+xml"/><item><title>Confidential GPU</title><link>https://testdev.ovh/security/confidential-gpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/security/confidential-gpu/</guid><description>&lt;p&gt;A &lt;strong&gt;Confidential GPU&lt;/strong&gt; is a GPU whose memory, computation state, and data transfers are hardware-encrypted and isolated from the host system — extending the Trusted Execution Environment (TEE) boundary that technologies like TDX and SEV-SNP provide at the CPU level to encompass the GPU accelerator as well. Without that extension, a GPU sits outside the CPU TEE: it cannot read TEE memory, and any data offloaded to it would leave the confidential boundary. The primary implementation today is &lt;strong&gt;NVIDIA Confidential Computing&lt;/strong&gt;, introduced on &lt;strong&gt;Hopper&lt;/strong&gt; (H100), continued on &lt;strong&gt;Blackwell&lt;/strong&gt;, and planned for &lt;strong&gt;Rubin&lt;/strong&gt; as the third generation. In Confidential Computing mode, the GPU encrypts data resident in High Bandwidth Memory (HBM) under keys managed by its on-die security processor, and three hardware properties close the trust gap to the CPU TEE: an &lt;strong&gt;on-die Root of Trust&lt;/strong&gt; that verifies GPU firmware authenticity before the OS can talk to the device; &lt;strong&gt;device attestation&lt;/strong&gt; via the &lt;strong&gt;NVIDIA Remote Attestation Service (NRAS)&lt;/strong&gt;, which produces signed evidence that the GPU is genuine NVIDIA hardware in Confidential Computing mode with unmodified firmware (analogous to Intel DCAP or AMD KDS for CPU TEEs); and &lt;strong&gt;encrypted PCIe transfers&lt;/strong&gt; between the CPU TEE and GPU TEE at line rate using a hardware &lt;strong&gt;AES-256-GCM&lt;/strong&gt; implementation, so a host administrator, hypervisor, or co-tenant with DMA access to the bus sees only ciphertext. GPU attestation is verified alongside CPU attestation before secrets (model decryption keys, dataset credentials) are released to the combined CPU+GPU TEE. The technology requires no application code changes — existing TensorFlow, PyTorch, and CUDA workloads run unmodified inside the confidential boundary. The primary threat model is the same as CPU-level confidential computing (protecting data-in-use from the infrastructure operator) but applied to AI workloads: model intellectual property theft, training data exfiltration, and inference input/output interception during GPU computation.&lt;/p&gt;</description></item><item><title>CUDA (Compute Unified Device Architecture)</title><link>https://testdev.ovh/ai/cuda/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/cuda/</guid><description>&lt;p&gt;&lt;strong&gt;CUDA (Compute Unified Device Architecture)&lt;/strong&gt; is NVIDIA&amp;rsquo;s software platform for general-purpose computing on GPUs. Its objective is to give developers a familiar C/C++ (and Fortran, Python bindings) programming model with explicit control over &lt;strong&gt;kernels&lt;/strong&gt; (functions that run on the device), &lt;strong&gt;streams&lt;/strong&gt; (ordered queues of work), and &lt;strong&gt;memory spaces&lt;/strong&gt; (host, device, unified). CUDA sits above the GPU driver and below frameworks such as cuDNN, NCCL, and higher-level ML stacks; it is the layer that makes it practical to implement custom operators, HPC solvers, and inference engines that are not covered by a closed library. For AI, virtually every major training and inference stack ultimately depends on CUDA (or a CUDA-compatible runtime) on NVIDIA hardware.&lt;/p&gt;</description></item><item><title>cuDNN (CUDA Deep Neural Network library)</title><link>https://testdev.ovh/ai/cudnn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/cudnn/</guid><description>&lt;p&gt;&lt;strong&gt;cuDNN (CUDA Deep Neural Network library)&lt;/strong&gt; is NVIDIA’s library of highly optimized &lt;strong&gt;GPU kernels&lt;/strong&gt; for operations that dominate deep learning: convolutions, matrix multiplies used in attention, pooling, normalization (batch/layer), activations, and recurrent cells. Its objective is to deliver near-peak performance on &lt;strong&gt;CUDA&lt;/strong&gt;-capable GPUs without every framework author hand-writing assembly-tuned kernels. &lt;strong&gt;PyTorch&lt;/strong&gt;, TensorFlow, and many &lt;strong&gt;inference&lt;/strong&gt; engines call cuDNN (directly or via cuBLAS) under the hood for training and serving. cuDNN sits between raw &lt;strong&gt;CUDA&lt;/strong&gt; and application code; version alignment with the CUDA toolkit and driver is mandatory for supported deployments.&lt;/p&gt;</description></item><item><title>DOCA (Data Center Infrastructure on a Chip Architecture)</title><link>https://testdev.ovh/ai/doca/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/doca/</guid><description>&lt;p&gt;&lt;strong&gt;DOCA (Data Center Infrastructure on a Chip Architecture)&lt;/strong&gt; is NVIDIA&amp;rsquo;s software framework for building and operating services on &lt;strong&gt;BlueField DPUs&lt;/strong&gt;. Its objective is to standardize how operators and ISVs develop &lt;strong&gt;infrastructure applications&lt;/strong&gt;—OVS offload, firewall/VNF, storage targets, RDMA/RoCE control, TLS inspection, telemetry agents—on Arm cores and hardware accelerators embedded in the NIC, using a consistent set of libraries instead of ad hoc kernel modules on the host. DOCA spans drivers, userspace APIs, reference pipelines, and marketplace-packaged applications; it is the DPU counterpart to CUDA on GPUs, oriented toward I/O and packet processing rather than tensor math.&lt;/p&gt;</description></item><item><title>MIG (Multi-Instance GPU)</title><link>https://testdev.ovh/ai/mig/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/mig/</guid><description>&lt;p&gt;&lt;strong&gt;MIG (Multi-Instance GPU)&lt;/strong&gt; is an NVIDIA &lt;strong&gt;GPU&lt;/strong&gt; partitioning mode on datacenter accelerators (e.g. &lt;strong&gt;A100&lt;/strong&gt;, &lt;strong&gt;H100&lt;/strong&gt;) that splits one physical card into up to seven &lt;strong&gt;GPU instances (GIs)&lt;/strong&gt;, each with isolated &lt;strong&gt;streaming multiprocessors&lt;/strong&gt;, memory bandwidth, and &lt;strong&gt;HBM&lt;/strong&gt; capacity. The objective is &lt;strong&gt;higher utilization&lt;/strong&gt; in multi-tenant environments: several smaller models or dev/test workloads share one expensive GPU without time-slicing contention as severe as full-card sharing. Each MIG instance appears to the OS and &lt;strong&gt;CUDA&lt;/strong&gt; as a separate GPU with fixed resources; workloads cannot oversubscribe another instance’s memory. MIG suits &lt;strong&gt;inference&lt;/strong&gt; and modest training more often than massive single-job training that needs the entire GPU and &lt;strong&gt;NVLink&lt;/strong&gt; domain.&lt;/p&gt;</description></item><item><title>NCCL (NVIDIA Collective Communications Library)</title><link>https://testdev.ovh/ai/nccl/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/nccl/</guid><description>&lt;p&gt;&lt;strong&gt;NCCL (NVIDIA Collective Communications Library)&lt;/strong&gt; implements &lt;strong&gt;collective operations&lt;/strong&gt;—&lt;strong&gt;all-reduce&lt;/strong&gt;, &lt;strong&gt;broadcast&lt;/strong&gt;, &lt;strong&gt;reduce-scatter&lt;/strong&gt;, &lt;strong&gt;all-gather&lt;/strong&gt;, and others—optimized for &lt;strong&gt;NVIDIA GPUs&lt;/strong&gt; across &lt;strong&gt;NVLink&lt;/strong&gt; within a node and &lt;strong&gt;RDMA&lt;/strong&gt; (&lt;strong&gt;InfiniBand&lt;/strong&gt; or &lt;strong&gt;RoCE&lt;/strong&gt;) across nodes. Its objective in &lt;strong&gt;AI&lt;/strong&gt; is to make &lt;strong&gt;distributed training&lt;/strong&gt; and multi-GPU &lt;strong&gt;inference&lt;/strong&gt; (tensor parallelism) scale: gradient shards must merge every step; attention and MLP partitions must exchange activations with minimal latency. Frameworks (&lt;strong&gt;PyTorch&lt;/strong&gt; DDP/FSDP, &lt;strong&gt;vLLM&lt;/strong&gt; tensor parallel) call NCCL (or delegate to it) rather than hand-rolling socket code.&lt;/p&gt;</description></item><item><title>NIM (NVIDIA Inference Microservices)</title><link>https://testdev.ovh/ai/nim/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/nim/</guid><description>&lt;p&gt;&lt;strong&gt;NIM (NVIDIA Inference Microservices)&lt;/strong&gt; are &lt;strong&gt;container images&lt;/strong&gt; and Helm charts that deliver ready-to-run &lt;strong&gt;inference endpoints&lt;/strong&gt; for specific models (LLMs, vision, embedding, reranking, and more). The objective is to shrink time-to-production: instead of assembling CUDA drivers, frameworks, model weights, and an OpenAI-compatible server yourself, operators pull a NIM that bundles a performance-tuned engine (often &lt;strong&gt;TensorRT-LLM&lt;/strong&gt; or Triton-backed paths), default model artifacts or download hooks, health checks, and a stable HTTP/gRPC API. NIMs are sized for &lt;strong&gt;GPU&lt;/strong&gt; deployment and target enterprise MLOps teams that want versioned, scannable containers with predictable resource requests rather than bespoke notebooks turned into scripts.&lt;/p&gt;</description></item><item><title>NVLink</title><link>https://testdev.ovh/ai/nvlink/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/nvlink/</guid><description>&lt;p&gt;&lt;strong&gt;NVLink&lt;/strong&gt; is NVIDIA’s proprietary &lt;strong&gt;high-speed interconnect&lt;/strong&gt; between GPUs (and, on some platforms, between GPUs and CPUs) inside a server or across an &lt;strong&gt;NVLink switch&lt;/strong&gt; system (e.g. NVL72-class racks). Its objective is to move tensors—activations, gradients, &lt;strong&gt;KV cache&lt;/strong&gt; shards, or partial attention results—at much higher bandwidth and lower latency than &lt;strong&gt;PCIe&lt;/strong&gt; or general &lt;strong&gt;Ethernet&lt;/strong&gt;, so multi-GPU &lt;strong&gt;training&lt;/strong&gt; and large-model &lt;strong&gt;inference&lt;/strong&gt; (tensor parallelism) are not bottlenecked on the bus. NVLink domains define which GPUs can treat each other’s memory as peer-accessible for &lt;strong&gt;CUDA&lt;/strong&gt; and &lt;strong&gt;NCCL&lt;/strong&gt; without leaving the box.&lt;/p&gt;</description></item><item><title>TensorRT</title><link>https://testdev.ovh/ai/tensorrt/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://testdev.ovh/ai/tensorrt/</guid><description>&lt;p&gt;&lt;strong&gt;TensorRT&lt;/strong&gt; is NVIDIA’s SDK for &lt;strong&gt;optimizing and deploying&lt;/strong&gt; trained neural networks for &lt;strong&gt;inference&lt;/strong&gt; on NVIDIA &lt;strong&gt;GPUs&lt;/strong&gt;. Its objective is minimum latency and maximum throughput: it ingests a model (ONNX, TensorFlow, PyTorch export, or framework-specific parsers), applies graph optimizations (layer fusion, constant folding, kernel autotuning), selects precisions (&lt;strong&gt;FP32&lt;/strong&gt;, &lt;strong&gt;FP16&lt;/strong&gt;, &lt;strong&gt;INT8&lt;/strong&gt;, &lt;strong&gt;FP8&lt;/strong&gt;), and produces a &lt;strong&gt;serialized engine&lt;/strong&gt; executed by a lightweight runtime. For LLMs, &lt;strong&gt;TensorRT-LLM&lt;/strong&gt; extends this with attention-specific fusions, inflight batching, and multi-GPU serving patterns; many &lt;strong&gt;NIM&lt;/strong&gt; microservices bundle TensorRT-LLM–optimized engines rather than raw PyTorch loops.&lt;/p&gt;</description></item></channel></rss>