Skip to main content

From the first
GPU to an AI
Factory at scale

Design, deploy, and validate NVIDIA DGX and HGX systems—from a single node to an AI factory with high-speed network fabric, storage, and a production software stack.

COMPUTE / DGX-HGXSCALE UNITHPC STORAGE
ILLUSTRATIVE SYSTEM VIEW · NOT A FINAL BOM
PLATFORMSDGX · HGX
DEPLOYMENT SCALE1 → 256+ Nodes
H100HopperH200HopperB200BlackwellB300Blackwell UltraGB300NVL72

Choose the right deployment scale

Support for current NVIDIA GPU generations

H100, H200, B200, B300, and GB300 on DGX or HGX platforms. Final architecture and node count follow the selected generation, data center, and workload.

Start with one GPU server and expand to a multi-rack cluster that unifies compute, networking, storage, and software—with validation before handover.

PRODUCTION BASELINE01 GPU NODEPOWER · NETWORK · SOFTWARE
SYSTEM VIEW · SINGLE NODE
01 / SINGLE NODE

Production-Ready DGX or HGX

We help you bring your first NVIDIA DGX or HGX into production, from data center readiness to an LLM your apps can call.

  • Site ReadinessWe confirm your data center is ready before you buy, checking rack power budget, cooling capacity, weight, and network uplinks.
  • Rack & StackWe mount the system, spread power across multiple PDUs, and cable it to a labeled plan, so your team can maintain it easily after handover.
  • System SoftwareWe configure firmware, BIOS, BMC, OS, drivers, CUDA, and container runtime as one version-matched stack. That starts with tuning BIOS to NVIDIA and OEM guidance for GPUDirect and placing the BMC on an isolated management network. We then install the OS on RAID 1 and lock the kernel to the driver, so updates keep the system stable, and finish by enabling NVIDIA Fabric Manager for full-bandwidth NVSwitch and NVLink, plus DOCA for BlueField-3 and ConnectX-7/8.
  • Integration & PlatformWe set up Kubernetes, vLLM, and GPU monitoring so your apps can call LLMs on your own system through an OpenAI-compatible API. We configure the NVIDIA GPU Operator to schedule GPUs for each workload, tune vLLM to split large models across GPUs over NVLink with tensor parallelism, and connect DCGM Exporter to Prometheus and Grafana so you can see utilization, memory, and temperature for every GPU.
  • ValidationWe test every system before handover, from a hardware stress test to a live LLM API call. We run DCGM diagnostics, measure GPU-to-GPU bandwidth over NVLink with NCCL tests, and hand over vLLM throughput and latency as your production baseline.
RDMAFABRIC8–256+ NODESNCCL VERIFIED
SYSTEM VIEW · BasePOD & SuperPOD
02 / BasePOD & SuperPOD

SuperPOD for Scalable AI Factory

We build an integrated GPU cluster for distributed AI and HPC with compute, RDMA fabric, management, storage, and workload orchestration, so your team can train and serve models across nodes as one system.

Final node count and architecture depend on the selected DGX/HGX generation and NVIDIA reference architecture. The NVIDIA DGX SuperPOD name applies only to systems that meet its reference architecture.

  • Architecture & BOMWe design compute, control plane, fabric, storage, rack layout, and cabling. We lay out a rail-optimized compute fabric that gives every GPU its own NIC, separate from storage and management networks, and size power, cooling, and optics per rack, delivering a BOM and rack elevations ready for purchasing.
  • RDMA FabricWe deploy InfiniBand or Ethernet with RoCEv2 and lossless configuration. On InfiniBand we set up the Subnet Manager and adaptive routing; on Ethernet we tune PFC, ECN, and DCQCN for full-bandwidth GPU traffic, then check the link, cabling, and firmware of every NIC port before go-live.
  • Cluster SoftwareWe prepare OS images, drivers, NCCL, containers, and Slurm or Kubernetes. We roll out one golden image to every node through NVIDIA Base Command Manager or your provisioning system, then set up Slurm with Pyxis and Enroot for HPC, or Kubernetes with the GPU and Network Operators, so your team can submit distributed training from day one.
  • NCCL VerificationWe validate correctness, latency, and collective bandwidth against the agreed baseline. We run NCCL all-reduce and all-gather tests across every node against reference-architecture targets, follow with HPL and a real training job, and hand over the results as your baseline for future expansion.
DUAL FABRIC A / BNVL72 RACKSMULTI-RACK SCALEMANAGEMENT + WORKLOAD PLANEHPC STORAGE PATH
SYSTEM VIEW · RACK-SCALE NVL72
03 / RACK-SCALE NVL72

Grace Blackwell and Vera Rubin NVL72

We design and deploy rack-scale NVIDIA Grace Blackwell (GB300 NVL72) and Vera Rubin NVL72 systems, where 72 GPUs share one NVLink domain—from data center power, liquid cooling, and floorplan to networking, software, and acceptance—so your AI factory is ready for today's platform and the next.

Vera Rubin specifications and availability follow NVIDIA's announcements and vary by OEM. Final architecture, rack count, and power and cooling requirements are confirmed with NVIDIA, the OEM, and your data center.

  • Scalable-Unit PlanningWe plan scalable units, power, and liquid cooling for high-density NVL72 racks. With your data center team, we design the floorplan and point-to-point (P2P) cabling, from rack positions, hot and cold aisles, and expansion space to every cable route between racks and rows, and size CDUs and water loops for Vera Rubin's fully liquid-cooled design with 45°C inlet water.
  • Rack-Scale IntegrationWe install and bring up rack-scale systems where 72 GPUs share one NVLink domain—GB300 NVL72 on fifth-generation NVLink and Vera Rubin NVL72 on NVLink 6. We verify every NVLink switch tray and cable cartridge through NVIDIA NMX Manager, connect liquid cooling and power shelves to the data center's systems, and check flow, leaks, and power rack by rack before full load.
  • Vera Rubin ReadinessWe plan the path from Grace Blackwell to Vera Rubin in the same data hall, preparing power, cooling, and networking for third-generation MGX racks, ConnectX-9 SuperNICs, and BlueField-4, and set power budgets around rack-level power smoothing so you use the facility's full capacity.
  • Workload PlaneWe deploy management, Slurm, containers, monitoring, and secure access. We separate management, compute, storage, and in-band networks, set up Slurm or Kubernetes with per-team quotas, feed GPU, fabric, and storage metrics into central monitoring, and give each user group role-based secure access.
  • System ValidationWe validate compute, fabric, storage, and workload scheduling end to end. We test node, rack, and full cluster with DCGM diagnostics, NCCL tests within each NVLink domain and across scalable units, HPL, and storage throughput, then finish with a real training or inference workload before handover.
LOSSLESS GPU FABRICRoCEv2 · INFINIBAND
SYSTEM VIEW · NETWORK FABRIC
04 / GPU FABRIC

High-Speed GPU Network Fabric

We design topology, NICs or SuperNICs, switching, optics, cabling, and RDMA for distributed AI/HPC over Ethernet with RoCEv2 or InfiniBand, so every GPU across your nodes communicates at full bandwidth.

NVLink and NVSwitch connect GPUs within a system. GPUDirect RDMA moves data between GPU nodes across the network fabric with less CPU overhead.

  • Ethernet with RoCEv2We configure QoS, PFC/ECN, routing, and telemetry for a lossless fabric. We build a non-blocking leaf-spine, tune DCQCN and switch buffers for full-bandwidth all-reduce traffic, and stream per-port telemetry into monitoring so congestion shows up early.
  • InfiniBandWe configure subnet management, adaptive routing or SHARP, and monitoring on NVIDIA Mellanox QM Series switches. We design the fat-tree topology, deploy NVIDIA UFM or a subnet manager, enable SHARP so the switches offload collective operations from the GPUs, and check cabling, firmware, and error counters on every port.
  • GPU-Aware TuningWe tune PCIe, NIC, and GPU locality, GPUDirect RDMA, and NCCL topology. We pair each GPU with the NIC under the same PCIe switch, set ACS and IOMMU so GPUDirect RDMA moves data straight between GPU memory and the NIC, and configure NCCL to use every rail of the fabric.
  • Fabric AcceptanceWe validate health, RDMA latency, and NCCL bandwidth against the agreed baseline. We run perftest (ib_write_bw and ib_write_lat) and NCCL tests across message sizes alongside a full-fabric stress test, then hand over the results as the baseline for future troubleshooting.
HIGH-PERFORMANCE DATA PATHSTORAGEGPUWEKA · DDN · GDS
SYSTEM VIEW · HPC STORAGE
05 / HPC STORAGE

High-Performance Data for Every GPU

We deploy WEKA or DDN for AI and HPC data paths—from ingestion and training checkpoints to model data and simulation output—so data reaches your GPUs as fast as they can process it.

  • Data-Path AssessmentWe assess datasets, I/O patterns, concurrency, metadata, and growth. We measure file sizes, read and write rates, and checkpoint cadence from your real jobs to size capacity and performance for today, with room for more data.
  • Storage ArchitectureWe design performance tiers, protection, network paths, and namespace. We place training data on an NVMe hot tier with a capacity or object tier for colder data, plan the storage network around the compute fabric, and give every GPU node one shared namespace.
  • WEKA or DDNWe deploy the platform, clients, access controls, protection, and monitoring. We roll out WEKA or DDN EXAScaler with clients on every GPU node, set per-team quotas and permissions, configure snapshots or replication to your data policy, and feed storage metrics into central monitoring.
  • GPU AccelerationWe enable GPUDirect Storage when the hardware and software stack support it, opening a direct RDMA path from storage into GPU memory that eases CPU load and lifts data-loader throughput during training.
  • ValidationWe validate bandwidth, IOPS, metadata performance, and pipeline throughput. We run fio, mdtest, and elbencho against your real I/O patterns, time checkpoint writes and dataset reads from every GPU node, and hand over the results as your baseline.
AI + HPC USE CASES

GPU cluster and AI factory use cases

What organizations run on GPU clusters and AI factories, from AI to scientific computing. Each one stresses compute, networking, and storage differently, so we design from your real workload.

01

Private LLM & Sovereign AI

Run LLMs and Thai-language models on your own systems, with data kept inside your organization and country.

Inference
02

Enterprise AI Assistants

RAG and AI agents that answer from internal documents and carry out multi-step work.

Agentic AI
03

Training & Fine-Tuning

Train and fine-tune models on your own data across many GPU nodes, with fast checkpoints.

Training
04

Vision & Video Analytics

Analyze images and video from many cameras in real time, from factory inspection to site safety.

Vision AI
05

Simulation & Digital Twins

CFD, weather, seismic, and digital twins of plants or cities that need many GPUs working together.

Simulation
06

Life Sciences & Materials

Genomics, protein modeling, and AI-driven drug and materials discovery alongside HPC.

Research

Technology ecosystem for enterprise systems

Connect compute, networking, storage, and the software stack across leading enterprise platforms, aligned with your architecture standards and existing environment.

Start with your workload and target

Share the GPU generation, node count, data center, network, storage, and target workload. Zotect will help define the right scope and system direction.

Talk to our AI Infrastructure team →