Skip to main content

LLMOps& AI Platforms

Model Serving & Inference on your GPU infrastructure

Extend AI Infrastructure into model APIs, from DGX Spark and DGX Station GB300 to SuperPOD, with Kubernetes / Slurm, RDMA and shared GPU resource management.

Application APIModel endpointModel serving & inferenceGPUGPUGPURDMA fabric · GPU workers
GPU CLUSTERKubernetes / Slurm
PROVISIONING / QUOTASBCM / Run:ai
K8s operatorsRDMA networkMonitoring
System components · Adapted to your use case and data

Model Serving
& Inference on GPU

Extend your installed GPUs with a software stack for model APIs, team resource allocation and measured acceptance before handover

Connect with AI Infrastructure ↗
Model endpoint

An API, model version and configuration for application teams

Resource policy

Reviewable access, quotas and tenant boundaries

Acceptance baseline

Benchmark report, dashboards, runbook and rollback procedure

From a desktop to an inference cluster

Start with models and usage patterns, then size GPUs, networking and replicas

Local inferenceGPU clusterMulti-rack
Conceptual growth from one system to a cluster and multiple racks

Which models can you start serving on each DGX?

Starting candidates based on memory and precision. Confirm the checkpoint, engine and license before deployment.

DGX Spark

1 × GB10 Grace Blackwell

Memory
128 GB LPDDR5x unified
Bandwidth
273 GB/s
Peak compute
1 PFLOP · FP4 sparse
NVIDIA specifications ↗

Prototype RAG and private assistants · 1 system

Starting models / precision

Qwen3-32B · FP4 / gpt-oss-120b · MXFP4

Start with NVIDIA Spark playbooks and a GB10-compatible runtime. Reserve unified memory for the OS and KV cache.

Model card / deployment guide ↗
DGX Station GB300

1 × Blackwell Ultra + Grace CPU

Memory
252 GB HBM3e + 496 GB CPU
Bandwidth
7.1 TB/s · GPU HBM
Peak compute
15 / 20 PFLOPS · FP4 dense / sparse
NVIDIA specifications ↗

Team development and model evaluation · 1 system

Starting models / precision

Qwen3-32B · BF16 / Llama 3.3 70B · BF16 / FP8

Single-GPU evaluation candidates: 70B BF16 weights use roughly 140 GB before KV cache. CPU memory has different bandwidth from HBM.

Model card / deployment guide ↗
DGX H200

8 × H200 · Hopper

Memory
141 GB / GPU · 1,128 GB total
Bandwidth
4.8 TB/s / GPU
NVIDIA specifications ↗

Pilot a chatbot or coding assistant · 1 node

Starting models / precision

Llama 3.3 70B · BF16 / DeepSeek-R1 671B · FP8

Evaluation candidates: 70B across 2–8 GPUs; start R1 FP8 across eight GPUs with tensor / expert parallelism, then measure KV cache and concurrency.

Model card / deployment guide ↗
DGX B200

8 × B200 · Blackwell

Memory
180 GB / GPU · 1,440 GB total
Bandwidth
8 TB/s / GPU · 64 TB/s total
Peak compute
72 / 144 PFLOPS · FP4 dense / sparse
NVIDIA specifications ↗

Evaluate large models or separate serving replicas · 1 node

Starting models / precision

Llama 3.3 70B · BF16 / DeepSeek-R1 671B · FP8

Evaluate R1 FP8 across eight GPUs or separate 70B replicas. Choose GPUs per replica from the latency target.

Model card / deployment guide ↗
DGX B300

8 × B300 · Blackwell Ultra

Memory
288 GB / GPU · 2,304 GB total
Bandwidth
8 TB/s / GPU · 64 TB/s total
NVIDIA specifications ↗

More memory for model weights and KV cache · 1 node

Starting models / precision

Llama 3.3 70B · BF16 / DeepSeek-R1 671B · FP8 / BF16 (converted)

Start with R1 FP8. Converting to BF16 with a compatible engine uses roughly 1.34 TB of weights across eight GPUs before KV cache.

Model card / deployment guide ↗
GB300 NVL72

72 × Blackwell Ultra · 36 Grace CPUs

Memory
20 TB GPU HBM · rack total
Bandwidth
576 TB/s · rack aggregate
Peak compute
1,080 / 1,440 PFLOPS · FP4 dense / sparse
NVIDIA specifications ↗

Inference pool for multiple models and teams · 1 rack

Starting models / precision

DeepSeek-R1 671B · FP8 / BF16 (converted) · multi-replica

Partition GPUs into serving groups for replicas or disaggregated prefill/decode. Benchmark one group before scaling across the rack.

Model card / deployment guide ↗
DGX Vera Rubin NVL72

72 × Rubin · 36 Vera CPUs

Memory
20.7 TB HBM4 · rack total
Bandwidth
1,400 TB/s · rack aggregate
Peak compute
3,600 PFLOPS · NVFP4 inference
NVIDIA specifications ↗

Plan reasoning and long-context inference · 1 rack

Starting models / precision

DeepSeek-R1 671B · FP8 · evaluation candidate

Preliminary NVIDIA specifications; the referenced table does not specify sparsity for NVFP4 inference. The model is an evaluation candidate: confirm Rubin-compatible runtimes and kernels before deployment.

Model card / deployment guide ↗

Peak compute uses the stated precision and sparsity; it is not measured model tokens/s. Aggregate GPU memory requires parallelism and room for weights, KV cache and runtime. These examples are not Zotect benchmark results.

Specifications reviewed · H200 / B200 / B300 bandwidth ↗ · Spark model support ↗

Scale beyond one system as workloads grow

Add replicas for multiple teams or distribute a model across nodes. These are benchmark starting points; concurrency is a test load to validate with real workloads.

DGX / HGX H200 · B200 · B300

8–256 nodes / 64–2,048 GPUsReplicas & service continuity

Shared inference with replicas and spare capacity for maintenance and releases

If a replica fits one node: example 8 nodes = 4 serving + 2 spare + 2 batch; a two-node replica needs a different reserve plan

H200 · B200 · B300 cluster

8–256 nodes / 64–2,048 GPUsDistributed inference

Evaluate 1–3T-parameter frontier MoE models or disaggregated prefill/decode for long-context agents

This is total cluster capacity across replicas; start each serving-group benchmark at supported FP4/FP8, 32–64K context and 8–32 requests; budget all expert weights, KV cache, communication buffers and spare replicas

GB300 NVL72 → SuperPOD

1–48 racks / 72–3,456 GPUs18 compute trays per rack · 4 GPUs per tray

Large inference pools for multiple models and tenants, with online and batch capacity separated

Plan in racks, not eight-GPU nodes: 1–48 racks = 18–864 four-GPU compute nodes. SuperPOD naming requires the NVIDIA reference architecture

DGX Vera Rubin NVL72

1–48 racks / 72–3,456 GPUs72 Rubin GPUs + 36 Vera CPUs per rack

Plan a Rubin inference pool for reasoning, long-context workloads and multiple tenants

The rack range is a planning example. NVIDIA specifications are preliminary; confirm configuration, software support, power and cooling before final design

How we turn examples into a deployment size

Measure weights + KV cache + runtime overhead, then benchmark representative prompts, output lengths and concurrency. Add replicas for throughput and complete-replica failure capacity. Use multi-node parallelism when memory or computation requires it; tokens/s does not scale linearly with node count

Budget all MoE expert weights, KV cache and runtime. Measure p95 TTFT, p95 inter-token latency and aggregate tokens/s using representative context, output and concurrency before setting capacity and HA.

Connect GPUs, networking and scheduling

Choose orchestration for the workload; every deployment does not need every tool

GPU + NetworkKubernetes / SlurmModel endpoint
Conceptual connection between hardware, orchestration and model endpoints

Kubernetes + NVIDIA Operators

Install NVIDIA GPU Operator for drivers, container runtime, device plugins and telemetry against the support matrix, with NVIDIA Network Operator for network drivers, RDMA device plugins and secondary networks such as Multus / SR-IOV as the topology requires

Documentation ↗

RDMA & distributed serving

Validate GPU–NIC locality, InfiniBand or RoCE fabric, GPUDirect RDMA and inter-node NCCL before inference. Choose tensor, pipeline or expert parallelism for the model. RDMA requires compatible hardware, drivers and fabric, not just a plugin

Documentation ↗

Slurm for scheduled workloads

Use Slurm partitions, GRES, QoS and accounting for batch inference, evaluation and fine-tuning. Use Kubernetes for persistent APIs, or separate resource pools when both schedulers are deployed so they do not allocate the same GPUs

Documentation ↗

Spark / Station are development and local-inference entry points. Cluster deployments must align GPU architecture, Arm64/x86, OS, drivers, NICs and Kubernetes/Slurm versions. Confirm NIM/operator support for each platform before proposing the stack

Shared GPUs. Clear team boundaries.

Administrators allocate through UI/API; users see and consume resources within their team permissions

Team AShared GPU poolTeam B
Conceptual GPU allocation from a shared pool to teams under policy

NVIDIA Base Command Manager

Provision OS images, configure nodes and monitor cluster health centrally, integrating Kubernetes or Slurm management for the selected design

Platform capabilities ↗

NVIDIA Run:ai · GPU / CPU / RAM

Administrators manage departments, projects, node pools, quotas and borrowing policies through UI/API. Users submit within their permissions; teams receive GPU, CPU and CPU-memory budgets with scheduling and resource visibility

Platform capabilities ↗

Storage / disk & tenancy

Combine PVCs / storage classes, requests.storage and ephemeral-storage quotas with Kubernetes ResourceQuota and storage-backend policy. Separate namespaces, RBAC, secrets and network policies per tenant; quotas alone are not security isolation

Platform capabilities ↗

Monitoring & capacity trends

Connect DCGM Exporter, Prometheus and Grafana for GPU utilization and memory, CPU/RAM, storage, fabric errors and per-team quota usage. Retain history for trends alongside TTFT, latency, tokens/s and queue depth

Platform capabilities ↗

Choose a serving stack your team can operate

Combine NVIDIA enterprise and open-source components around compatibility and support needs

NVIDIA AI Enterprise + NIMNVIDIA enterprise

Inference microservices with an enterprise-support path. Match containers and GPUs to the support matrix; verify production entitlements separately from model licenses

Tool documentation ↗
NVIDIA Dynamo + TensorRT-LLMNVIDIA open source

Distributed serving, prefill/decode and KV-aware routing with Dynamo; select TensorRT-LLM where the model and backend are supported

Tool documentation ↗
vLLM / SGLangOpen source

API serving and continuous batching; choose the runtime from model architecture, quantization and measured performance

Tool documentation ↗
KV cache + LMCacheOpen source

KV cache stores attention state; LMCache supports cache reuse, offload and transfer with compatible backends. Validate memory budgets, hit rates and isolation

Tool documentation ↗
LiteLLMOpen source + enterprise options

API gateway for routing, keys, rate limits and token budgets, depending on edition; not a GPU scheduler, and token budgets are not GPU quotas

Tool documentation ↗
DCGM Exporter · Prometheus · GrafanaOpen tooling

Resource and endpoint telemetry, dashboards and alerts, with retention suited to capacity planning

Tool documentation ↗

Start with one serving engine. Add a gateway, KV-cache layer or distributed serving when measurements justify it. Verify NVIDIA AI Enterprise entitlements and commercial features for the delivered versions; open weights do not imply identical open-source licensing

Models to evaluate with your workloads

Selected high-ranking LLM Stats examples, linked to publisher weights. These are candidates, not configurations already certified by Zotect

LLM Stats · Checked · Overall ranks across all models at review date

#11LLM Stats Score · 51.8

Qwen3.8 Max / open checkpoint

Score is for Qwen3.8 Max. The related open checkpoint is not the identical hosted API; re-evaluate quality and features

Qwen/Qwen3.8-2.4T-A95B

Benchmark rankings are not adoption statistics and do not establish Thai-language or domain quality. Recheck quality, licensing, engine support and checkpoint memory; benchmark real context lengths rather than treating model-card maximum context as guaranteed serving capacity

How do open weights compare with frontier AI?

Compare by task to shortlist models for evaluation on your GPUs.

Snapshot checked September 23, 2026. Frontier references are the versions in the source tables, not a live ranking. Open-weight describes access under each model’s license; open-weight models can also be frontier models.

Kimi K3 (max)compared withGPT-5.6 Sol (max)

Reported by Moonshot AI ↗
GPQA Diamond
Kimi K3 (max)93.5
GPT-5.6 Sol (max)94.1
Terminal-Bench 2.1
Kimi K3 (max)88.3
GPT-5.6 Sol (max)88.8

Scores are close on these two benchmarks. Terminal results use different agent harnesses; repeat evaluation in a shared workflow.

GLM-5.3compared withGPT-5.6 Sol

Reported by Z.AI ↗
Terminal-Bench 2.1
GLM-5.388.2
GPT-5.6 Sol88.8
DeepSWE v1.1
GLM-5.366.9
GPT-5.6 Sol72.7

Close on terminal tasks, with a wider gap on DeepSWE. Coding selection depends on the task and agent harness.

DeepSeek-V4-Pro-0813compared withClaude Fable 5 (w/ fallback)

Reported by DeepSeek ↗
Terminal-Bench 2.1
DeepSeek-V4-Pro-081387.9
Claude Fable 5 (w/ fallback)88.0
DeepSWE
DeepSeek-V4-Pro-081362.7
Claude Fable 5 (w/ fallback)70.0

Terminal scores are close; DeepSWE differs by 7.3 points. The source labels the Claude results as including fallback.

Higher is better on these benchmarks. Bars share a 0–100 scale; do not average across benchmarks. Publisher results may use different reasoning budgets, tools and harnesses. Similar scores do not establish statistical equivalence, universal interchangeability, Thai-language quality, or speed and cost on your GPUs.

Already have GPUs? Start with a serving baseline

Share GPU types and node counts, target models, context, concurrency and latency targets to scope deployment and benchmarking

Plan model serving ↗

Extend the core platform

Add integrations and operational services to your serving platform. Baseline monitoring and resource policies are already part of the core scope.

Deployment and Release

Control versions, promotion, and rollback for models and configuration.

  • Model version
  • Promotion
  • Rollback

RAG Production Integration

Connect retrieval sources and applications with defined inspection points.

  • Data sources
  • Retrieval
  • Application

Performance and Cost

Capture throughput, latency, and resource use for operational decisions.

  • Throughput
  • Latency
  • Resource usage

Observability and Governance

Record operating signals, changes, and access.

  • Operating signals
  • Change records
  • Access

Managed AI Platform Operations

Provide proactive platform care after technical-baseline acceptance.

  • Technical baseline
  • Platform care
  • Service agreement

What to track when systems change

Define owners, release records, rollback paths and operating signals for each service.

Release Control

Record versions and approvals for models, prompts, and configuration.

Observability

View application, model, and infrastructure signals in one context.

Security

Control access to data, model endpoints, and operational tools.

Start with the state of your system

Each stage is a separate scope, starting with a shared review of the technical baseline.

Discover

Assess serving, data paths, and operations, then prioritize gaps.

Build

Implement the platform, integrations, observability, and release path.

Support

Provide reactive help with incidents, configuration reviews, and model changes.

Operate

Provide proactive platform care under a service agreement.

Start with your models, data and platform

Review serving, data paths and operations to define the scope together.

Explore AI Infrastructure