NVIDIA DGX Spark Quad AI Cluster

Professional four-node AI infrastructure designed for distributed inference, large language models (LLMs), and high-speed AI communication, featuring a 200GbE RoCEv2 compute network.

NVIDIA DGX Spark Quad AI Cluster

Professional four-node AI infrastructure designed for distributed inference, large language models (LLMs), and high-speed AI communication, featuring a 200GbE RoCEv2 compute network.

The DGX Spark Quad AI Cluster combines four Grace Blackwell systems through a switch-based high-speed networking architecture. Each node is connected to a dedicated 200GbE compute network via NVIDIA ConnectX-7, while a separate 10GbE Ethernet network handles system management and standard network traffic. The compute network supports RoCEv2-based RDMA communication, enabling high-bandwidth, low-latency data transfer between nodes for NCCL and distributed AI workloads. This architecture isolates AI communication from internet and management traffic, ensuring maximum performance and network efficiency.

The DGX Spark Quad AI Cluster combines four Grace Blackwell systems through a switch-based high-speed networking architecture. Each node is connected to a dedicated 200GbE compute network via NVIDIA ConnectX-7, while a separate 10GbE Ethernet network handles system management and standard network traffic. The compute network supports RoCEv2-based RDMA communication, enabling high-bandwidth, low-latency data transfer between nodes for NCCL and distributed AI workloads. This architecture isolates AI communication from internet and management traffic, ensuring maximum performance and network efficiency.

Key Features

  • 4 × NVIDIA GB10 Grace Blackwell systems
  • 128 GB unified system memory per node
  • 512 GB total distributed memory capacity
  • 200GbE NVIDIA ConnectX-7 compute connectivity per node
  • RoCEv2-based RDMA compute networking
  • 10GbE in-band management network per node
  • 16 TB total local NVMe storage in the standard 4 TB-per-node configuration
  • Optimized for tensor parallelism and pipeline parallelism
  • Supports multi-node AI workloads with NCCL, MPI, and RoCEv2/RDMA
  • Optional 64 TB RAW NAS storage
  • Dual-network architecture designed by OpenZeka
  • Available as a plug-and-play solution or with complimentary pre-deployment configuration by OpenZeka
  • Pre-shipment validation of compute, management, and core cluster connectivity

Öne çıkan özellikler

  • 4 × NVIDIA GB10 Grace Blackwell sistemi

  • Node başına 128 GB unified system memory

  • Toplam 512 GB dağıtık bellek kapasitesi

  • Node başına 200GbE ConnectX-7 compute bağlantısı

  • RoCEv2 tabanlı RDMA compute iletişimi

  • Node başına 10GbE in-band management bağlantısı

  • Standart 4 TB konfigürasyonda toplam 16 TB yerel NVMe

  • Tensor parallel ve pipeline parallel çalışma senaryolarına uygun

  • NCCL, MPI ve RoCEv2/RDMA tabanlı çok node’lu AI iş yükleri

  • Opsiyonel 64 TB RAW NAS depolama

  • OpenZeka tarafından tasarlanmış çift ağ mimarisi

  • Kurulumsuz teslimat veya OpenZeka tarafından ücretsiz ön kurulum seçeneği

  • Compute, management ve temel cluster bağlantılarının sevkiyat öncesi doğrulanabilmesi

Öne çıkan özellikler

  • 4 × NVIDIA GB10 Grace Blackwell sistemi

  • Node başına 128 GB unified system memory

  • Toplam 512 GB dağıtık bellek kapasitesi

  • Node başına 200GbE ConnectX-7 compute bağlantısı

  • RoCEv2 tabanlı RDMA compute iletişimi

  • Node başına 10GbE in-band management bağlantısı

  • Standart 4 TB konfigürasyonda toplam 16 TB yerel NVMe

  • Tensor parallel ve pipeline parallel çalışma senaryolarına uygun

  • NCCL, MPI ve RoCEv2/RDMA tabanlı çok node’lu AI iş yükleri

  • Opsiyonel 64 TB RAW NAS depolama

  • OpenZeka tarafından tasarlanmış çift ağ mimarisi

  • Kurulumsuz teslimat veya OpenZeka tarafından ücretsiz ön kurulum seçeneği

  • Compute, management ve temel cluster bağlantılarının sevkiyat öncesi doğrulanabilmesi

Key Features

  • 4 × NVIDIA GB10 Grace Blackwell systems
  • 128 GB unified system memory per node
  • 512 GB total distributed memory capacity
  • 200GbE NVIDIA ConnectX-7 compute connectivity per node
  • RoCEv2-based RDMA compute networking
  • 10GbE in-band management network per node
  • 16 TB total local NVMe storage in the standard 4 TB-per-node configuration
  • Optimized for tensor parallelism and pipeline parallelism
  • Supports multi-node AI workloads with NCCL, MPI, and RoCEv2/RDMA
  • Optional 64 TB RAW NAS storage
  • Dual-network architecture designed by OpenZeka
  • Available as a plug-and-play solution or with complimentary pre-deployment configuration by OpenZeka
  • Pre-shipment validation of compute, management, and core cluster connectivity

What’s Included in the 4× NVIDIA DGX Spark Switch Package

The package can be shipped either as a standard boxed, unconfigured solution or with complimentary pre-installation completed by OpenZeka, depending on customer requirements. For pre-configured deliveries, compute and management connections are prepared, and basic system validation checks are performed.

Standard Package

  • 4 × NVIDIA DGX Spark systems or selected compatible GB10 systems
  • 4 × System power adapters
  • 1 × MikroTik CRS812 compute network switch
  • 1 × CRS812 power connection kit
  • 2 × 1.5m FS 400G QSFP-DD → 2 × 200G QSFP56 passive copper breakout cables (1m option available)
  • 1 × MikroTik CRS312 10GbE in-band management switch
  • 1 × CRS312 power cable
  • 6 × Cat6 10GbE Ethernet cables

Optional NAS Package

  • 1 × ASUSTOR Lockerstor AS6808T NAS
  • 8 × 8 TB NAS HDDs
  • 2 × Cat6 Ethernet cables
  • Selected RAID configuration

What’s Included in the 4× NVIDIA DGX Spark Switch Package

The package can be shipped either as a standard boxed, unconfigured solution or with complimentary pre-installation completed by OpenZeka, depending on customer requirements. For pre-configured deliveries, compute and management connections are prepared, and basic system validation checks are performed.

Standard Package

  • 4 × NVIDIA DGX Spark systems or selected compatible GB10 systems
  • 4 × System power adapters
  • 1 × MikroTik CRS812 compute network switch
  • 1 × CRS812 power connection kit
  • 2 × 1.5m FS 400G QSFP-DD → 2 × 200G QSFP56 passive copper breakout cables (1m option available)
  • 1 × MikroTik CRS312 10GbE in-band management switch
  • 1 × CRS312 power cable
  • 6 × Cat6 10GbE Ethernet cables

Optional NAS Package

  • 1 × ASUSTOR Lockerstor AS6808T NAS
  • 8 × 8 TB NAS HDDs
  • 2 × Cat6 Ethernet cables
  • Selected RAID configuration

4× NVIDIA DGX Spark Switch Connection Diagram

4× NVIDIA DGX Spark Switch Connection Diagram

Features

Installation

The installation process of the DGX Spark Quad AI Cluster consists of the following stages: establishing physical connections, preparing compute switch ports, validating 200 Gbps link speeds, separating management and compute networks, configuring node IP addresses, setting up SSH access, and testing multi-node communication.

Standard Package

In the standard unconfigured delivery option, DGX Spark systems, switches, and connection cables are shipped with the standard package contents. Physical installation, switch configuration, IP addressing, RoCEv2/RDMA setup, SSH access configuration, and cluster validation are performed by the customer.

Complimentary Pre-Installation

The customer can perform the cluster installation independently or take advantage of OpenZeka’s complimentary pre-installation service.

When complimentary pre-installation is selected, OpenZeka performs the following:

  • Compute and management networks are configured
  • Basic configurations of the MikroTik CRS812 and CRS312 switches are completed
  • 200GbE compute links are validated
  • ConnectX-7 and RoCEv2/RDMA device visibility is verified
  • Basic node-to-node connectivity is prepared
  • SSH access and fundamental cluster communication are tested
  • NCCL communication validation is performed

The system is delivered ready for completion of the customer’s final environment settings, including IP configuration, VLAN, DNS, NTP, internet access, and security policies.

OpenZeka’s complimentary pre-installation service includes standard network preparation, node access configuration, 200GbE connectivity validation, RoCEv2/RDMA device verification, NCCL communication testing, and basic cluster health checks. Customer-specific software, model deployment, security configurations, and integration activities are scoped separately.

Installation

The installation process of the DGX Spark Quad AI Cluster consists of the following stages: establishing physical connections, preparing compute switch ports, validating 200 Gbps link speeds, separating management and compute networks, configuring node IP addresses, setting up SSH access, and testing multi-node communication.

Standard Package

In the standard unconfigured delivery option, DGX Spark systems, switches, and connection cables are shipped with the standard package contents. Physical installation, switch configuration, IP addressing, RoCEv2/RDMA setup, SSH access configuration, and cluster validation are performed by the customer.

Complimentary Pre-Installation

The customer can perform the cluster installation independently or take advantage of OpenZeka’s complimentary pre-installation service.

When complimentary pre-installation is selected, OpenZeka performs the following:

  • Compute and management networks are configured
  • Basic configurations of the MikroTik CRS812 and CRS312 switches are completed
  • 200GbE compute links are validated
  • ConnectX-7 and RoCEv2/RDMA device visibility is verified
  • Basic node-to-node connectivity is prepared
  • SSH access and fundamental cluster communication are tested
  • NCCL communication validation is performed

The system is delivered ready for completion of the customer’s final environment settings, including IP configuration, VLAN, DNS, NTP, internet access, and security policies.

OpenZeka’s complimentary pre-installation service includes standard network preparation, node access configuration, 200GbE connectivity validation, RoCEv2/RDMA device verification, NCCL communication testing, and basic cluster health checks. Customer-specific software, model deployment, security configurations, and integration activities are scoped separately.

Which Cluster Size?

1x NVIDIA DGX Spark

Models that fit on a single node, development workloads, and low-concurrency inference

2x NVIDIA DGX Spark Bundle

Two-node tensor/pipeline parallel scenarios and higher total memory capacity

3x NVIDIA DGX Spark Triple

Three-node TP/PP scenarios and a compact configuration without requiring a switch

4x NVIDIA DGX Spark Quad

Four-node TP/PP scenarios, larger models, and centralized compute fabric

8x NVIDIA DGX Spark 

Higher total memory requirements and distributed workloads supporting up to eight nodes

Which Cluster Size?

1x DGX Spark

Models that fit on a single node, development workloads, and low-concurrency inference

2x DGX Spark Bundle

Two-node tensor/pipeline parallel scenarios and higher total memory capacity

3x DGX Spark Triple

Three-node TP/PP scenarios and a compact configuration without requiring a switch

 

4x DGX Spark Quad

Four-node TP/PP scenarios, larger models, and centralized compute fabric

8x DGX Spark Quad

Higher total memory requirements and distributed workloads supporting up to eight nodes

Performance Tests

The results below are obtained from tests conducted by OpenZeka across different NVIDIA DGX Spark topologies. All measurements are reported with Concurrency = 1 using Mean statistics.

The purpose of presenting the results for 1, 2, 3, 4, and 8 NVIDIA DGX Spark configurations together is to allow users to compare suitable cluster sizes based on the models they plan to run, their parallelization strategy, and their expected performance requirements.

TP refers to Tensor Parallelism, and PP refers to Pipeline Parallelism. Empty fields indicate that no measurement is available for the respective model and topology, or that the configuration could not be applied to that model.

Tokens Per Second, TPS

Higher values indicate better performance.

Model1× DGX Spark2× DGX Spark Bundle3× DGX Spark Triple4× DGX Spark Quad8× DGX Spark
gemma-4-31B-it-NVFP410.8119.42N/T30.5336.93
Qwen3.6-27B-NVFP412.6322.57N/T33.1133.99
Qwen3.6-35B-A3B-NVFP464.5291.95N/T91.16N/T
MiniMax-M2.7-NVFP4, PPOOM17.20, PP-218.04, PP-3N/TN/T
MiniMax-M2.7-NVFP4, TPOOM26.64TP3-N/AN/TN/T
Qwen3.5-397B-A17B-int4OOMOOM17.0538.14N/T
glm-5.2-int4OOMOOMOOM20.16N/T
glm-5.2-nvfp4OOMOOMOOMOOM24.86

*OOM = Out of Memory (Insufficient memory)
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable

Time to First Token, TTFT

Lower values indicate better performance.

Model1× DGX Spark2× DGX Spark Bundle3× DGX Spark Triple4× DGX Spark Quad8× DGX Spark
gemma-4-31B-it-NVFP4201.96 ms136.35 msN/T279.77 ms326.54 ms
Qwen3.6-27B-NVFP4233.30 ms148.87 msN/T163.59 ms263.59 ms
Qwen3.6-35B-A3B-NVFP4500.24 ms165.31 msN/T230.82 msN/T
MiniMax-M2.7-NVFP4, PPOOM396.35 ms, PP-2211.39 ms, PP-3N/TN/T
MiniMax-M2.7-NVFP4, TPOOM288.98 msTP3-N/AN/TN/T
Qwen3.5-397B-A17B-int4OOMOOM499.65 ms324.61 msN/T
glm-5.2-int4OOMOOMOOM571.82 msN/T
glm-5.2-nvfp4OOMOOMOOMOOM451.23 ms

*OOM = Out of Memory (Insufficient memory)
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable

The results are specific to the tested model, quantization method, software version, and runtime parameters. Factors such as context length, prompt length, generated token count, batch size, KV cache requirements, parallelization strategy, and network communication overhead can affect performance results.

Using more Spark nodes does not guarantee linear TPS scaling or lower TTFT for every model. While some models can effectively utilize additional compute capacity, communication overhead between nodes may become the dominant factor for certain workloads.

Cluster size should not be selected solely based on the highest tokens-per-second value. Model memory requirements, supported parallelization methods, target context length, and the expected number of concurrent users should also be considered.

These values are not guaranteed minimum performance figures; they represent reference results from configurations tested by OpenZeka.

PRODUCT GUIDES

4-Node DGX Spark Cluster Installation Guide

NVIDIA DGX Spark Product Documentation

NVIDIA Multi-Node DGX Spark Switch Installation Guide

MikroTik CRS812 Product Documentation

MikroTik CRS312 Product Documentation

RouterOS Documentation

ASUSTOR AS6808T User Guide

FS Breakout Cable Product Datasheet

Western Digital Datasheet

PRODUCT GUIDES

NVIDIA DGX Spark Product Documentation

NVIDIA Multi-Node DGX Spark Switch Installation Guide

4-Node DGX Spark Cluster Installation Guide

MikroTik CRS812 Product Documentation

MikroTik CRS312 Product Documentation

RouterOS Documentation

ASUSTOR AS6808T User Guide

FS Breakout Cable Product Datasheet

Western Digital Datasheet

Software and Documentation

For additional resources such as documentation and software downloads, please contact us.