The DGX Spark Quad AI Cluster combines four Grace Blackwell systems through a switch-based high-speed networking architecture. Each node is connected to a dedicated 200GbE compute network via NVIDIA ConnectX-7, while a separate 10GbE Ethernet network handles system management and standard network traffic. The compute network supports RoCEv2-based RDMA communication, enabling high-bandwidth, low-latency data transfer between nodes for NCCL and distributed AI workloads. This architecture isolates AI communication from internet and management traffic, ensuring maximum performance and network efficiency.
The DGX Spark Quad AI Cluster combines four Grace Blackwell systems through a switch-based high-speed networking architecture. Each node is connected to a dedicated 200GbE compute network via NVIDIA ConnectX-7, while a separate 10GbE Ethernet network handles system management and standard network traffic. The compute network supports RoCEv2-based RDMA communication, enabling high-bandwidth, low-latency data transfer between nodes for NCCL and distributed AI workloads. This architecture isolates AI communication from internet and management traffic, ensuring maximum performance and network efficiency.




Key Features
- 4 × NVIDIA GB10 Grace Blackwell systems
- 128 GB unified system memory per node
- 512 GB total distributed memory capacity
- 200GbE NVIDIA ConnectX-7 compute connectivity per node
- RoCEv2-based RDMA compute networking
- 10GbE in-band management network per node
- 16 TB total local NVMe storage in the standard 4 TB-per-node configuration
- Optimized for tensor parallelism and pipeline parallelism
- Supports multi-node AI workloads with NCCL, MPI, and RoCEv2/RDMA
- Optional 64 TB RAW NAS storage
- Dual-network architecture designed by OpenZeka
- Available as a plug-and-play solution or with complimentary pre-deployment configuration by OpenZeka
- Pre-shipment validation of compute, management, and core cluster connectivity
Öne çıkan özellikler
-
4 × NVIDIA GB10 Grace Blackwell sistemi
-
Node başına 128 GB unified system memory
-
Toplam 512 GB dağıtık bellek kapasitesi
-
Node başına 200GbE ConnectX-7 compute bağlantısı
-
RoCEv2 tabanlı RDMA compute iletişimi
-
Node başına 10GbE in-band management bağlantısı
-
Standart 4 TB konfigürasyonda toplam 16 TB yerel NVMe
-
Tensor parallel ve pipeline parallel çalışma senaryolarına uygun
-
NCCL, MPI ve RoCEv2/RDMA tabanlı çok node’lu AI iş yükleri
-
Opsiyonel 64 TB RAW NAS depolama
-
OpenZeka tarafından tasarlanmış çift ağ mimarisi
-
Kurulumsuz teslimat veya OpenZeka tarafından ücretsiz ön kurulum seçeneği
-
Compute, management ve temel cluster bağlantılarının sevkiyat öncesi doğrulanabilmesi

Öne çıkan özellikler
-
4 × NVIDIA GB10 Grace Blackwell sistemi
-
Node başına 128 GB unified system memory
-
Toplam 512 GB dağıtık bellek kapasitesi
-
Node başına 200GbE ConnectX-7 compute bağlantısı
-
RoCEv2 tabanlı RDMA compute iletişimi
-
Node başına 10GbE in-band management bağlantısı
-
Standart 4 TB konfigürasyonda toplam 16 TB yerel NVMe
-
Tensor parallel ve pipeline parallel çalışma senaryolarına uygun
-
NCCL, MPI ve RoCEv2/RDMA tabanlı çok node’lu AI iş yükleri
-
Opsiyonel 64 TB RAW NAS depolama
-
OpenZeka tarafından tasarlanmış çift ağ mimarisi
-
Kurulumsuz teslimat veya OpenZeka tarafından ücretsiz ön kurulum seçeneği
-
Compute, management ve temel cluster bağlantılarının sevkiyat öncesi doğrulanabilmesi
Key Features
- 4 × NVIDIA GB10 Grace Blackwell systems
- 128 GB unified system memory per node
- 512 GB total distributed memory capacity
- 200GbE NVIDIA ConnectX-7 compute connectivity per node
- RoCEv2-based RDMA compute networking
- 10GbE in-band management network per node
- 16 TB total local NVMe storage in the standard 4 TB-per-node configuration
- Optimized for tensor parallelism and pipeline parallelism
- Supports multi-node AI workloads with NCCL, MPI, and RoCEv2/RDMA
- Optional 64 TB RAW NAS storage
- Dual-network architecture designed by OpenZeka
- Available as a plug-and-play solution or with complimentary pre-deployment configuration by OpenZeka
- Pre-shipment validation of compute, management, and core cluster connectivity
What’s Included in the 4× NVIDIA DGX Spark Switch Package
The package can be shipped either as a standard boxed, unconfigured solution or with complimentary pre-installation completed by OpenZeka, depending on customer requirements. For pre-configured deliveries, compute and management connections are prepared, and basic system validation checks are performed.
What’s Included in the 4× NVIDIA DGX Spark Switch Package
The package can be shipped either as a standard boxed, unconfigured solution or with complimentary pre-installation completed by OpenZeka, depending on customer requirements. For pre-configured deliveries, compute and management connections are prepared, and basic system validation checks are performed.
4× NVIDIA DGX Spark Switch Connection Diagram
4× NVIDIA DGX Spark Switch Connection Diagram

Features
| Feature | Value |
|---|---|
| Node Count | 4 |
| System Platform | NVIDIA DGX Spark or selected compatible GB10 system |
| Total System Memory | 512 GB distributed unified memory |
| Memory per Node | 128 GB LPDDR5x |
| Total Local Storage | 16 TB NVMe in a 4 × 4 TB configuration |
| AI Performance per Node | Up to 1 PFLOP FP4 |
| Compute Connectivity | 200GbE per node |
| ConnectX-7 Interface Structure | May appear as multiple logical interfaces depending on operating system and driver configuration |
| In-Band Management | 10GbE per node |
| Compute Topology | Ethernet switch-based star topology |
| Optional Shared Storage | 64 TB RAW NAS |
| Parallel Execution | Tensor parallelism, pipeline parallelism, distributed inference |
| Operating System | NVIDIA DGX OS or the supported NVIDIA software environment of the selected system |
| Compute Protocol | RDMA-enabled Ethernet over RoCEv2 |
| RDMA Adapter | NVIDIA ConnectX-7 |
| Compute Usage | NCCL, model parallelism, and distributed inference communication |
| Network Separation | Dedicated 200GbE compute and 10GbE management networks |
| Feature | Per Node |
|---|---|
| Architecture | NVIDIA Grace Blackwell |
| Superchip | NVIDIA GB10 |
| GPU | NVIDIA Blackwell |
| CPU | 20-core Arm CPU, 10 × Cortex-X925 + 10 × Cortex-A725 |
| Tensor Cores | 5th Generation |
| RT Cores | 4th Generation |
| AI Performance | Up to 1 PFLOP FP4 |
| System Memory | 128 GB LPDDR5x coherent unified memory |
| Memory Bandwidth | Up to 273 GB/s |
| Storage | 4 TB NVMe M.2 |
| Standard Ethernet | 10GbE RJ45 |
| High-Speed NIC | ConnectX-7 |
| Wireless | Wi-Fi 7 |
| Bluetooth | Bluetooth 5.4 |
| USB | 4 × USB Type-C |
| Display Output | HDMI 2.1a |
| Operating System | NVIDIA DGX OS |
| Dimensions | 150 × 150 × 50.5 mm |
| Weight | Approximately 1.2 kg |
| Component | Function |
|---|---|
| MikroTik CRS812 | High-speed Ethernet switch carrying RoCEv2 compute traffic |
| FS QDD-400G-2QPC015 | 400G QSFP-DD → 2 × 200G QSFP56 breakout cable |
| Cable Quantity | 2 |
| Cable Length | 1.5 m (1 m optional) |
| MikroTik CRS312 | Management, internet, SSH, and NAS traffic |
| Spark Management Cables | 6 × Cat6 |
| Compute Connectivity | 200GbE per Spark node |
| Management Connectivity | 10GbE per Spark node |
| RoCEv2/RDMA | High-speed compute communication between Spark nodes |
| NVIDIA ConnectX-7 | Node network adapter supporting RoCEv2 and RDMA |
| Feature | Example Configuration |
|---|---|
| NAS | ASUSTOR Lockerstor AS6808T |
| Drive Bays | 8 |
| Standard Drive Configuration | 8 × 8 TB |
| Example Drive Model | Western Digital WD80EFPX-68C4ZN0 |
| RAW Capacity | 64 TB |
| RAID Options | RAID 0, RAID 5, RAID 6, and other levels supported by the device |
| Network Connectivity | 2 × Ethernet, IEEE 802.3ad LACP bonding |
| Connected Network | MikroTik CRS312 management/storage network |
| Customization | Drive brand, model number, and capacity can be customized |
Installation
The installation process of the DGX Spark Quad AI Cluster consists of the following stages: establishing physical connections, preparing compute switch ports, validating 200 Gbps link speeds, separating management and compute networks, configuring node IP addresses, setting up SSH access, and testing multi-node communication.
When complimentary pre-installation is selected, OpenZeka performs the following:
- Compute and management networks are configured
- Basic configurations of the MikroTik CRS812 and CRS312 switches are completed
- 200GbE compute links are validated
- ConnectX-7 and RoCEv2/RDMA device visibility is verified
- Basic node-to-node connectivity is prepared
- SSH access and fundamental cluster communication are tested
- NCCL communication validation is performed
Installation
The installation process of the DGX Spark Quad AI Cluster consists of the following stages: establishing physical connections, preparing compute switch ports, validating 200 Gbps link speeds, separating management and compute networks, configuring node IP addresses, setting up SSH access, and testing multi-node communication.
When complimentary pre-installation is selected, OpenZeka performs the following:
- Compute and management networks are configured
- Basic configurations of the MikroTik CRS812 and CRS312 switches are completed
- 200GbE compute links are validated
- ConnectX-7 and RoCEv2/RDMA device visibility is verified
- Basic node-to-node connectivity is prepared
- SSH access and fundamental cluster communication are tested
- NCCL communication validation is performed
Which Cluster Size?
Which Cluster Size?
Performance Tests
The results below are obtained from tests conducted by OpenZeka across different NVIDIA DGX Spark topologies. All measurements are reported with Concurrency = 1 using Mean statistics.
The purpose of presenting the results for 1, 2, 3, 4, and 8 NVIDIA DGX Spark configurations together is to allow users to compare suitable cluster sizes based on the models they plan to run, their parallelization strategy, and their expected performance requirements.
TP refers to Tensor Parallelism, and PP refers to Pipeline Parallelism. Empty fields indicate that no measurement is available for the respective model and topology, or that the configuration could not be applied to that model.
Tokens Per Second, TPS
Higher values indicate better performance.
| Model | 1× DGX Spark | 2× DGX Spark Bundle | 3× DGX Spark Triple | 4× DGX Spark Quad | 8× DGX Spark |
|---|---|---|---|---|---|
| gemma-4-31B-it-NVFP4 | 10.81 | 19.42 | N/T | 30.53 | 36.93 |
| Qwen3.6-27B-NVFP4 | 12.63 | 22.57 | N/T | 33.11 | 33.99 |
| Qwen3.6-35B-A3B-NVFP4 | 64.52 | 91.95 | N/T | 91.16 | N/T |
| MiniMax-M2.7-NVFP4, PP | OOM | 17.20, PP-2 | 18.04, PP-3 | N/T | N/T |
| MiniMax-M2.7-NVFP4, TP | OOM | 26.64 | TP3-N/A | N/T | N/T |
| Qwen3.5-397B-A17B-int4 | OOM | OOM | 17.05 | 38.14 | N/T |
| glm-5.2-int4 | OOM | OOM | OOM | 20.16 | N/T |
| glm-5.2-nvfp4 | OOM | OOM | OOM | OOM | 24.86 |
*OOM = Out of Memory (Insufficient memory)
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable
Time to First Token, TTFT
Lower values indicate better performance.
| Model | 1× DGX Spark | 2× DGX Spark Bundle | 3× DGX Spark Triple | 4× DGX Spark Quad | 8× DGX Spark |
|---|---|---|---|---|---|
| gemma-4-31B-it-NVFP4 | 201.96 ms | 136.35 ms | N/T | 279.77 ms | 326.54 ms |
| Qwen3.6-27B-NVFP4 | 233.30 ms | 148.87 ms | N/T | 163.59 ms | 263.59 ms |
| Qwen3.6-35B-A3B-NVFP4 | 500.24 ms | 165.31 ms | N/T | 230.82 ms | N/T |
| MiniMax-M2.7-NVFP4, PP | OOM | 396.35 ms, PP-2 | 211.39 ms, PP-3 | N/T | N/T |
| MiniMax-M2.7-NVFP4, TP | OOM | 288.98 ms | TP3-N/A | N/T | N/T |
| Qwen3.5-397B-A17B-int4 | OOM | OOM | 499.65 ms | 324.61 ms | N/T |
| glm-5.2-int4 | OOM | OOM | OOM | 571.82 ms | N/T |
| glm-5.2-nvfp4 | OOM | OOM | OOM | OOM | 451.23 ms |
*OOM = Out of Memory (Insufficient memory)
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable
The results are specific to the tested model, quantization method, software version, and runtime parameters. Factors such as context length, prompt length, generated token count, batch size, KV cache requirements, parallelization strategy, and network communication overhead can affect performance results.
Using more Spark nodes does not guarantee linear TPS scaling or lower TTFT for every model. While some models can effectively utilize additional compute capacity, communication overhead between nodes may become the dominant factor for certain workloads.
Cluster size should not be selected solely based on the highest tokens-per-second value. Model memory requirements, supported parallelization methods, target context length, and the expected number of concurrent users should also be considered.
These values are not guaranteed minimum performance figures; they represent reference results from configurations tested by OpenZeka.











