This package includes eight NVIDIA DGX Spark systems, a MikroTik CRS804 high-speed compute switch, a MikroTik CRS312 10GbE in-band management switch, and the required compute and management network cables.
The DGX Spark 8-Node AI Cluster combines eight Grace Blackwell systems through a switch-based high-speed network architecture. Each node is connected to the 200GbE compute network via NVIDIA ConnectX-7, while a separate 10GbE Ethernet network is used for system management and standard network traffic.
The compute network supports RoCEv2-based RDMA communication, enabling high-bandwidth data transfer between nodes for NCCL and distributed AI workloads. This keeps AI communication isolated from internet and management traffic.
In OpenZeka’s tests, large language models from the Gemma, Qwen3.6, and GLM-5.2 model families were run on this eight-node cluster. The tested models include Gemma-4-31B-it-NVFP4, Qwen3.6-27B-NVFP4, and GLM-5.2-NVFP4.
NVIDIA DGX Spark 8-Node AI Cluster is a compact and scalable AI cluster that combines eight NVIDIA DGX Spark systems through a high-speed switch-based network architecture.
The package includes eight NVIDIA DGX Spark systems, one MikroTik CRS804 compute switch, four 400G QSFP-DD to 2 × 200G QSFP56 breakout cables, one MikroTik CRS312 in-band management switch, and ten Cat6 10GbE Ethernet cables.
Depending on the customer’s preference, the DGX Spark 8-Node AI Cluster can either be delivered unconfigured or shipped with free pre-installation by OpenZeka. With the unconfigured option, all components are provided as part of the standard package, and the installation is performed by the customer. If free pre-installation is selected, the compute and management networks are configured, connections are checked, and basic cluster communication is verified before shipment.
Each DGX Spark node is connected to the compute switch via a 200GbE connection through NVIDIA ConnectX-7. This compute network can be used for NCCL collective operations, model parallelization, and distributed AI data transfer using RoCEv2-based RDMA communication. Management, SSH, internet access, model downloads, and standard network traffic are handled through a separate 10GbE management network. This configuration keeps the high-speed compute network dedicated to AI communication between nodes.
Scalable AI Infrastructure with Eight Nodes
Each NVIDIA DGX Spark includes 128 GB of coherent unified system memory. A four-node configuration provides a total of:
1 TB of distributed unified memory capacity
32 TB of local NVMe storage
The eight-node configuration is particularly well suited for:
Running large models that exceed the total memory capacity of a single, two, or three Spark systems by distributing them across four nodes using tensor parallelism, pipeline parallelism, or other supported distributed methods
Running models that can operate on fewer Spark systems at higher throughput when appropriate parallelization support is available
Providing greater total capacity for model weights, KV cache, and working memory when using long context lengths
Increasing overall system capacity in inference scenarios where a larger number of concurrent users or requests are served by the same model
Running large language models using distributed inference methods
Multi-node communication based on NCCL and MPI
RAG and agentic AI applications
High-speed node-to-node data communication over RoCEv2/RDMA
RDMA infrastructure designed to reduce network processing overhead on host CPUs during GPU communication
Multimodal generative AI workloads
Model quantization and performance testing
Prototyping distributed architectures before transitioning to a data center environment
Model compatibility and achievable performance vary depending on the model architecture, quantization format, context length, KV cache requirements, framework support, batch size, and parallelization method.
Switch-Based Compute Network
The compute network is built on a MikroTik CRS812 switch.
The CRS812 provides the following high-speed connectivity:
4 × 400G QSFP56-DD
The switch also features RouterOS v7, a quad-core 2 GHz ARM processor, 4 GB of RAM, redundant hot-swappable power supplies, and hot-swappable fans. MikroTik also positions this model for AI clusters and laboratory environments requiring high-speed east-west traffic.
In the DGX Spark 8-Node configuration, all four 400G ports on the CRS804 are used. Each 400G port is split into two 200G QSFP56 connections using a passive breakout cable, providing a 200GbE physical connection for each Spark.
The DGX Spark 8-Node compute network is configured to take advantage of the RoCEv2 and RDMA capabilities of the NVIDIA ConnectX-7 adapters. RoCEv2 carries RDMA communication over Ethernet and IP infrastructure, enabling high-bandwidth, low-overhead data transfer between nodes.
This architecture is particularly suitable for intensive inter-node data transfers during NCCL collective operations, tensor parallelism, pipeline parallelism, and distributed inference.
RoCE performance can be affected by switch queueing, MTU settings, traffic classification, PFC/flow control, ECN, and software configuration. As part of the OpenZeka installation, link speeds, RDMA devices, and inter-node communication are also verified.
RoCE performance can be affected by switch queues, MTU configuration, traffic classification, PFC/flow control, ECN, and software-level settings. In OpenZeka deployments, link speeds, RDMA devices, and inter-node communication are additionally verified.
Dedicated 10GbE Management Network
The in-band management network is provided through the MikroTik CRS312-4C+8XG-RM.
The CRS312 features:
8 × 10G RJ45 Ethernet ports
4 × 10G RJ45/SFP+ combo ports
120 Gbps non-blocking throughput
240 Gbps switching capacity
RouterOS v7 or SwitchOS support
Each Spark system is connected to the CRS312 using a single Cat6 Ethernet cable. This network is used for:
SSH access
System management
Software updates
Model downloads
Internet access
Monitoring and log traffic
Optional NAS access
The separation of compute and management networks prevents standard management and internet traffic from traversing the ConnectX-7 compute links, ensuring that the high-speed network is dedicated to distributed AI communication.
Optional Shared NAS Storage
The cluster can optionally be delivered with an ASUSTOR Lockerstor AS6808T NAS and eight NAS-grade drives. In the standard example configuration, 8 × 8 TB Western Digital WD80EFPX-68C4ZN0 drives are used. The disk brand, model number, and capacity can be customized based on stock availability or project requirements.
Example 8 × 8 TB Configuration:
Configuration
Approximate Raw Capacity
Disk Protection
RAW/JBOD
64 TB
Depends on configuration
RAID 0
64 TB
None
RAID 5
56 TB
1 disk
RAID 6
48 TB
2 disks
The actual usable capacity will be lower than the raw values shown in the table due to the disk manufacturer’s capacity calculation method, RAID metadata, file system overhead, and system-reserved space.
The NAS is connected to the MikroTik CRS312 switch using dual LACP-supported links with two Cat6 Ethernet cables.
Bonding can provide advantages such as:
Link redundancy
Load balancing across multiple clients
Distribution of multiple simultaneous data streams
The disk brand, disk capacity, and RAID configuration can be customized according to project requirements. Depending on stock availability, technically equivalent NAS-grade drives with the same capacity and usage class may be provided instead of the specified disk model.
OpenZeka Cluster Solution
The DGX Spark 8-Node AI Cluster is not simply a hardware package consisting of eight computers placed side by side. By combining a high-speed compute network, a separate management network, breakout connectivity infrastructure, and optional shared storage, it provides a comprehensive infrastructure platform for distributed AI workloads.
Depending on the customer’s preference, the system can either be delivered unconfigured with all components included or shipped with free pre-installation by OpenZeka. As part of the free pre-installation, the compute and management connections are configured, the 200GbE links are checked, basic node access is established, and cluster communication is verified.
Additional services such as customer-specific model installation, custom network policies, enterprise integrations, data transfer, application installation, and performance optimization can be evaluated separately as part of the project.
Explore the detailed specifications, architecture, and configuration details of the NVIDIA DGX Spark Quad AI Cluster, including hardware components, networking capabilities, and optional storage features.
A. Cluster General Features:
Features
Value
Node Count
8
System Platform
NVIDIA DGX Spark or selected compatible GB10 system
Total System Memory
1 TB distributed unified memory
Memory per Node
128 GB LPDDR5x
Total Local Storage
32 TB NVMe in a 4 × 4 TB configuration
AI Performance per Node
Up to 1 PFLOP FP4
Compute Connectivity
200GbE per node
ConnectX-7 Interface Architecture
May appear as multiple logical interfaces depending on operating system and driver configuration
Handles management, internet, SSH, and NAS traffic
Spark Management Cables
10 × Cat6
Compute Connectivity
200GbE per Spark node
Management Connectivity
10GbE per Spark node
RoCEv2/RDMA
High-speed compute communication between Spark nodes
NVIDIA ConnectX-7
RoCEv2 and RDMA-enabled node network adapter
D. Optional NAS Features:
Specification
Example Configuration
NAS
ASUSTOR Lockerstor AS6808T
Drive Bays
8
Standard Drive Configuration
8 × 8 TB
Example Drive Model
Western Digital WD80EFPX-68C4ZN0
RAW Capacity
64 TB
RAID Options
RAID 0, RAID 5, RAID 6, and other levels supported by the device
Network Connectivity
2 × Ethernet, IEEE 802.3ad LACP bonding
Connected Network
MikroTik CRS312 management/storage network
Customization
Disk brand, model number, and capacity can be customized
Performance Tests
The results below are obtained from tests conducted by OpenZeka on different DGX Spark topologies. All measurements are reported with Concurrency = 1 and using Mean statistics.
The purpose of presenting the results for 1, 2, 3, 4, and 8 Spark configurations together is to allow users to compare suitable cluster sizes based on the model they plan to run, the parallelization method, and their expected performance requirements.
TP stands for Tensor Parallelism, and PP stands for Pipeline Parallelism. Empty fields indicate that no test result is currently available for the related model and topology, or that the configuration could not be applied due to the model’s memory requirements and parallelization constraints.
Generation Throughput (TPS)
Higher values are better.
Model
1x DGX Spark
2x DGX Spark Bundle
3x DGX Spark Triple
4x DGX Spark Quad
8x DGX Spark
gemma-4-31B-it-NVFP4
10,81
19,42
N/T
30,53
36,93
Qwen3.6-27B-NVFP4
12,63
22,57
N/T
33,11
33,99
Qwen3.6-35B-A3B-NVFP4
64,52
91,95
N/T
91,16
N/T
MiniMax-M2.7-NVFP4, PP
OOM
17,20, PP-2
18,04, PP-3
N/T
N/T
MiniMax-M2.7-NVFP4, TP
OOM
26,64
TP3-N/A
N/T
N/T
Qwen3.5-397B-A17B-int4
OOM
OOM
17,05
38,14
N/T
glm-5.2-int4
OOM
OOM
OOM
20,16
N/T
glm-5.2-nvfp4
OOM
OOM
OOM
OOM
24,86
*OOM = Out of Memory
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable
Time to First Token (TTFT)
Lower values are better.
Model
1x DGX Spark
2x DGX Spark Bundle
3x DGX Spark Triple
4x DGX Spark Quad
8x DGX Spark
gemma-4-31B-it-NVFP4
201,96 ms
136,35 ms
N/T
279,77 ms
326,54 ms
Qwen3.6-27B-NVFP4
233,30 ms
148,87 ms
N/T
163,59 ms
263,59 ms
Qwen3.6-35B-A3B-NVFP4
500,24 ms
165,31 ms
N/T
230,82 ms
N/T
MiniMax-M2.7-NVFP4, PP
OOM
396,35 ms, PP-2
211,39 ms, PP-3
N/T
N/T
MiniMax-M2.7-NVFP4, TP
OOM
288,98 ms
TP3-N/A
N/T
N/T
Qwen3.5-397B-A17B-int4
OOM
OOM
499,65 ms
324,61 ms
N/T
glm-5.2-int4
OOM
OOM
OOM
571,82 ms
N/T
glm-5.2-nvfp4
OOM
OOM
OOM
OOM
451,23 ms
*OOM = Out of Memory
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable
The results are specific to the tested model, quantization method, software version, and runtime parameters. Factors such as context length, prompt length, number of generated tokens, batch size, KV cache requirements, parallelization method, and inter-node communication overhead may affect the results.
Adding more Spark nodes does not guarantee linear TPS scaling or lower TTFT for every model. While some models can benefit from additional compute capacity, inter-node communication overhead may become the dominant factor for certain workloads.
Cluster size should not be selected solely based on the highest tokens-per-second value. Model memory fit, supported parallelization method, target context length, and the expected number of concurrent users should also be taken into consideration.
These values are not guaranteed minimum performance levels; they are reference results from configurations tested by OpenZeka.
Which Cluster Size?
Configuration
General Usage Approach
1x DGX Spark
Models that fit on a single node, development workloads, and low-concurrency inference
2x DGX Spark Bundle
Two-node tensor/pipeline parallel scenarios and higher total memory capacity
3x DGX Spark Triple
Three-node TP/PP scenarios and a compact architecture without a switch
4x DGX Spark Quad
Four-node TP/PP scenarios, larger models, and centralized compute fabric
8x DGX Spark
Distributed workloads requiring higher total memory capacity and support for eight nodes
The appropriate cluster size varies depending on the model. Before purchase, the target model, quantization method, context length, and intended usage scenario can be shared with OpenZeka for evaluation.
Box Contents
Standard Package
8× NVIDIA DGX Spark or selected compatible GB10 system
8× System power adapters
1× MikroTik CRS804 compute network switch
1× CRS804 power connection kit
4× 1.5 m FS 400G QSFP-DD → 2 × 200G QSFP56 passive copper breakout cables (1 m optional)
The package can be shipped either in standard boxed form without installation or with free pre-installation completed by OpenZeka, depending on the customer’s preference. For pre-installed deliveries, the compute and management connections are configured and basic system checks are performed.
The NVIDIA DGX Spark 8-Node AI Cluster installation consists of several stages, including making the physical connections, preparing the compute switch ports, verifying the 200 Gbps link speeds, separating the management and compute networks, configuring node IP addresses, setting up SSH access, and testing multi-node communication.
The customer can perform the cluster installation themselves or take advantage of OpenZeka’s free pre-installation service.
With the unconfigured delivery option, the DGX Spark systems, switches, and connection cables are shipped as part of the standard package. Physical installation, switch configuration, IP addressing, RoCEv2/RDMA configuration, SSH access, and cluster validation are performed by the customer.
If free pre-installation is selected, OpenZeka will:
Configure the compute and management networks
Perform the basic configuration of the MikroTik CRS812 and CRS312
Verify the 200GbE compute connections
Check ConnectX-7 and RoCEv2/RDMA device visibility
Establish basic connectivity between nodes
Test SSH access and basic cluster communication
Perform NCCL communication validation
The system is shipped ready for the customer to complete the final IP, VLAN, DNS, NTP, internet access, and security configurations within their environment.
OpenZeka’s free pre-installation service includes standard network preparation, node access setup, 200GbE link verification, RoCEv2/RDMA device validation, NCCL communication testing, and basic cluster health checks. Customer-specific software, model, security, and integration work is scoped separately.