This package includes four NVIDIA DGX Spark systems, one MikroTik CRS812 high-speed compute switch, one MikroTik CRS312 10GbE in-band management switch, and the required compute and management connection cables.
The DGX Spark Quad AI Cluster combines four Grace Blackwell systems into a switch-based high-speed network architecture. Each node connects to a 200GbE compute network via NVIDIA ConnectX-7, while a separate 10GbE Ethernet network is used for system management and standard network traffic.
The compute network supports RoCEv2-based RDMA communication, enabling high-bandwidth data transfer between nodes for NCCL and distributed AI workloads. This architecture isolates AI communication from internet and management traffic.
In OpenZeka benchmarks, large language models from the GLM-5.2, Qwen3.5, Qwen3.6, and Gemma model families were deployed on this four-node cluster. The evaluated models included GLM-5.2-INT4, Qwen3.5-397B-A17B-INT4, Gemma-4-31B-it-NVFP4, Qwen3.6-27B-NVFP4, and Qwen3.6-35B-A3B-NVFP4.
The NVIDIA DGX Spark Quad AI Cluster is a compact and scalable AI cluster that integrates four NVIDIA DGX Spark systems through a high-speed, switch-based network architecture.
The package includes four NVIDIA DGX Spark systems, one MikroTik CRS812 compute switch, two 400G QSFP-DD to 2 × 200G QSFP56 breakout cables, one MikroTik CRS312 in-band management switch, and six Cat6 10GbE Ethernet cables.
The DGX Spark Quad AI Cluster can be delivered either without installation or with complimentary pre-installation by OpenZeka, depending on the customer’s preference. With the uninstalled option, all components are shipped in their standard packaging, and installation is performed by the customer. If complimentary pre-installation is selected, the compute and management networks are configured, all connections are verified, and basic cluster communication is validated prior to shipment.
Each DGX Spark node connects to the compute switch via a 200GbE link using the NVIDIA ConnectX-7 network interface. The compute network supports RoCEv2-based RDMA communication, enabling NCCL collective operations, model parallelism, and high-speed data transfer for distributed AI workloads. Management, SSH access, internet connectivity, model downloads, and standard network traffic are handled over a separate 10GbE management network. This architecture dedicates the high-speed compute network exclusively to inter-node AI communication.
Scalable AI Infrastructure with Four Nodes
Each NVIDIA DGX Spark system includes 128 GB of coherent unified system memory. The four-node configuration provides a total of:
512 GB of distributed unified memory capacity
16 TB of local NVMe storage
The four-node architecture provides a suitable development environment for:
Running large models that cannot fit into the total memory capacity of one, two, or three Spark systems by distributing them across four nodes using tensor parallelism, pipeline parallelism, or other supported distributed methods
Running models that can fit on fewer Spark systems with higher throughput when appropriate parallelization support is available
Providing greater total capacity for model weights, KV cache, and runtime memory in long context length scenarios
Increasing overall system capacity in inference scenarios where more concurrent users or requests are served by the same model
Running large language models through distributed inference methods
Multi-node communication based on NCCL and MPI
RAG and agentic AI applications
High-speed node-to-node data communication over RoCEv2/RDMA
RDMA-based infrastructure designed to reduce network processing overhead on host CPUs during GPU communication
Multimodal generative AI workloads
Model quantization and performance testing
Prototyping distributed architectures before transitioning to data center environments
Model compatibility and achievable performance vary depending on the model architecture, quantization format, context length, KV cache requirements, framework support, batch size, and parallelization method.
Switch-Based Compute Network
The compute network is built on the MikroTik CRS812 switch.
The CRS812 provides the following high-speed connectivity ports:
2 × 400G QSFP56-DD ports
2 × 200G QSFP56 ports
8 × 50G SFP56 ports
The switch also features RouterOS v7, a quad-core 2 GHz ARM processor, 4 GB of RAM, redundant hot-swappable power supplies, and hot-swappable fans. MikroTik positions this model for AI cluster environments and laboratory deployments that require high-speed east-west traffic.
In the DGX Spark Quad configuration, two 400G ports of the CRS812 are used. Each 400G port is split into two 200G QSFP56 connections using a passive breakout cable. This provides a dedicated 200GbE physical connection for each Spark system.
The DGX Spark Quad compute network is configured to leverage the RoCEv2 and RDMA capabilities of NVIDIA ConnectX-7 adapters. RoCEv2 enables RDMA communication over Ethernet and IP infrastructure, providing high-bandwidth data transfer between nodes with low communication overhead.
This architecture is particularly used for intensive data transfers between nodes during NCCL collective operations, tensor parallelism, pipeline parallelism, and distributed inference workloads.
RoCE performance can be affected by switch queues, MTU configuration, traffic classification, PFC/flow control, ECN, and software-level settings. In OpenZeka deployments, link speeds, RDMA devices, and inter-node communication are additionally verified.
Dedicated 10GbE Management Network
The in-band management network is provided through the MikroTik CRS312-4C+8XG-RM.
The CRS312 features:
8 × 10G RJ45 Ethernet ports
4 × 10G RJ45/SFP+ combo ports
120 Gbps non-blocking throughput
240 Gbps switching capacity
RouterOS v7 or SwitchOS support
Each Spark system is connected to the CRS312 using a single Cat6 Ethernet cable. This network is used for:
SSH access
System management
Software updates
Model downloads
Internet access
Monitoring and log traffic
Optional NAS access
The separation of compute and management networks prevents standard management and internet traffic from traversing the ConnectX-7 compute links, ensuring that the high-speed network is dedicated to distributed AI communication.
Optional Shared NAS Storage
The cluster can optionally be delivered with an ASUSTOR Lockerstor AS6808T NAS and eight NAS-grade drives. In the standard example configuration, 8 × 8 TB Western Digital WD80EFPX-68C4ZN0 drives are used. The disk brand, model number, and capacity can be customized based on stock availability or project requirements.
Example 8 × 8 TB Configuration:
Configuration
Approximate Raw Capacity
Disk Protection
RAW/JBOD
64 TB
Depends on configuration
RAID 0
64 TB
None
RAID 5
56 TB
1 disk
RAID 6
48 TB
2 disks
The actual usable capacity will be lower than the raw values shown in the table due to the disk manufacturer’s capacity calculation method, RAID metadata, file system overhead, and system-reserved space.
The NAS is connected to the MikroTik CRS312 switch using dual LACP-supported links with two Cat6 Ethernet cables.
Bonding can provide advantages such as:
Link redundancy
Load balancing across multiple clients
Distribution of multiple simultaneous data streams
The disk brand, disk capacity, and RAID configuration can be customized according to project requirements. Depending on stock availability, technically equivalent NAS-grade drives with the same capacity and usage class may be provided instead of the specified disk model.
OpenZeka Cluster Solution
The DGX Spark Quad AI Cluster is not simply a hardware package consisting of four computers placed together. By integrating a high-speed compute network, a dedicated management network, breakout connectivity infrastructure, and optional shared storage, it provides a comprehensive infrastructure designed for distributed AI workloads.
The system can be delivered either as a complete uninstalled package with all components included or shipped with complimentary pre-installation performed by OpenZeka, depending on customer preference. As part of the complimentary pre-installation service, compute and management connections are configured, 200GbE links are verified, basic node access is established, and cluster communication is validated.
Additional services such as customer-specific model deployment, custom network policies, enterprise integrations, data migration, application installation, and performance optimization are evaluated separately within the scope of each project.
Explore the detailed specifications, architecture, and configuration details of the NVIDIA DGX Spark Quad AI Cluster, including hardware components, networking capabilities, and optional storage features.
A. Cluster General Features:
Features
Value
Node Count
4
System Platform
NVIDIA DGX Spark or selected compatible GB10 system
Total System Memory
512 GB distributed unified memory
Memory per Node
128 GB LPDDR5x
Total Local Storage
16 TB NVMe in a 4 × 4 TB configuration
AI Performance per Node
Up to 1 PFLOP FP4
Compute Connectivity
200GbE per node
ConnectX-7 Interface Architecture
May appear as multiple logical interfaces depending on operating system and driver configuration
Handles management, internet, SSH, and NAS traffic
Spark Management Cables
6 × Cat6
Compute Connectivity
200GbE per Spark node
Management Connectivity
10GbE per Spark node
RoCEv2/RDMA
High-speed compute communication between Spark nodes
NVIDIA ConnectX-7
RoCEv2 and RDMA-enabled node network adapter
D. Optional NAS Features:
Specification
Example Configuration
NAS
ASUSTOR Lockerstor AS6808T
Drive Bays
8
Standard Drive Configuration
8 × 8 TB
Example Drive Model
Western Digital WD80EFPX-68C4ZN0
RAW Capacity
64 TB
RAID Options
RAID 0, RAID 5, RAID 6, and other levels supported by the device
Network Connectivity
2 × Ethernet, IEEE 802.3ad LACP bonding
Connected Network
MikroTik CRS312 management/storage network
Customization
Disk brand, model number, and capacity can be customized
Performance Tests
The results below are obtained from tests conducted by OpenZeka on different DGX Spark topologies. All measurements are reported with Concurrency = 1 and using Mean statistics.
The purpose of presenting the results for 1, 2, 3, 4, and 8 Spark configurations together is to allow users to compare suitable cluster sizes based on the model they plan to run, the parallelization method, and their expected performance requirements.
TP stands for Tensor Parallelism, and PP stands for Pipeline Parallelism. Empty fields indicate that no test result is currently available for the related model and topology, or that the configuration could not be applied due to the model’s memory requirements and parallelization constraints.
Generation Throughput (TPS)
Higher values are better.
Model
1x DGX Spark
2x DGX Spark Bundle
3x DGX Spark Triple
4x DGX Spark Quad
8x DGX Spark
gemma-4-31B-it-NVFP4
10,81
19,42
N/T
30,53
36,93
Qwen3.6-27B-NVFP4
12,63
22,57
N/T
33,11
33,99
Qwen3.6-35B-A3B-NVFP4
64,52
91,95
N/T
91,16
N/T
MiniMax-M2.7-NVFP4, PP
OOM
17,20, PP-2
18,04, PP-3
N/T
N/T
MiniMax-M2.7-NVFP4, TP
OOM
26,64
TP3-N/A
N/T
N/T
Qwen3.5-397B-A17B-int4
OOM
OOM
17,05
38,14
N/T
glm-5.2-int4
OOM
OOM
OOM
20,16
N/T
glm-5.2-nvfp4
OOM
OOM
OOM
OOM
24,86
*OOM = Out of Memory
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable
Time to First Token (TTFT)
Lower values are better.
Model
1x DGX Spark
2x DGX Spark Bundle
3x DGX Spark Triple
4x DGX Spark Quad
8x DGX Spark
gemma-4-31B-it-NVFP4
201,96 ms
136,35 ms
N/T
279,77 ms
326,54 ms
Qwen3.6-27B-NVFP4
233,30 ms
148,87 ms
N/T
163,59 ms
263,59 ms
Qwen3.6-35B-A3B-NVFP4
500,24 ms
165,31 ms
N/T
230,82 ms
N/T
MiniMax-M2.7-NVFP4, PP
OOM
396,35 ms, PP-2
211,39 ms, PP-3
N/T
N/T
MiniMax-M2.7-NVFP4, TP
OOM
288,98 ms
TP3-N/A
N/T
N/T
Qwen3.5-397B-A17B-int4
OOM
OOM
499,65 ms
324,61 ms
N/T
glm-5.2-int4
OOM
OOM
OOM
571,82 ms
N/T
glm-5.2-nvfp4
OOM
OOM
OOM
OOM
451,23 ms
*OOM = Out of Memory
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable
The results are specific to the tested model, quantization method, software version, and runtime parameters. Factors such as context length, prompt length, number of generated tokens, batch size, KV cache requirements, parallelization method, and inter-node communication overhead may affect the results.
Adding more Spark nodes does not guarantee linear TPS scaling or lower TTFT for every model. While some models can benefit from additional compute capacity, inter-node communication overhead may become the dominant factor for certain workloads.
Cluster size should not be selected solely based on the highest tokens-per-second value. Model memory fit, supported parallelization method, target context length, and the expected number of concurrent users should also be taken into consideration.
These values are not guaranteed minimum performance levels; they are reference results from configurations tested by OpenZeka.
Which Cluster Size?
Configuration
General Usage Approach
1x DGX Spark
Models that fit on a single node, development workloads, and low-concurrency inference
2x DGX Spark Bundle
Two-node tensor/pipeline parallel scenarios and higher total memory capacity
3x DGX Spark Triple
Three-node TP/PP scenarios and a compact architecture without a switch
4x DGX Spark Quad
Four-node TP/PP scenarios, larger models, and centralized compute fabric
8x DGX Spark
Distributed workloads requiring higher total memory capacity and support for eight nodes
The appropriate cluster size varies depending on the model. Before purchase, the target model, quantization method, context length, and intended usage scenario can be shared with OpenZeka for evaluation.
Box Contents
Standard Package
4 × NVIDIA DGX Spark systems or selected compatible GB10 systems
4 × System power adapters
1 × MikroTik CRS812 compute network switch
1 × CRS812 power connection kit
2 × 1.5 m FS 400G QSFP-DD to 2 × 200G QSFP56 passive copper breakout cables (1 m option available)
The package can be delivered either as a standard boxed, uninstalled solution or with complimentary pre-installation completed by OpenZeka, depending on customer requirements. For pre-installed deliveries, compute and management connections are configured, and basic system checks are performed.
The DGX Spark Quad AI Cluster installation process consists of physical hardware connections, compute switch port preparation, verification of 200Gbps link speeds, separation of management and compute networks, node IP address configuration, SSH access setup, and multi-node communication testing.
Customers can perform the cluster installation themselves or take advantage of OpenZeka’s complimentary pre-installation service.
With the uninstalled delivery option, DGX Spark systems, switches, and connection cables are shipped as standard package contents. Physical installation, switch configuration, IP addressing, RoCEv2/RDMA configuration, SSH access setup, and cluster validation are performed by the customer.
When complimentary pre-installation is selected, OpenZeka performs the following:
Prepares the compute and management networks
Performs basic configuration of the MikroTik CRS812 and CRS312 switches
Verifies 200GbE compute connections
Checks ConnectX-7 and RoCEv2/RDMA device visibility
Establishes basic inter-node access
Tests SSH access and basic cluster communication
Performs NCCL communication validation
The system is delivered ready for final customer-specific configuration, including IP addressing, VLAN, DNS, NTP, internet access, and security settings within the customer environment.
OpenZeka’s complimentary pre-installation service includes standard network preparation, node access setup, 200GbE connection verification, RoCEv2/RDMA device validation, NCCL communication testing, and basic cluster health checks. Customer-specific software deployment, model installation, security configuration, and integration services are scoped separately.