NVIDIA Official Embedded Compute Distributor (MEA)
الموزع الرسمي للحوسبة المدمجة من NVIDIA (الشرق الأوسط وأفريقيا)
Официальный дистрибьютор встроенных вычислительных систем NVIDIA (MEA)
NVIDIA Resmi Gömülü Sistem Distribütörü (MEA)
NVIDIA Official Embedded Compute Distributor (MEA)
الموزع الرسمي للحوسبة المدمجة من NVIDIA (الشرق الأوسط وأفريقيا)
Официальный дистрибьютор встроенных вычислительных систем NVIDIA (MEA)
NVIDIA Resmi Gömülü Sistem Distribütörü (MEA)

NVIDIA DGX Spark Quad AI Cluster – 4 Node, 512 GB, 200GbE

NVIDIA DGX Spark Quad AI Cluster

This package includes four NVIDIA DGX Spark systems, one MikroTik CRS812 high-speed compute switch, one MikroTik CRS312 10GbE in-band management switch, and the required compute and management connection cables.

The DGX Spark Quad AI Cluster combines four Grace Blackwell systems into a switch-based high-speed network architecture. Each node connects to a 200GbE compute network via NVIDIA ConnectX-7, while a separate 10GbE Ethernet network is used for system management and standard network traffic.

The compute network supports RoCEv2-based RDMA communication, enabling high-bandwidth data transfer between nodes for NCCL and distributed AI workloads. This architecture isolates AI communication from internet and management traffic.

In OpenZeka benchmarks, large language models from the GLM-5.2, Qwen3.5, Qwen3.6, and Gemma model families were deployed on this four-node cluster. The evaluated models included GLM-5.2-INT4, Qwen3.5-397B-A17B-INT4, Gemma-4-31B-it-NVFP4, Qwen3.6-27B-NVFP4, and Qwen3.6-35B-A3B-NVFP4.
REQUEST FOR QUOTE
SKU: DGXSPARK-QUAD Category: Tags: , OpenZeka Elite Partner

Description

NVIDIA DGX Spark Quad AI Cluster

The NVIDIA DGX Spark Quad AI Cluster is a compact and scalable AI cluster that integrates four NVIDIA DGX Spark systems through a high-speed, switch-based network architecture.

The package includes four NVIDIA DGX Spark systems, one MikroTik CRS812 compute switch, two 400G QSFP-DD to 2 × 200G QSFP56 breakout cables, one MikroTik CRS312 in-band management switch, and six Cat6 10GbE Ethernet cables.

The DGX Spark Quad AI Cluster can be delivered either without installation or with complimentary pre-installation by OpenZeka, depending on the customer’s preference. With the uninstalled option, all components are shipped in their standard packaging, and installation is performed by the customer. If complimentary pre-installation is selected, the compute and management networks are configured, all connections are verified, and basic cluster communication is validated prior to shipment.

Each DGX Spark node connects to the compute switch via a 200GbE link using the NVIDIA ConnectX-7 network interface. The compute network supports RoCEv2-based RDMA communication, enabling NCCL collective operations, model parallelism, and high-speed data transfer for distributed AI workloads. Management, SSH access, internet connectivity, model downloads, and standard network traffic are handled over a separate 10GbE management network. This architecture dedicates the high-speed compute network exclusively to inter-node AI communication.

Scalable AI Infrastructure with Four Nodes

Each NVIDIA DGX Spark system includes 128 GB of coherent unified system memory. The four-node configuration provides a total of:

  • 512 GB of distributed unified memory capacity
  •  16 TB of local NVMe storage

The four-node architecture provides a suitable development environment for:

  • Running large models that cannot fit into the total memory capacity of one, two, or three Spark systems by distributing them across four nodes using tensor parallelism, pipeline parallelism, or other supported distributed methods
  • Running models that can fit on fewer Spark systems with higher throughput when appropriate parallelization support is available
  • Providing greater total capacity for model weights, KV cache, and runtime memory in long context length scenarios
  • Increasing overall system capacity in inference scenarios where more concurrent users or requests are served by the same model
  • Running large language models through distributed inference methods
  • Multi-node communication based on NCCL and MPI
  • RAG and agentic AI applications
  • High-speed node-to-node data communication over RoCEv2/RDMA
  • RDMA-based infrastructure designed to reduce network processing overhead on host CPUs during GPU communication
  • Multimodal generative AI workloads
  • Model quantization and performance testing
  • Prototyping distributed architectures before transitioning to data center environments

Model compatibility and achievable performance vary depending on the model architecture, quantization format, context length, KV cache requirements, framework support, batch size, and parallelization method.

Switch-Based Compute Network

The compute network is built on the MikroTik CRS812 switch.

The CRS812 provides the following high-speed connectivity ports:

  • 2 × 400G QSFP56-DD ports
  • 2 × 200G QSFP56 ports
  • 8 × 50G SFP56 ports

The switch also features RouterOS v7, a quad-core 2 GHz ARM processor, 4 GB of RAM, redundant hot-swappable power supplies, and hot-swappable fans. MikroTik positions this model for AI cluster environments and laboratory deployments that require high-speed east-west traffic.

In the DGX Spark Quad configuration, two 400G ports of the CRS812 are used. Each 400G port is split into two 200G QSFP56 connections using a passive breakout cable. This provides a dedicated 200GbE physical connection for each Spark system.

The DGX Spark Quad compute network is configured to leverage the RoCEv2 and RDMA capabilities of NVIDIA ConnectX-7 adapters. RoCEv2 enables RDMA communication over Ethernet and IP infrastructure, providing high-bandwidth data transfer between nodes with low communication overhead.

This architecture is particularly used for intensive data transfers between nodes during NCCL collective operations, tensor parallelism, pipeline parallelism, and distributed inference workloads.

RoCE performance can be affected by switch queues, MTU configuration, traffic classification, PFC/flow control, ECN, and software-level settings. In OpenZeka deployments, link speeds, RDMA devices, and inter-node communication are additionally verified.

Dedicated 10GbE Management Network

The in-band management network is provided through the MikroTik CRS312-4C+8XG-RM.

The CRS312 features:

  • 8 × 10G RJ45 Ethernet ports
  • 4 × 10G RJ45/SFP+ combo ports
  • 120 Gbps non-blocking throughput
  • 240 Gbps switching capacity
  • RouterOS v7 or SwitchOS support

Each Spark system is connected to the CRS312 using a single Cat6 Ethernet cable. This network is used for:

  • SSH access
  • System management
  • Software updates
  • Model downloads
  • Internet access
  • Monitoring and log traffic
  • Optional NAS access

The separation of compute and management networks prevents standard management and internet traffic from traversing the ConnectX-7 compute links, ensuring that the high-speed network is dedicated to distributed AI communication.

Optional Shared NAS Storage

The cluster can optionally be delivered with an ASUSTOR Lockerstor AS6808T NAS and eight NAS-grade drives. In the standard example configuration, 8 × 8 TB Western Digital WD80EFPX-68C4ZN0 drives are used. The disk brand, model number, and capacity can be customized based on stock availability or project requirements.

Example 8 × 8 TB Configuration:

Configuration

Approximate Raw Capacity

Disk Protection

RAW/JBOD

64 TB

Depends on configuration

RAID 0

64 TB

None

RAID 5

56 TB

1 disk

RAID 6

48 TB

2 disks

The actual usable capacity will be lower than the raw values shown in the table due to the disk manufacturer’s capacity calculation method, RAID metadata, file system overhead, and system-reserved space.

The NAS is connected to the MikroTik CRS312 switch using dual LACP-supported links with two Cat6 Ethernet cables.

Bonding can provide advantages such as:

  • Link redundancy
  • Load balancing across multiple clients
  • Distribution of multiple simultaneous data streams

The disk brand, disk capacity, and RAID configuration can be customized according to project requirements. Depending on stock availability, technically equivalent NAS-grade drives with the same capacity and usage class may be provided instead of the specified disk model.

OpenZeka Cluster Solution

The DGX Spark Quad AI Cluster is not simply a hardware package consisting of four computers placed together. By integrating a high-speed compute network, a dedicated management network, breakout connectivity infrastructure, and optional shared storage, it provides a comprehensive infrastructure designed for distributed AI workloads.

The system can be delivered either as a complete uninstalled package with all components included or shipped with complimentary pre-installation performed by OpenZeka, depending on customer preference. As part of the complimentary pre-installation service, compute and management connections are configured, 200GbE links are verified, basic node access is established, and cluster communication is validated.

Additional services such as customer-specific model deployment, custom network policies, enterprise integrations, data migration, application installation, and performance optimization are evaluated separately within the scope of each project.

Additional information

Architecture

NVIDIA Grace Blackwell

GPU

NVIDIA Blackwell

CPU

CUDA Cores

NVIDIA Blackwell Generation

Tensor Cores

5th Generation

RT Cores

4th Generation

Tensor performance

1 PFLOP

Memory

128 GB 256-bit LPDDR5X 273 GB/s

Memory Interface

256-bit

Memory bandwidth

273 GB/s

Storage

4 TB NVME.M2 with self-encryption

USB

4x USB TypeC

Ethernet

1x RJ-45 connector 10 GbE

NIC

ConnectX-7 NIC @ 200 Gbps

WiFi

WIFI 7

Bluetooth

BT 5.4 w/LE

Audio

HDMI multichannel audio output

Power consumption

240 W

Display

1.4a/HDMI 2.1

OS

NVIDIA DGX™ OS

Dimension

150 mm L x 150 mm W x 50.5 mm H

Weight

1.2 kg

Detailed Information

Explore the detailed specifications, architecture, and configuration details of the NVIDIA DGX Spark Quad AI Cluster, including hardware components, networking capabilities, and optional storage features.

A. Cluster General Features:

FeaturesValue
Node Count4
System PlatformNVIDIA DGX Spark or selected compatible GB10 system
Total System Memory512 GB distributed unified memory
Memory per Node128 GB LPDDR5x
Total Local Storage16 TB NVMe in a 4 × 4 TB configuration
AI Performance per NodeUp to 1 PFLOP FP4
Compute Connectivity200GbE per node
ConnectX-7 Interface ArchitectureMay appear as multiple logical interfaces depending on operating system and driver configuration
In-Band Management10GbE per node
Compute TopologyEthernet switch-based star topology
Optional Shared Storage64 TB RAW NAS
Parallel ExecutionTensor parallelism, pipeline parallelism, distributed inference
Operating SystemNVIDIA DGX OS or supported NVIDIA software environment of the selected system
Compute ProtocolRDMA-enabled Ethernet over RoCEv2
RDMA AdapterNVIDIA ConnectX-7
Compute UsageNCCL, model parallelism, and distributed inference communication
Network SeparationDedicated 200GbE compute and 10GbE management networks

B. Node Specifications (For NVIDIA DGX Spark):

SpecificationPer Node
ArchitectureNVIDIA Grace Blackwell
SuperchipNVIDIA GB10
GPUNVIDIA Blackwell
CPU20-core Arm, 10 × Cortex-X925 + 10 × Cortex-A725
Tensor Cores5th Generation
RT Cores4th Generation
AI PerformanceUp to 1 PFLOP FP4
System Memory128 GB LPDDR5x coherent unified memory
Memory BandwidthUp to 273 GB/s
Storage4 TB NVMe M.2
Standard Ethernet10GbE RJ45
High-Speed NICConnectX-7
WirelessWi-Fi 7
BluetoothBluetooth 5.4
USB4 × USB Type-C
Display OutputHDMI 2.1a
Operating SystemNVIDIA DGX OS
Dimensions150 × 150 × 50.5 mm
WeightApproximately 1.2 kg

C. Network Components:

ComponentFunction
MikroTik CRS812High-speed Ethernet switch carrying RoCEv2 compute traffic
FS QDD-400G-2QPC015400G QSFP-DD to 2 × 200G QSFP56 breakout
Cable Count2
Cable Length1.5 m (1 m optional)
MikroTik CRS312Handles management, internet, SSH, and NAS traffic
Spark Management Cables6 × Cat6
Compute Connectivity200GbE per Spark node
Management Connectivity10GbE per Spark node
RoCEv2/RDMAHigh-speed compute communication between Spark nodes
NVIDIA ConnectX-7RoCEv2 and RDMA-enabled node network adapter

D. Optional NAS Features:

SpecificationExample Configuration
NASASUSTOR Lockerstor AS6808T
Drive Bays8
Standard Drive Configuration8 × 8 TB
Example Drive ModelWestern Digital WD80EFPX-68C4ZN0
RAW Capacity64 TB
RAID OptionsRAID 0, RAID 5, RAID 6, and other levels supported by the device
Network Connectivity2 × Ethernet, IEEE 802.3ad LACP bonding
Connected NetworkMikroTik CRS312 management/storage network
CustomizationDisk brand, model number, and capacity can be customized

Performance Tests

The results below are obtained from tests conducted by OpenZeka on different DGX Spark topologies. All measurements are reported with Concurrency = 1 and using Mean statistics.

The purpose of presenting the results for 1, 2, 3, 4, and 8 Spark configurations together is to allow users to compare suitable cluster sizes based on the model they plan to run, the parallelization method, and their expected performance requirements.

TP stands for Tensor Parallelism, and PP stands for Pipeline Parallelism. Empty fields indicate that no test result is currently available for the related model and topology, or that the configuration could not be applied due to the model’s memory requirements and parallelization constraints.

Generation Throughput (TPS)

Higher values are better.

Model1x DGX Spark2x DGX Spark Bundle3x DGX Spark Triple4x DGX Spark Quad8x DGX Spark
gemma-4-31B-it-NVFP410,8119,42N/T30,5336,93
Qwen3.6-27B-NVFP412,6322,57N/T33,1133,99
Qwen3.6-35B-A3B-NVFP464,5291,95N/T91,16N/T
MiniMax-M2.7-NVFP4, PPOOM17,20, PP-218,04, PP-3N/TN/T
MiniMax-M2.7-NVFP4, TPOOM26,64TP3-N/AN/TN/T
Qwen3.5-397B-A17B-int4OOMOOM17,0538,14N/T
glm-5.2-int4OOMOOMOOM20,16N/T
glm-5.2-nvfp4OOMOOMOOMOOM24,86
*OOM = Out of Memory
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable

Time to First Token (TTFT)

Lower values are better.

Model1x DGX Spark2x DGX Spark Bundle3x DGX Spark Triple4x DGX Spark Quad8x DGX Spark
gemma-4-31B-it-NVFP4201,96 ms136,35 msN/T279,77 ms326,54 ms
Qwen3.6-27B-NVFP4233,30 ms148,87 msN/T163,59 ms263,59 ms
Qwen3.6-35B-A3B-NVFP4500,24 ms165,31 msN/T230,82 msN/T
MiniMax-M2.7-NVFP4, PPOOM396,35 ms, PP-2211,39 ms, PP-3N/TN/T
MiniMax-M2.7-NVFP4, TPOOM288,98 msTP3-N/AN/TN/T
Qwen3.5-397B-A17B-int4OOMOOM499,65 ms324,61 msN/T
glm-5.2-int4OOMOOMOOM571,82 msN/T
glm-5.2-nvfp4OOMOOMOOMOOM451,23 ms

*OOM = Out of Memory
*N/T = Not Tested
*TP3-N/A = TP3 Not Applicable

The results are specific to the tested model, quantization method, software version, and runtime parameters. Factors such as context length, prompt length, number of generated tokens, batch size, KV cache requirements, parallelization method, and inter-node communication overhead may affect the results.

Adding more Spark nodes does not guarantee linear TPS scaling or lower TTFT for every model. While some models can benefit from additional compute capacity, inter-node communication overhead may become the dominant factor for certain workloads.

Cluster size should not be selected solely based on the highest tokens-per-second value. Model memory fit, supported parallelization method, target context length, and the expected number of concurrent users should also be taken into consideration.

These values are not guaranteed minimum performance levels; they are reference results from configurations tested by OpenZeka.

Which Cluster Size?

ConfigurationGeneral Usage Approach
1x DGX SparkModels that fit on a single node, development workloads, and low-concurrency inference
2x DGX Spark BundleTwo-node tensor/pipeline parallel scenarios and higher total memory capacity
3x DGX Spark TripleThree-node TP/PP scenarios and a compact architecture without a switch
4x DGX Spark QuadFour-node TP/PP scenarios, larger models, and centralized compute fabric
8x DGX SparkDistributed workloads requiring higher total memory capacity and support for eight nodes

The appropriate cluster size varies depending on the model. Before purchase, the target model, quantization method, context length, and intended usage scenario can be shared with OpenZeka for evaluation.

Box Contents

Standard Package

  • 4 × NVIDIA DGX Spark systems or selected compatible GB10 systems
  • 4 × System power adapters
  • 1 × MikroTik CRS812 compute network switch
  • 1 × CRS812 power connection kit
  • 2 × 1.5 m FS 400G QSFP-DD to 2 × 200G QSFP56 passive copper breakout cables (1 m option available)
  • 1 × MikroTik CRS312 10GbE in-band management switch
  • 1 × CRS312 power cable
  • 6 × Cat6 10GbE Ethernet cables

Optional NAS Package

  • 1 × ASUSTOR Lockerstor AS6808T NAS
  • 8 × 8 TB NAS HDDs
  • 2 × Cat6 Ethernet cables
  • Selected RAID configuration

The package can be delivered either as a standard boxed, uninstalled solution or with complimentary pre-installation completed by OpenZeka, depending on customer requirements. For pre-installed deliveries, compute and management connections are configured, and basic system checks are performed.

Product Manuals

Installation

The DGX Spark Quad AI Cluster installation process consists of physical hardware connections, compute switch port preparation, verification of 200Gbps link speeds, separation of management and compute networks, node IP address configuration, SSH access setup, and multi-node communication testing.

Customers can perform the cluster installation themselves or take advantage of OpenZeka’s complimentary pre-installation service.

With the uninstalled delivery option, DGX Spark systems, switches, and connection cables are shipped as standard package contents. Physical installation, switch configuration, IP addressing, RoCEv2/RDMA configuration, SSH access setup, and cluster validation are performed by the customer.

When complimentary pre-installation is selected, OpenZeka performs the following:

  • Prepares the compute and management networks
  • Performs basic configuration of the MikroTik CRS812 and CRS312 switches
  • Verifies 200GbE compute connections
  • Checks ConnectX-7 and RoCEv2/RDMA device visibility
  • Establishes basic inter-node access
  • Tests SSH access and basic cluster communication
  • Performs NCCL communication validation

The system is delivered ready for final customer-specific configuration, including IP addressing, VLAN, DNS, NTP, internet access, and security settings within the customer environment.

OpenZeka’s complimentary pre-installation service includes standard network preparation, node access setup, 200GbE connection verification, RoCEv2/RDMA device validation, NCCL communication testing, and basic cluster health checks. Customer-specific software deployment, model installation, security configuration, and integration services are scoped separately.

Title

Go to Top