Frontend Fabric Topology
The validated frontend fabric topology follows a three-stage Clos leaf-spine IP fabric architecture with a 3:1 subscription factor.
Figure 3 shows the high-level frontend fabric topology used during testing.
Figure 3: AI Inference Frontend Fabric Topology
The topology used to validate the design included 4 leaf nodes and 2 spine nodes. As described in Table 6, we validated QFX5130-32CD, QFX5140-24CD8O and QFX5240-64OD in the leaf node role, and QFX5220-32CD, QFX5230-64CD and QFX5240-64OD in the spine role.
Each frontend leaf connects to both spine nodes using 2 x 400GbE Ethernet links, providing redundant and scalable connectivity across the frontend fabric.
The AMD Instinct MI300X GPU servers connect to leaf nodes 3 and 4 using 400GbE Ethernet links with Connect X7 NICs.
The Lambda scaler devices running Envoy Proxy and GenAI-Perf connect to leaf nodes 1 and 2 using 100GbE Ethernet links with ConnectX-6 NICs.
Note: Additional NICs and link speeds may be added in future updates of this JVD.
HPE Juniper Apstra Data Center Director assigns the fabric IP addressing, autonomous system numbers, and other network parameters, and then creates and deploys the fabric configuration. The point-to-point links between leaf and spine nodes are assigned /31 addresses from the 10.0.5.0/24 address range, as shown in Figure 3.
Figure 3: AI Inference Frontend Fabric Implementation Details
The fabric uses eBGP between leaf and spine nodes to provide IP reachability, path redundancy, and equal-cost forwarding across the Clos topology.