Help us improve your experience.

Let us know what you think.

Do you have time for a two-minute survey?

 
 

Frontend Fabric Topology

The validated frontend fabric topology follows a three-stage Clos leaf-spine IP fabric architecture with a 3:1 subscription factor.

Figure 3 shows the high-level frontend fabric topology used during testing.

Figure 3: AI Inference Frontend Fabric Topology

The topology used to validate the design included 4 leaf nodes and 2 spine nodes. As described in Table 6, we validated QFX5130-32CD, QFX5140-24CD8O and QFX5240-64OD in the leaf node role, and QFX5220-32CD, QFX5230-64CD and QFX5240-64OD in the spine role.

Each frontend leaf connects to both spine nodes using 2 x 400GbE Ethernet links, providing redundant and scalable connectivity across the frontend fabric.

The AMD Instinct MI300X GPU servers connect to leaf nodes 3 and 4 using 400GbE Ethernet links with Connect X7 NICs.

The Lambda scaler devices running Envoy Proxy and GenAI-Perf connect to leaf nodes 1 and 2 using 100GbE Ethernet links with ConnectX-6 NICs.

HPE Juniper Apstra Data Center Director assigns the fabric IP addressing, autonomous system numbers, and other network parameters, and then creates and deploys the fabric configuration. The point-to-point links between leaf and spine nodes are assigned /31 addresses from the 10.0.5.0/24 address range, as shown in Figure 3.

Figure 3: AI Inference Frontend Fabric Implementation Details

The fabric uses eBGP between leaf and spine nodes to provide IP reachability, path redundancy, and equal-cost forwarding across the Clos topology.