Help us improve your experience.

Let us know what you think.

Do you have time for a two-minute survey?

 
 

Monitor Fabric Health

Use this topic to determine how Routing Director monitors fabric health and generates alerts when fabric queue drop or fabric destination errors occurs.

Fabric health monitoring helps you detect and triage packet loss that originates inside the switching fabric of MX Series devices rather than issues related to ingress or egress interface drops, traffic policing, or general device health problems.

Routing Director uses AI-ML to monitor fabric-specific telemetry, such as priority-aware fabric queue drop behavior and fabric destination delivery failures and correlate multiple fabric-related signals to enable a network operator to identify whether traffic loss is originating because of congestion, hardware issues, or fabric-path degradation.

Fabric health monitoring complements existing blackhole and traffic loss detection capabilities by providing additional insight into packet-loss symptoms that occur inside a device fabric. A network operator can view the following issues in a device fabric:

  • Fabric queue drops

  • Fabric destination errors

You can view these fabric issues on the Events for Device-Name page as shown in Figure 1. You can access the Events for Device-Name page by clicking the x Unhealthy link displayed for the Fabric field on the Hardware accordion on the Device-Name page (Observability > Troubleshoot Devices > click a Device-Name).

You can also view alert for fabric queue drops and fabric destination errors under Related Events section of the Hardware accordion. For more details, see Table 1.

Figure 1: Fabric Issue Listed on Events for Device-Name Page Fabric Issue Listed on Events for Device-Name Page

The fabric health issues detected are classified as:

  • Critical—Indicates confirmed traffic loss or fabric destination delivery failure within the fabric.

  • Major—Indicates presence of fabric health issues but traffic loss is not confirmed (for example, symptoms that occur but do not persist long enough to confirm sustained loss).

  • Info—Provides information for fabric operational awareness such as fabric health states, historical records, and basic observations.

For Routing Director to detect fabric health:

  • AI-ML must be enabled in the Routing Director cluster.

  • Fabric Health Detection must be enabled in the device profile assigned to the device. See Fields in the Analytics tab.

  • Rules to collect fabric queue drop counters, fabric destination delivery failure indicators, and related fabric error signals must be configured.

Note:

Enabling AI-ML operations require additional CPU, memory, and storage resources. For detailed information on the required capacity for AI-ML use cases, see Hardware Requirements.