Help us improve your experience.

Let us know what you think.

Do you have time for a two-minute survey?

 
 

Detect Anomalies for Device Hardware, Interface, and Protocol KPIs

Use this topic to understand how Routing Director monitors key performance indicators (KPIs) fpr device hardware, interfaces, and protocols and detects anomalies, and how you can use the GUI to view anomalies related the KPIs.

Detect Anomalies Overview

To assess the overall health of a network, you must monitor the health of devices, their interfaces, and the operation of network protocols. Routing Director uses artificial intelligence and machine learning (AI-ML) to continuously monitor KPIs related to device hardware, interface health, and protocol performance. Using AI-ML, Routing Director learns about the dynamic behavior of the KPIs, identifies changes in threshold patterns, and generates alerts when anomalies are detected.

In addition, Routing Director performs root-cause analysis (RCA) for device temperature anomalies that occur during device operation, helping you to quickly identify and address underlying issues.

After a device is onboarded successfully, Routing Director monitors the KPI, forecasts the range, and detects any anomalies that occur. During device operation, Routing Director detects device health anomalies (within 30 minutes) based on historical data for that device and the forecasted range. After a KPI value changes, the forecasted range takes approximately three hours to adapt and stabilize, after which anomaly detection continues based on the updated forecasted range.

To detect anomalies based on the dynamic thresholds for KPIs:

  • AI-ML must be enabled in the Routing Director cluster.

    See Deploy the Cluster for details.

  • Dynamic Threshold must be enabled in the device profile assigned to the device. See Table 8.

    Note:

    Enabling AI-ML operations require additional CPU, memory, and storage resources. For detailed information on the required capacity for AI-ML use cases, see Hardware Requirements.

    When you upgrade to Routing Director 2.9.0 from an older release, Dynamic Threshold is disabled even if it is enabled in the older release. To view dynamic thresholds for KPIs, ensure to enable Dynamic Threshold in the device profile, after the upgrade. See Table 8

Detecting anomalies based on dynamic threshold of KPIs is supported on the following device families—ACX, MX, and PTX.

RCA of Temperature Anomalies

When a device is in operation, Routing Director provides RCA for issues related to the Routing Engine temperature and Routing Engine CPU temperature. Routing Director analyzes the different attributes (CPU utilization percentage, fan RPM percentage, and inlet temperature) that could cause a temperature issue. Routing Director also compares the device's temperature to an expected range. Based on the analysis and comparison, Routing Director provides an alert, an expected reason for the issue, and details on the events that might have caused the issue. Figure 1 displays a sample page showing the RCA logs for an anomaly in the Routing Engine temperature.

Figure 1: Sample Page Showing RCA for Device Temperature Anomaly Line graph showing temperature from March 22-28 with thresholds at 100°C Critical and 95°C High. March 26, 44°C; CPU alert 50°C outside 30.77-55.57°C range.
  1

Device Temperature RCA Details

Device Hardware, Interface, and Protocol KPIs

Table 1 displays the device health KPIs that Routing Director monitors for each device.

Table 1: KPIs Related to Device Health
KPI Component Parameters
CPU

Routing Engine

Line card

CPU Utilization Percentage (%)
Memory

Routing Engine

Line card

Memory Utilization Percentage (%)
Fan Not applicable

RPM Percentage (%)

Temperature
  • Routing Engine (RE)

  • Routing Engine CPU

  • Line card

  • Line card CPU

Current temperature
Power supply unit (PSU) Power Power usage percentage (%) of the PSU.

For more information on the device hardware KPIs, see Hardware Data and Test Results.

Table 2 lists the KPIs related to interface health that Routing Director monitors for each interface.

Table 2: KPIs Related to Interface Health
KPI Description

Optics Rx power

Optics Tx power

Current optics power level in dBm.

Input traffic

Output traffic

Current traffic in Mbps.

Optical/Module temperature

Current optics temperature in ℃.

For more information on interface KPIs, see Interfaces Data and Test Results.

Table 3 lists the KPIs related to protocol performance detected using AI-ML.

Table 3: KPIs Related to Protocol Performance
KPI Parameters
BGP
  • Advertised routes

  • Received routes

RIB

Route count

You can view dynamic threshold for route count.

FIB

Route count

You can view dynamic threshold for route count.

For more information on Protocol KPIs, see Routing and MPLS Data and Test Results.

View Device Hardware, Interface, and Protocol KPI Anomalies on the GUI

You can view and monitor the device hardware, interface, and protocol KPI anomalies for a device respectively on the Hardware accordion, Interface accordion, and the Routing and MPLS accordion of the Device-Name page.

To view and monitor device hardware, interface, and protocol KPI anomalies:

  1. Select Observability > Troubleshoot Devices > Device-Name .

    The Device-Name page appears.

  2. To view:
    • KPIs related to device hardware, scroll to the Hardware accordion and click > to expand the accordion.

      • The Chassis section of the accordion displays the health status of the following KPIs monitored by Routing Director:

        • PSUs

        • Fans

        • CPUs

        • Memory

        • Temperature

      • Device events appear under Relevant Events with the following information:

        • Event notification message

        • Date and time that the last event was received by Routing Director.

    • KPIs related to Interface health, scroll to the Interface accordion and click > to expand the accordion. You can view the below KPIs:

      • Optical temperature, optical Tx power, and optical Rx power of pluggables

      • Input traffic

      • Output traffic

      • Interface events appear under Relevant Events with the following information:

        • Event notification message

        • Date and time that the last event was received by Routing Director.

    • KPIs related to protocol performance, scroll to the Routing and MPLS accordion and click > to expand the accordion. You can view KPIs for:

      • Routing Protocols—BGP (advertised routes and received routes)

      • Routing tables (RIB) and forwarding tables (FIB)

  3. Under Relevant Events, hover over or click View Details to view the details of the event, including the number of times that the event recurred.
  4. (Optional) Click View All Relevant Events to view all the health-related events for the device.

    The events appear on the Events for Device-Name page.

  5. You can view detailed information about each KPI related to device hardware, interface, or protocol KPI by doing the following:
    1. Click the health status link for the KPI. For example, Fans or Temperature.

      The Hardware details for Device-Name page appears, displaying the section for the KPI that you clicked in the preceding page.

      For example, if you click the link for Fans, then the Fans section is expanded and the graphs related to the fans are displayed.

      Figure 2 shows a sample section (Temperature) of the Hardware Details for Device-Name page.

      For interfaces, click the status link of the KPI; for example, Input Traffic or Output Traffic, Respective page appears displaying the graph related to the KPI. For example, clicking Input Traffic link opens the Input Traffic details for Device-Name page displaying graphs related to input traffic and details of alerts if any. Figure shows a sample graph for input traffic.

      See Interfaces Data and Test Results for details on the graphs and KPIs related to interfaces.

      Similarly, for protocols, click the status link of a protocol; for example, BGP or RSVP. The respective details page appears displaying the graph related to the protocol KPIs. For example, BGP opens the BGP Routing Details for Device-Name page. You can view graphs for:

      • Number of BGP routes advertised to neighbors

      • Number of BGP routes received from neighbors

      See Routing and MPLS Data and Test Results for details on the graphs and KPIs related to protocols.

    2. To view the details of an anomaly, click the yellow triangle icon on the graph.

      The details of the anomaly appear in a pop-up, as shown in Figure 2 for hardware and as shown in Figure for interfaces.

  6. Click Close or the X icon to go to the Device-Name page.

For more information on the hardware accordion, see Hardware Data and Test Results and on interface accordion, see Interfaces Data and Test Results.

Figure 2: Sample Hardware Details for Device-Name Page Graph showing fan speed monitoring in a system with fan list, critical thresholds, and alerts for performance issues.
  1

KPI

  5

High threshold marker

  2

Circle icons indicating that the KPI is normal

  6

Pop-up showing details of device health anomaly.

  3

Upper and lower boundaries (dynamic thresholds) for the data displayed in the graph

  7

Triangle icons indicating an anomaly when the dynamic threshold is breached.

  4

Critical threshold marker

  8

Legend showing the colors for different sub-components used in the graphs

Figure shows the input traffic through et-0/0/1 interface on a device during a 30 minute interval. A warning (indicated by the yellow triangle icon) is raised to indicate an anomaly in the projected input traffic rate,

Figure 3: Input Traffic Page Showing Anomalies in measured Input Traffic Input Traffic Page Showing Anomalies in measured Input Traffic

A KPI value is considered:

  • Anomalous if the KPI value is outside the dynamic threshold (shaded area of the map) for nine continuous minutes of data collection.

  • Normal if the KPI value falls within the dynamic threshold for six continuous minutes of data collection.

If the KPI value continues to be outside the dynamic threshold for more than nine consecutive minutes, the dynamic threshold adapts to the new values and a new dynamic threshold is created. An alert is raised if the KPI value crosses the High or Critical values irrespective of whether the value falls within the dynamic threshold or not.

Note:

For an oscillating KPI pattern, the dynamic threshold boundaries initially adapts correctly to KPI, but the boundaries begin to re-adjust and oscillate after a few hours, even when the KPI pattern remains unchanged, triggering false-positive alerts.