4.2 Network Analytics Engine (NAE) and Automated Telemetry

Key Takeaways

  • The Network Analytics Engine (NAE) is an on-box telemetry and diagnostics framework embedded directly within AOS-CX switches, eliminating the blind spots of periodic SNMP polling.

  • NAE combines a local Time-Series Database (TSDB) with event-driven Python scripts to monitor hardware and protocol states in real time.

  • When an anomaly or threshold condition occurs, NAE executes automated actions including CLI diagnostics snapshots, syslog generation, packet captures, and webhook notifications.

  • The AOS-CX Web UI dashboard visually correlates real-time time-series telemetry charts with system events and administrative configuration changes on a unified timeline.

  • NAE scripts can be deployed from Aruba Solution Exchange (ASE), customized, or authored from scratch using the native AOS-CX Python framework.

Last updated: October 2026

4.2 Network Analytics Engine (NAE) and Automated Telemetry

In traditional enterprise network operations, diagnosing intermittent network degradation—such as transient micro-burst congestion, sporadic packet drops, interface CRC errors, or Spanning Tree topology flaps—is notoriously difficult. Conventional monitoring relies primarily on Simple Network Management Protocol (SNMP) polling and centralized Syslog collectors. However, standard SNMP polling cycles typically occur at 5-minute intervals. A network anomaly lasting 10 or 15 seconds begins, causes severe application degradation, and clears entirely within the polling gap, leaving network engineers with zero diagnostic visibility.

Furthermore, when a failure occurs, administrators must scramble to log into multiple switches via CLI to capture operational state before the evidence disappears. Aruba Network Analytics Engine (NAE) fundamentally solves this problem by embedding a real-time analytics framework and automated diagnostics engine directly inside the AOS-CX operating system.


Limitations of Traditional Monitoring vs. Aruba NAE

Monitoring DimensionTraditional SNMP / Syslog PollingAruba Network Analytics Engine (NAE)
Data CollectionExternal NMS polls switch every 1 to 5 minutesReal-time, continuous on-box telemetry from state database
Detection LatencyPolling interval delay (minutes)Event-driven / sub-second threshold detection
Transient VisibilityCompletely misses micro-bursts and intermittent flapsCaptures millisecond-level spikes and state transitions
Root-Cause AnalysisManual CLI intervention required long after the incidentAutomated instantaneous diagnostic snapshot at moment of fault
Configuration CorrelationDisconnected audit logs; no correlation with telemetryUnified visual timeline correlating config changes with metrics
Network OverheadHigh bandwidth consumption from polling thousands of OIDsZero polling overhead; telemetry processed locally in hardware/OS

Core Architectural Components of NAE

Aruba NAE operates as a native extension of the AOS-CX state database architecture. It consists of four integrated pillars: the Time-Series Database (TSDB), NAE Python Scripts, Monitors and Conditions, and Automated Actions.

+-------------------------------------------------------------------------+
|                 NETWORK ANALYTICS ENGINE (NAE) ENGINE                   |
|                                                                         |
|   +=================================================================+   |
|   |           AOS-CX STATE DATABASE (OVSDB Telemetry Tables)        |   |
|   +=================================================================+   |
|                                    |                                    |
|                          Continuous Streaming                           |
|                                    v                                    |
|   +-----------------------------------------------------------------+   |
|   |                      NAE PYTHON AGENT                           |   |
|   |                                                                 |   |
|   |  1. MONITOR: Subscribes to OVSDB URIs (e.g. CRC drops, Tx load) |   |
|   |  2. TSDB: Writes high-resolution metrics to local time-series db|   |
|   |  3. CONDITION: Evaluates logic (e.g. Rate > 100 drops/sec)      |   |
|   +--------------------------------+--------------------------------+   |
|                                    | (When Condition Triggers)          |
|                                    v                                    |
|   +-----------------------------------------------------------------+   |
|   |                      AUTOMATED ACTIONS                          |   |
|   |  - Capture CLI State: 'show interface', 'show mac-address-table'|   |
|   |  - Pin Alert Marker to Web UI Timeline                          |   |
|   |  - Generate Syslog / SNMP Trap                                  |   |
|   |  - Send Webhook / REST Alert to Aruba Central / ServiceNow      |   |
|   +-----------------------------------------------------------------+   |
+-------------------------------------------------------------------------+

1. The Time-Series Database (TSDB)

  • Embedded directly within the switch operating system storage (available on CX 6200, 6300, 6400, 8100, 8325, and 8360 platforms).
  • Automatically captures and records chronological telemetry metrics, counter states, and performance statistics without requiring an external server or telemetry collector.
  • Retains high-resolution historical data (seconds, minutes, hours, and days) for immediate local trend analysis and anomaly detection.

2. NAE Scripts (Python Framework)

  • NAE analytics logic is written in standard Python 3 using the proprietary Aruba nae module.
  • Scripts are human-readable, modular, and open. Engineers can download pre-built scripts from the Aruba Solution Exchange (ASE) or write custom Python scripts to address environment-specific diagnostic needs.
  • Every script defines the parameters it intends to watch, the calculation logic for identifying an anomaly, and the precise diagnostic steps to execute when an issue is detected.

3. Monitors and Conditions

  • Monitors: Define the data streams to collect from OVSDB. A monitor can track physical interface metrics (input errors, CRC errors, transmit utilization, link flaps), environment metrics (fan speed, temperature, power supply status, SFP optical DOM parameters), or protocol states (OSPF neighbor transitions, BGP state changes, STP topology change notifications).
  • Conditions: Define mathematical rules or logical criteria applied to the monitored data. Examples include:
    • Threshold Crossing: Metric exceeds a static value (e.g., CPU utilization > 85%).
    • Rate of Change: Metric increases rapidly within a sliding window (e.g., interface discards increase by more than 50 packets per second over a 10-second interval).
    • State Transition: Protocol daemon changes operational state (e.g., OSPF neighbor changes from Full to Init or Down).

4. Automated Actions and Root-Cause Diagnostics

When a monitored condition is met, the NAE agent executes pre-programmed automated actions at the exact millisecond of the fault, long before human operators are aware of an incident:

  • Diagnostic CLI Command Execution: The agent automatically runs diagnostic commands (such as show interface <port> extended, show mac-address-table, or show ip route) and saves the exact terminal output into a diagnostic snapshot package stored in the TSDB.
  • Timeline Pinning: Places a visual alert indicator directly onto the Web UI time-series graph, highlighting the exact moment the threshold was breached.
  • Alerting: Predefined actions include syslog messages (ActionSyslog), captured CLI or shell output (ActionCLI, ActionShell), and custom reports. Alerts appear in the switch Web UI and can be consumed by external tools such as HPE Aruba Networking Central or a syslog/SIEM platform.
  • Automated Remediation: ActionCLI can run configuration commands as well as show commands (for example config, interface 1/1/1, shutdown), so a script can react to a condition. Use this carefully, because command strings are not validated before they run.

Web UI Visual Correlation and Root-Cause Analysis

A critical operational feature of NAE is its deep integration into the AOS-CX Web User Interface (Web UI). When administrators access the switch dashboard, NAE presents an intuitive visual analytics interface:

+-------------------------------------------------------------------------+
|                  AOS-CX WEB UI - NAE ANALYTICS TIMELINE                 |
|                                                                         |
|  Metric: Interface 1/1/1 CRC Error Rate                                 |
|                                                                         |
|  Errors/s                                                               |
|   150 |                                                                 |
|   100 |                        [!] ALERT: CRC Spike > 100/s             |
|    50 |                         *---*                                   |
|     0 |------------------------*     *----------------------------      |
|       +-------------------------+-----+---------------------------+     |
|       10:00                   10:15  10:20                      10:35   |
|                                   ^                                     |
|                                   |                                     |
|                           [i] Config Change Event:                      |
|                           Admin modified port speed on 1/1/1            |
+-------------------------------------------------------------------------+

The Unified Timeline Advantage

In the Web UI, NAE overlays three distinct categories of historical data onto a single synchronized timeline:

  1. Time-Series Metric Line: Continuous visual graph of the monitored metric (e.g., bandwidth, transceiver optical power, buffer queue depth).
  2. Alert Trigger Markers: High-visibility indicators (Normal, Minor, Major, Critical) showing exactly when an NAE condition was breached and when it returned to normal.
  3. Configuration Audit Trail: A chronological record of administrative actions. If an engineer modified the port MTU, changed an interface speed, or added a VLAN 30 seconds before an error storm began, the configuration event appears directly below the telemetry spike.

This immediate visual correlation allows network engineers to instantly determine root cause—differentiating between an external physical fiber fault and an administrative configuration error in seconds rather than hours.


Pre-Built vs. Custom NAE Scripts

Aruba provides a library of NAE scripts through the Aruba Solution Exchange (ASE) and the NAE scripts repository on GitHub (both referenced in the AOS-CX 10.14 NAE Guide). Common operational use cases include:

  • Faulty SFP / Transceiver Monitoring: Continuously tracks Digital Optical Monitoring (DOM) Rx/Tx optical power levels. If optical signal drops below transceiver thresholds, NAE flags a degraded cable or dirty fiber connector before the link drops entirely.
  • Loop and Broadcast Storm Detection: Monitors Layer 2 broadcast/multicast packets and Spanning Tree Topology Change Notifications (TCNs), instantly identifying the ingress port causing a network loop.
  • MTU Mismatch Identifier: Detects drops associated with oversized frames and alerts administrators to mismatched MTU settings across point-to-point links.
  • Routing Neighbor Flap Analytics: Diagnoses intermittent OSPF/BGP session resets, capturing routing table states and neighbor hello timers at the exact moment of adjacency collapse.
Loading diagram...
Aruba Network Analytics Engine (NAE) Telemetry and Action Flow
Test Your Knowledge

Which combination of architectural elements powers the on-box monitoring and automated diagnostics in Aruba Network Analytics Engine (NAE)?

A

External NetFlow collectors combined with centralized SNMP trap managers

B

Hardware ASIC loopback interfaces running proprietary bash shell scripts

C

An on-box time-series database, Python agents, and the OVSDB state database

D

An active-standby Linux kernel mirror with FTP diagnostic log shipping

Test Your Knowledge

How does the AOS-CX Web UI dashboard assist network engineers in diagnosing network anomalies detected by NAE?

A

It automatically initiates a remote desktop session to connected client workstations experiencing connectivity issues

B

It visually correlates time-series telemetry graphs with alert markers and configuration change events on a unified timeline

C

It blocks all web browser management access until the underlying network fault is resolved

D

It converts all switch configuration commands into legacy SNMP MIB syntax

Test Your Knowledge

Why is Aruba NAE fundamentally more effective at identifying intermittent network anomalies than traditional SNMP polling?

A

SNMP polling requires a manual console authorization for every metric that is queried by the NMS

B

SNMP cannot monitor physical port bandwidth, interface drop counters, or transceiver statistics

C

NAE watches state on the switch continuously, catching short spikes that fall between SNMP polls

D

NAE runs only on external cloud servers, so it never uses switch CPU while collecting data

Sections you finish are checked off in the contents.