Skip to main content
User guide

Network

The network overview, the device inventory with per-interface alert policy, the monitoring-coverage baseline, and interface flapping and alert noise reduction.

The Network group in the sidebar has four pages: Network overview, Devices, Coverage and Capacity.

Network overview

Network overview shows the current state of the network on one screen.

Network overview

PanelContents
Device reachabilityDevices reachable by ICMP probes
SNMP collectionDevices being collected normally
Link statusInterfaces with a description and admin up that are down
Current alertsCritical and warning alerts
MTTA, MTTRMean time to acknowledge and to resolve in the selected window
Top 10Interfaces with the highest bandwidth utilization and the highest loss
ProblemsLinks that are down, devices that are unreachable or not being collected, rules with the most alerts

Flapping and noise

Needs a Professional license. Detects interface flapping (number of state changes and the peer) and makes two kinds of suggestions:

  • Merging duplicate alerts: when the same alert fires many times on one device, or on several devices at once (possibly one upstream cause), it suggests an Alertmanager group_by.
  • Inhibit rules: for example, inhibit the other alerts on a device while it is unreachable, or the flapping, error and loss alerts on an interface while it is down. The suggestions can be pasted into the Alertmanager configuration; the page shows whether each is already configured and how many current alerts it would inhibit.

Device inventory

Devices lists devices and interfaces. The network rule pack uses it to decide which interfaces should alert.

Devices

  • Discover: finds devices and interfaces from Prometheus scrape targets and SNMP metrics and recognises vendors by sysObjectID. It runs hourly by default and can be run by hand. Devices whose recognised vendor disagrees with their scrape labels, or that return no sysObjectID, are listed separately.
  • Import CSV, Export CSV and adding devices: maintain the device name, management IP, vendor, model, role, site, rack, owner and whether HA is configured. Fields you edit by hand are not overwritten by discovery.
  • Interface alert policy: each interface is default, key or ignored. Default: only interfaces with a description alert when they go down. Key: alerts even without a description. Ignored: no down, flapping, bandwidth or error alerts for the interface.
  • Session capacity by model: the session limit for each firewall model. The preset values come from vendor specifications; check them.
  • HA unit problems are only reported for devices marked as having HA configured.

The inventory is exported to Prometheus at /metrics/inventory, and the rule pack relies on it; see Data sources.

Monitoring coverage

Coverage compares the inventory, the scrape targets and the loaded alert rules to find gaps in monitoring. Check again refreshes it.

Coverage

CheckWhat to do
In the inventory but not collectedAdd the device to the snmp_exporter scrape configuration. Devices marked as managed but not collected are the most serious
Collected but not in the inventoryRun discovery on the Devices page
Interfaces without a descriptionUndescribed interfaces do not alert when they go down by default. Add an ifAlias, or set the interface to key or ignored in the inventory
Key interfaces without alert rulesInterfaces with a description and admin up, and key interfaces, should all have down and utilization rules
Alert rule qualityRules without for fire on every blip; rules without a runbook leave on-call guessing; severities other than critical or warning; labels containing $value cause high cardinality

On this page