# Support for Minor and Major alarm status and generate alarm as SNMP trap notification

**URL:** <https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616>\
**Category:** Help\
**Tags:** agent\
**Created:** [August 12, 2021, 4:17am UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616 "2021-08-12T04:17:53Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vasanthakumarmurugan](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/vasanthakumarmurugan/32/1066_2.png) [@Vasanthakumarmurugan](https://community.netdata.cloud/u/Vasanthakumarmurugan)\
**Post date:** [August 12, 2021, 4:17am UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/1 "2021-08-12T04:17:53Z")

</div>

## Environment

Linux - Redhat

## Problem/Question

Currently, we’re looking for an alternative to IBM’s Netcool/SSM (resource monitoring agent ) due to EOS. We come across Netdata and evaluating it

In “Configure health alarms”, could see the options to set conditions for warning and critical threshold limits.

In our existing alarm implementation, we have alarm statuses like “warning”, “minor”, “major”, “critical”, and “clear”, and we have the option to customize the alarm messages based on the status.

Also, we have the option to directly pass the alarm notifications as SNMP traps to the external backend.

## What I expected to happen

1. Options to trigger and clear minor and major alarms based on threshold limits
2. Options to configure custom alarm messages (for both raise and clear) based on the status
3. Options to generate an alarm as SNMP trap notification to external backend

---

<div class="post-metadata">

**Author:** ![andrewm4894](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/andrewm4894/32/125_2.png) [@andrewm4894](https://community.netdata.cloud/u/andrewm4894)\
**Post date:** [August 17, 2021, 10:20am UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/2 "2021-08-17T10:20:48Z")

</div>

I’m not sure if we have the ability to define custom alarm statuses but i know it’s come up in the past as something people might sometimes want/need.

I know there are some SNMP examples in here:

> **[SNMP devices | Learn Netdata](https://learn.netdata.cloud/docs/collecting-metrics/generic-collecting-metrics/snmp-devices)**
>
> Plugin: go.d.plugin

But i’m not quite sure if we have any example alarm configs based on this.

I do recall there may be some examples in some of these threads:

> [@Encountered multiple issues trying to set an alarm on bandwidth overuse using snmp fetched metric](https://community.netdata.cloud/t/encountered-multiple-issues-trying-to-set-an-alarm-on-bandwidth-overuse-using-snmp-fetched-metric/1103):
>
> Environment Running netdata/netdata:v1.29.3 docker image on top of aarch64 architecture Problem/Question This is the snmp.conf I use for collecting my router BW info (faked some data for security reasons): { "enable\_autodetect": false, "update\_every": 5, "max\_request\_size": 100, "servers": [ { "hostname": "192.161.1.1", "community": "my\_comunity", "update\_every": 10, "max\_request\_size": 50, "options": { …

> [@SNMP device monitoring](https://community.netdata.cloud/t/snmp-device-monitoring/774):
>
> Hi I followed the SNMP device monitoring document and created a /etc/netdata/node.d/snmp.conf file. { "enable\_autodetect": false, "update\_every": 10, "max\_request\_size": 100, "servers":[ { "hostname": "x.x.x.x", "community": "public", "update\_every": 10, "max\_request\_size": 50, "options": { "timeout": 10000 }, "charts": { "snmp\_router\_bandwith\_gateway": { "title": "60B Router Bandwidth for Gateway", "units": "kilobits/s", "type": "area", "…

> [@Reporting metrics for remote "dumb" hosts (printers, switches, etc) - seperate "node"?](https://community.netdata.cloud/t/reporting-metrics-for-remote-dumb-hosts-printers-switches-etc-seperate-node/65):
>
> I’m evaluating Netdata Cloud as a potential replacement for my LibreNMS setup. In that system, devices monitored by SNMP are added as seperate hosts with the metrics which were collected. I’ve managed to set up SNMP through the node plugin on one of my hosts, but I see these metrics (for example, printer toner level) are only available as a metric as a part of the host that collected the metric. Is there (or is there planned?) a way to have my netdata node report the health of another node, so t…

> [@Does Netdata support SNMP monitoring?](https://community.netdata.cloud/t/does-netdata-support-snmp-monitoring/1061):
>
> Netdata does support SNMP with the [SNMP collector](https://learn.netdata.cloud/docs/agent/collectors/node.d.plugin/snmp). If you’d like to see more SNMP capabilities, send us a [feature request](https://community.netdata.cloud/c/feature-requests/agent-fr/21) and we’ll be happy to discuss and learn about your use case!

---

<div class="post-metadata">

**Author:** ![Vasanthakumarmurugan](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/vasanthakumarmurugan/32/1066_2.png) [@Vasanthakumarmurugan](https://community.netdata.cloud/u/Vasanthakumarmurugan)\
**Post date:** [August 18, 2021, 11:44am UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/3 "2021-08-18T11:44:54Z")

</div>

Thanks, Andrew for the response and suggestions.

Regarding SNMP, I’ve gone the links shared. I could see examples for SNMP device monitoring. But couldn’t find an example to generate an alarm as SNMP trap notifications to external . Please suggest, if you have any links that could help me.

Also, I’d like to know, if there is a way to define custom messages for alarm rising and clearing.  
E.g.  
CPU usage has been above 85% for the last 5 minutes (currently 98%)  
CPU usage has returned to normal (less than 85%)

---

<div class="post-metadata">

**Author:** ![OdysLam](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/odyslam/32/98_2.png) [@OdysLam](https://community.netdata.cloud/u/OdysLam)\
**Post date:** [August 20, 2021, 10:11am UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/4 "2021-08-20T10:11:06Z")

</div>

Hey,

Sorry for the late response, I was in PTO and just returned!

In general, netdata is not idea for SNMP, but it can support it. SNMP is just another “collector”, a data source for Netdata. Whatever we can do to the other gathered dimensions, we can also do it to SNMP values.

1. Netdata supports 3 thresholds and their respective status: CLEAR, WARNING, CRITICAL.
2. It is possible to do that, but you will possibly need to create a custom notification method. Notification and alarms are 2 different modules in the architecture. Netdata’s health module will constantly evaluate alarms and raise them when needed, while it will call the notification script to issue any appropriate notification.
3. This is not currently possible and I can’t imagine that we move to implement it in the near future. Perhaps it’s a great first feature for you/your team to contribute to the open source netdata agent!

The notification script is a bash script that supports many different notification methods, including defining a custom one. In that custom notification method, you could code any logic you want for the message and then call an API to send out the notification (e.g Twillio, Pagerduty, etc.).

Some helpful documentation:

- [Creating your first health alarm in Netdata - YouTube](https://www.youtube.com/watch?v=aWYj9VT8I5A)
- [Alarm notifications | Learn Netdata](https://learn.netdata.cloud/docs/agent/health/notifications)
- [Configure health alarms | Learn Netdata](https://learn.netdata.cloud/docs/monitor/configure-alarms)

---

<div class="post-metadata">

**Author:** ![Vasanthakumarmurugan](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/vasanthakumarmurugan/32/1066_2.png) [@Vasanthakumarmurugan](https://community.netdata.cloud/u/Vasanthakumarmurugan)\
**Post date:** [September 6, 2021, 1:31pm UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/5 "2021-09-06T13:31:57Z")

</div>

Hi OdysLam,

Thanks for the response and suggestions.

I defined custom logic and used it in custom notification method. It works fine.

Thanks,  
Vasanth

---

<div class="post-metadata">

**Author:** ![Christopher\_Akritid1](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/christopher_akritid1/32/419_2.png) [@Christopher\_Akritid1](https://community.netdata.cloud/u/Christopher_Akritid1)\
**Post date:** [September 7, 2021, 2:30am UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/6 "2021-09-07T02:30:44Z")

</div>

So did you create a custom notification function to issue SNMP traps? It’s a bit of a niche need, but perhaps you can share the function so maybe it can eventually become a supported notification method.

---

<div class="post-metadata">

**Author:** ![Vasanthakumarmurugan](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/vasanthakumarmurugan/32/1066_2.png) [@Vasanthakumarmurugan](https://community.netdata.cloud/u/Vasanthakumarmurugan)\
**Post date:** [September 13, 2021, 11:38am UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/7 "2021-09-13T11:38:42Z")

</div>

Hi Christopher,

Sure, I can share it. However, it is not a production-ready solution. It would be great if Netdata could improve and make it a standard solution.

## Alarm definition health.d/cpu.conf

```
	alarm: 111111
	   on: system.cpu
	class: Utilization
	 type: System
component: CPU
	   os: linux
	hosts: *
   lookup: average -30s unaligned of user,system
	units: %
	every: 10s
	 warn: $this > 60
	 info: CPU Usage Warning
	   to: sysadmin

```

## INI file:

```
[111111]
info=CPU Usage Warning
moduleId=System Resources
<Skipping...>
raisetext="CPU Usage Warning, CPU usage has been above 60% for the last 5 minutes (currently ${value}%)"
falltext="CPU Usage Warning, CPU usage has returned to normal (less than 60%)"

```

## Wrapper script:

```
#!/bin/bash

inifile="</path/to/ini file>" # Replace here

# declare an associative array
declare -A conf

# Functions
#-------------

# Parse alarm translation INI file using regex matches
# Take out only the necessary infomation and create array
Get_INI_Section ()
{

		local filename="$1"
		local section="$2"
		if [-f "$filename"] && [-n "$section"]; then
				block_Start="^[\t]*\[${section}\][\t]*$"
				block_end='^[\t]*\[[^]]+\][\t]*$'

				while IFS='=' read -r key val ; do
				   conf[$key]="${val}"
				done < <( sed -nre "/${block_Start}/, /${block_end}/ {
												s/${block_Start}//; # Skip the start of the match pattern
												s/${block_end}//; # Skip the end of the match pattern
												s/^\s*//; # Trim white spaces before
												s/\s*$//; # Trim white spaces after
												s/\#.*//; # remove comments from the lines
												s/\s*=\s*/=/; # remove spaces before and after =
												/^$/ d; # Delete empty lines
												p; # print the content
				}" "$filename" )

				#for key in "${!conf[@]}"
				#do
				# echo "$key", "${conf[$key]}"
				#done
		else
				echo "Missing INI file and/or INI section"
				exit 1
		fi

}

# Main
#------
if ["$#" -eq "5"] && [-n "$1"] && [-n "$2"] && [-n "$3"] && [-n "$4"] && [-n "$5"] ; then
		host="$1"
		errorCode="$2"
		status="$3"
		value="$4"
		family="$5"

		Get_INI_Section "$inifile" "$errorCode"
		conf[raisetext]=$(eval echo $(echo ${conf[raisetext]}))
		conf[falltext]=$(eval echo $(echo ${conf[falltext]}))

		nOID="< Trap OID>" # replace here
		if ["${status}" == "WARNING"]
		then
				snmptrap -v 2c -c <Community> localhost:<port> '' ${nOID} ${nOID}.0 s "${conf[raisetext]}" # update here
		else
				snmptrap -v 2c -c <Community> localhost:<port> '' ${nOID} ${nOID}.0 s "${conf[falltext]}" # update here
				
		fi
		ec=$?
		if ["${ec}" == "0"]; then
		   echo "Sent notification \"${conf[info]}\" to Custom endpoint"
		   exit 0
		else
		   echo "Failed to send notification \"${conf[info]}\" to Custom endpoint"
		   exit 1
		fi

else
  echo "Missing arguments and/or arguments are not initialized"
  echo "Usage: $0 <host> <errorcode> <status> <current_value> <family>"
  exit 1
fi

```

## Inside custom sender function

```
/etc/netdata/netdata_alarm_notify.sh "${host}" "${name}" "${status}" "${value}" "${family}"
ec=$?
if ["${ec}" == "0"]; then
	info "Successfully sent notification to Custom endpoint"
else
	error "Failed to send notification to Custom endpoint"
fi

```

Thanks,  
Vasanth

---

<div class="post-metadata">

**Author:** ![Christopher\_Akritid1](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/christopher_akritid1/32/419_2.png) [@Christopher\_Akritid1](https://community.netdata.cloud/u/Christopher_Akritid1)\
**Post date:** [September 20, 2021, 5:08pm UTC](https://community.netdata.cloud/t/support-for-minor-and-major-alarm-status-and-generate-alarm-as-snmp-trap-notification/1616/8 "2021-09-20T17:08:22Z")

</div>

Ok, I took a look at this. A few comments, that might make your life easier.  
There are two things to accomplish here:

1. Send different messages when raising severity vs when decreasing severity.
2. Use `snmptrap` as a notification method.

For the first, you don’t have the logic quite right. Specifically, going to status “CRITICAL” will send `conf[falltext]`. We have `${status}` and `${old_status}`, so you’d want to do the check also the transition from CRITICAL to WARNING and probably have an additional text for that. See for example the code after [netdata/alarm-notify.sh.in at master · netdata/netdata · GitHub](https://github.com/netdata/netdata/blob/master/health/notifications/alarm-notify.sh.in#L2388)  
A more generic way to solve this would be to do a PR to add similar logic with the html sender in the example above around [netdata/alarm-notify.sh.in at master · netdata/netdata · GitHub](https://github.com/netdata/netdata/blob/master/health/notifications/alarm-notify.sh.in#L286) to define the type of state transition. Then any sender could optionally do different things (like define different messages) based on the type of the transition, without checking status and old\_status.

For the latter, I didn’t quite get why you needed the INI file and the wrapper, the custom sender function has every parameter at its disposal and is defined in you own installations’ `health_alarm_notify.conf`. So you could probably achieve everything more simply, if you just did `/etc/netdata/edit-config health_alarm_notify.conf` on one of your nodes and copied the file over to your other deployments.

To summarize and suggest the way forward to our dev team, in case we want to do a PR for this soon:  
Modify `https://github.com/netdata/netdata/blob/master/health/notifications/alarm-notify.sh.in` to:

- Implement generically the logic found in the html sender, that permits different messages to be sent, depending on the state transition.
- Add an snmptrap sender function that sends different messages based on the state transition (we can generalize what happens in the htmlsender and steal from there).
