# How to "soften" slow web service alarms

**URL:** <https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009>\
**Category:** Help\
**Tags:** agent\
**Created:** [March 2, 2021, 3:54pm UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009 "2021-03-02T15:54:04Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![jurgenhaas](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/jurgenhaas/32/126_2.png) [@jurgenhaas](https://community.netdata.cloud/u/jurgenhaas)\
**Post date:** [March 2, 2021, 3:54pm UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/1 "2021-03-02T15:54:04Z")

</div>

On some of our high performance web hosts we experience too many web\_service\_slow alarms and they often turn critical even for an average response time of 90ms or less. The health template is configured like this:

```
template: web_service_slow
families: *
      on: httpcheck.responsetime
  lookup: average -5m unaligned of time
   units: ms
   every: 10s
    warn: ($this > ($1h_web_service_response_time * 4) )
    crit: ($this > ($1h_web_service_response_time * 6) )
    info: average response time over the last 5 minutes, compared to the average over the last hour
   delay: down 5m multiplier 1.5 max 1h
      to: webmaster

```

We analysed the problem and it seems that we have regular response times between 0.01 and 0.05 most of the time. Every now and then, there is a single request that takes a couple of hundred milliseconds to respond, but that immediately “ruins” the average value and triggers the alarm.

Not sure if this is the proper thinking but we came up with the idea to ignore the lowest and highest value when calculating the average, that may should already provide a much more realistic picture in this context.

Is that the right way to think about it and if so, can this be done already in Netdata? Or is there better ways to approach this?

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [March 2, 2021, 7:04pm UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/2 "2021-03-02T19:04:08Z")

</div>

Hi @jurgenhaas

> ignore the lowest and highest value

Sounds like median. Unfortunately there is no such [lookup method](https://learn.netdata.cloud/docs/agent/health/reference#alarm-line-lookup) (available: average, min, max, sum).

Your use case is clear, i agree that `median` method would be a nice addition 🤔

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [March 2, 2021, 7:05pm UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/3 "2021-03-02T19:05:08Z")

</div>

@dim08 @Stelios_Fragkakis take a look

---

<div class="post-metadata">

**Author:** ![Christopher\_Akritid1](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/christopher_akritid1/32/419_2.png) [@Christopher\_Akritid1](https://community.netdata.cloud/u/Christopher_Akritid1)\
**Post date:** [March 2, 2021, 7:47pm UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/4 "2021-03-02T19:47:02Z")

</div>

Percentile charts would probably be more appropriate for latencies. Then we could just have normal alerts on top of those, instead of the ones with the avg response time. We have recently dealt with histograms that we convert to such charts, what do you think about adding some of those here @ilyam8 ?

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [March 3, 2021, 5:05pm UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/5 "2021-03-03T17:05:50Z")

</div>

That is true, i agree. I like that `median` lookup method idea, because [i’ve dealt with our alarms recently](https://github.com/netdata/netdata/pull/10688) and i had exactly same thought - `average` prone to false positives, especially when we lookup 10+minutes. Any spike _ruins_ the picture.

And not only _latencies_ would benefit from changing `average` to `median`. I guess it wasn’t implemented, because using `average` is much cheaper in terms of cpu usage.

---

<div class="post-metadata">

**Author:** ![jurgenhaas](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/jurgenhaas/32/126_2.png) [@jurgenhaas](https://community.netdata.cloud/u/jurgenhaas)\
**Post date:** [March 8, 2021, 9:14am UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/6 "2021-03-08T09:14:00Z")

</div>

OK, sounds like this is something for the product and my question is, what can we do to get this onto the roadmap?

---

<div class="post-metadata">

**Author:** ![OdysLam](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/odyslam/32/98_2.png) [@OdysLam](https://community.netdata.cloud/u/OdysLam)\
**Post date:** [March 8, 2021, 10:19am UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/7 "2021-03-08T10:19:16Z")

</div>

Hey @jurgenhaas,

The most immediate thing you can do is to contribute this to the Netdata Agent. Although this would require some C knowledge, we are more than eager to help you! cc @vlvkobal

The other route would be to create a topic on #feature-requests:agent-fr and wait for the product team to pick it up and prioritize it accordingly depending on our internal roadmap and the community. Of course, if we see a large interest by the community, it will be prioritized higher.

P.S Have a great week!

---

<div class="post-metadata">

**Author:** ![jurgenhaas](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/jurgenhaas/32/126_2.png) [@jurgenhaas](https://community.netdata.cloud/u/jurgenhaas)\
**Post date:** [March 8, 2021, 10:25am UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/8 "2021-03-08T10:25:53Z")

</div>

Thanks @OdysLam for your reply. Unfortunately we have no C knowledge whatsoever in our organisations, so we have to go for option 2 and do some promotion too so that people hopefully show their interest in this for you to better prioritize it then.

Have a great week too, thanks.

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [December 7, 2021, 11:20am UTC](https://community.netdata.cloud/t/how-to-soften-slow-web-service-alarms/1009/9 "2021-12-07T11:20:11Z")

</div>

> [@ilyam8](#):
>
> Sounds like median. Unfortunately there is no such [lookup method](https://learn.netdata.cloud/docs/agent/health/reference#alarm-line-lookup) (available: average, min, max, sum).

Actually seems there is - see [Median | Learn Netdata](https://learn.netdata.cloud/docs/agent/web/api/queries/median).
