# Critical - Netdata CPU Leak - had to shut it off on hundreds of nodes

**URL:** <https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062>\
**Category:** Help\
**Created:** [December 8, 2024, 4:35pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062 "2024-12-08T16:35:47Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 8, 2024, 4:35pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/1 "2024-12-08T16:35:47Z")

</div>

Recently Netdata has been going rogue and we had to completely turn it off on hundreds of nodes. It often climbs to using more than half the CPU resources available and the CPU PSI going from 1% to +10%.

 ![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/0/0f996a0a7ed2b08f4b89e38a21e96c38bcc2886d.png)

 ![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/1/1d35b4db2eef63d5c50600ff4543689005dd9c6c.png)

 ![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/1/18f62b46b39a8379230c8ec2397fe2681fc5bcf0.png)

Previous posts did not shed any useable information on how to identify the root cause for this instance.

---

<div class="post-metadata">

**Author:** ![Costa\_Tsaousis](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/costa_tsaousis/32/1851_2.png) [@Costa\_Tsaousis](https://community.netdata.cloud/u/Costa_Tsaousis)\
**Post date:** [December 8, 2024, 4:37pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/2 "2024-12-08T16:37:42Z")

</div>

Which netdata version you use?

---

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 8, 2024, 4:39pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/3 "2024-12-08T16:39:26Z")

</div>

v2.0.0-182-nightly mostly

---

<div class="post-metadata">

**Author:** ![Costa\_Tsaousis](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/costa_tsaousis/32/1851_2.png) [@Costa\_Tsaousis](https://community.netdata.cloud/u/Costa_Tsaousis)\
**Post date:** [December 8, 2024, 5:07pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/4 "2024-12-08T17:07:09Z")

</div>

hm… A couple of days ago we merged a big set of changes. We reworked everything about streaming to actually lower cpu consumption at scale and increase Netdata’s ability to handle a lot more children. We have already found a couple of problems and we fixed them. We are currently testing these changes.

If the tests finish without issues, today we will merge a new nightly with all the fixes.

Are you able to build from source? Do you want to try them somewhere before merging, or you prefer to wait for the next nightly (most likely tomorrow)?

---

<div class="post-metadata">

**Author:** ![Costa\_Tsaousis](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/costa_tsaousis/32/1851_2.png) [@Costa\_Tsaousis](https://community.netdata.cloud/u/Costa_Tsaousis)\
**Post date:** [December 8, 2024, 6:23pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/5 "2024-12-08T18:23:52Z")

</div>

btw, the new version has more replication threads. Prior there was only 1 replication thread (and you had to manually increase it in netdata.conf), but now there are more depending on the number of cores you have.

If you experience increased CPU because you restarted a parent or some children, this is normal. It will finish and settle back to normal.

btw, tests are going well. Most likely tonight the new fixes will be merged.

---

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 9, 2024, 1:04am UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/6 "2024-12-09T01:04:51Z")

</div>

We will be changing to stable and setting strict resource limitations for Netdata Agent. Is there a reason why it doesn’t restrict itself by default?

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [December 9, 2024, 9:38am UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/7 "2024-12-09T09:38:57Z")

</div>

@Slind14 Can you check which treads are using CPU? See [how to do it with htop](https://github.com/netdata/netdata/issues/14892#issuecomment-1505296000).

---

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 9, 2024, 3:40pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/8 "2024-12-09T15:40:53Z")

</div>

That doesn’t seem to work:

 ![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/8/8a427819ab11ec0593437fe77ed682dc38d3c2e5.png)

![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/e/e22b40bfb4662b82669a8470a7d4ede6d34efa68.png)

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [December 9, 2024, 5:51pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/9 "2024-12-09T17:51:25Z")

</div>

~~What doesn’t seem to work? Why did you decide to show ebpf.plugin threads?~~

I apologize for the confusion. I was simply using that `htop` link as an example to demonstrate how to identify CPU-intensive threads. I didn’t intend to ask for a breakdown of `ebpf.plugin` threads. Can you find which threads use a lot of CPU?

---

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 9, 2024, 8:53pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/10 "2024-12-09T20:53:55Z")

</div>

We already capped it and I can’t easily get a reading from before kernel priority adjustments and constraints. Here is the current snapshot where it is limited to idle resources and one core. With this I see it constantly spending +90 CPU% on health while the other threads fluctuate a lot and are averaged out low.

 ![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/1/15f8374bac1f37f26ac6c4d621e9ddb1654f24ca.png)

---

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 9, 2024, 10:48pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/11 "2024-12-09T22:48:31Z")

</div>

After some time:

 ![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/2/21874fa71eb84c11803b885161b5a4c31eaca2ae.png)

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [December 10, 2024, 9:28am UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/12 "2024-12-10T09:28:37Z")

</div>

> limited to idle resources and one cor

Are those screenshots from the Parent or Child instance?

---

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 10, 2024, 9:54pm UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/13 "2024-12-10T21:54:58Z")

</div>

These are on childs without parents. So single instance setups. We have decided to uninstall Netdata for now (we primarily use node\_exporter for these 250 instances anyway), unfortunately we could not reliably constrain/can’t put in the time for it now and it causes too much risk.

 ![image](https://canada1.discourse-cdn.com/flex029/uploads/netdata2/original/2X/8/86110d343a7758a43bffede027af8347b933457d.png)

---

<div class="post-metadata">

**Author:** ![Slind14](https://avatars.discourse-cdn.com/v4/letter/s/b19c9b/32.png) [@Slind14](https://community.netdata.cloud/u/Slind14)\
**Post date:** [December 12, 2024, 12:59am UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/14 "2024-12-12T00:59:13Z")

</div>

Just wanted to let you know that on the test node where we kept it `v2.0.0-205-nightly` is still consuming 5-6 cores consistently.

---

<div class="post-metadata">

**Author:** ![Costa\_Tsaousis](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/costa_tsaousis/32/1851_2.png) [@Costa\_Tsaousis](https://community.netdata.cloud/u/Costa_Tsaousis)\
**Post date:** [December 14, 2024, 3:43am UTC](https://community.netdata.cloud/t/critical-netdata-cpu-leak-had-to-shut-it-off-on-hundreds-of-nodes/6062/15 "2024-12-14T03:43:10Z")

</div>

All pending issues have been fixed. Can you please install the latest nightly? Is it working as it should after this?
