# Is it time for multiple streaming servers?

**URL:** <https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806>\
**Category:** General\
**Tags:** question\
**Created:** [January 28, 2023, 6:35am UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806 "2023-01-28T06:35:31Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [January 28, 2023, 6:35am UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/1 "2023-01-28T06:35:31Z")

</div>

I have a single Netdata server and a fairly powerful system, 32 CPU cores and 256 GB of RAM. It’s setup as a streaming server to “collect” stats from about 150 clients using the stream feature. I’ve followed the guide over at [Netdata daemon | Learn Netdata](https://learn.netdata.cloud/docs/agent/daemon#netdata-process-scheduling-policy) to get the most out of the dedicated system. I am still having some problems.

I’ve started to see (on the server) a “too many open files” error and Netdata is just halted. Not collecting stats and not responding to the clients.

```auto
2023-01-28 06:31:51: netdata ERROR : PLUGIN[diskspace] : PROCFILE: Cannot open file '/proc/self/mountinfo' (errno 24, Too many open files)
2023-01-28 06:31:51: netdata ERROR : PLUGIN[diskspace] : PROCFILE: Cannot open file '/proc/1/mountinfo' (errno 24, Too many open files)
2023-01-28 06:31:58: netdata ERROR : WEB_SERVER[static6] : POLLFD: LISTENER: too many open files - used by this thread 1, max for this thread 42 (similar messages repeated 55466 times in the last 10 secs) (sleeping for 1000 microseconds every time this happens) (errno 24, Too many open files)
2023-01-28 06:32:06: netdata ERROR : PLUGIN[diskspace] : PROCFILE: Cannot open file '/proc/self/mountinfo' (errno 24, Too many open files)
2023-01-28 06:32:06: netdata ERROR : PLUGIN[diskspace] : PROCFILE: Cannot open file '/proc/1/mountinfo' (errno 24, Too many open files)
2023-01-28 06:32:08: netdata ERROR : WEB_SERVER[static4] : POLLFD: LISTENER: too many open files - used by this thread 1, max for this thread 42 (similar messages repeated 55239 times in the last 10 secs) (sleeping for 1000 microseconds every time this happens) (errno 24, Too many open files)
2023-01-28 06:32:18: netdata ERROR : WEB_SERVER[static6] : POLLFD: LISTENER: too many open files - used by this thread 1, max for this thread 42 (similar messages repeated 55291 times in the last 10 secs) (sleeping for 1000 microseconds every time this happens) (errno 24, Too many open files)
2023-01-28 06:32:21: netdata ERROR : PLUGIN[diskspace] : PROCFILE: Cannot open file '/proc/self/mountinfo' (errno 24, Too many open files)
2023-01-28 06:32:21: netdata ERROR : PLUGIN[diskspace] : PROCFILE: Cannot open file '/proc/1/mountinfo' (errno 24, Too many open files)
2023-01-28 06:32:28: netdata ERROR : WEB_SERVER[static5] : POLLFD: LISTENER: too many open files - used by this thread 1, max for this thread 42 (similar messages repeated 55278 times in the last 10 secs) (sleeping for 1000 microseconds every time this happens) (errno 24, Too many open files)

```

Should I look into daisy-chaining multiple instances as described here: [Streaming and replication | Learn Netdata](https://learn.netdata.cloud/docs/agent/streaming#netdata-proxies). Would I be able to setup multiple proxy instances behind a loadbalancer?

The clients are configured with a`stream.conf`:

```auto
[stream]
	enabled = yes 
	destination = netdata.server.domain.tld
	api key = XXXXXXXXXX-XXXXXXXXXXX-XXXX-XXXXXXXXXXXXX

```

And `netdata.conf`

```auto
[global]
	web files owner = root
	web files group = netdata
	bind socket to IP = 0.0.0.0
	memory mode = none
	history = 86400

[health]
	enabled = no

```

The Netdata “server” has this in `stream.conf`

```auto
[XXXXXXXXXX-XXXXXXXXXXX-XXXX-XXXXXXXXXXXXX]
	enabled = yes
	default memory mode = dbengine

```

And this in `netdata.conf`

```auto
[global]
	web files owner = root
	web files group = netdata
	bind socket to IP = 0.0.0.0
	memory mode = dbengine
	page cache size MB = 32768
	dbengine multihost disk space = 81664
	
	# Systemd
	process scheduling policy = keep

```

On the server, the Systemd unit file has:

```auto
CPUSchedulingPolicy=rr
CPUSchedulingPriority=75

```

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [January 28, 2023, 7:37am UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/2 "2023-01-28T07:37:57Z")

</div>

Now the server is just crashing

```auto
free(): double free detected in tcache 2
2023-01-28 07:36:00: apps.plugin ERROR : APPS_READER : Received error on stdin.
EOF found in spawn pipe.
Shutting down spawn server event loop.
Shutting down spawn server loop complete.
2023-01-28 07:36:00: go.d INFO: main[main] received terminated signal (15). Terminating...
2023-01-28 07:36:00: go.d INFO: build[manager] instance is stopped
2023-01-28 07:36:00: go.d INFO: discovery[file manager] instance is stopped
2023-01-28 07:36:00: go.d INFO: run[manager] instance is stopped
2023-01-28 07:36:00: go.d INFO: discovery[manager] instance is stopped

```

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [January 31, 2023, 4:46pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/3 "2023-01-31T16:46:55Z")

</div>

Hi, @tuaris. What is your Netdata version (both parent and children instances)?

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [February 2, 2023, 4:51pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/4 "2023-02-02T16:51:19Z")

</div>

The majority of the child instances are on 1.37.1. There is probably 3 or 4 on a slightly lesser version, but still on the 1.3x series.

Parent is 1.37.1:

```auto
Every second, Netdata collects 4,250 metrics on xxxxxxxx, presents them in 633 charts and monitors them with 57 alarms.

netdata
v1.37.1

```

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [February 3, 2023, 9:35am UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/5 "2023-02-03T09:35:03Z")

</div>

We need an exact setup (parent version, children version) to try to replicate the problem. I suggest you wait for v1.38.0 (which will happen next week), update all your instances, and see if it helps. Additionally, if you get “too many open files” I suggest you to bump the limit for the parent instance.

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [February 15, 2023, 4:03pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/6 "2023-02-15T16:03:53Z")

</div>

Ah, no problem. I use Salt stack and am able to query all systems easily.

There are exactly 130 child instances at the moment of querying for version. All but one instance is using `v1.37.1`. There is a single instance That was provisioned only a few days ago that’s running `v1.38.0`.

The parent is running `v1.37.1`

I also noticed that the parent will occasionally cease collecting stats and start throwing alerts like “ram available = 0%” or “mysql galera cluster status = 0 status” even though the child instances are fine.

I’ll upgrade to `v1.38.0` and report back on how things go.

EDIT: looks like 1.38.1 is available… updating to that instead.

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [February 15, 2023, 4:18pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/7 "2023-02-15T16:18:52Z")

</div>

Oh, and this is all on Ubuntu 20.04.5 LTS on AWS EC2 if it matters.

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [March 6, 2023, 7:56pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/9 "2023-03-06T19:56:06Z")

</div>

Not seeing much improvements with v1.38.1. The parent server is locking up, but the utilization is very small. I am not getting any too open files errors this time and am unsure of where the problem is.

```auto
top - 18:14:08 up 11 days, 3:08, 1 user, load average: 7.30, 9.16, 8.86
Tasks: 377 total, 1 running, 376 sleeping, 0 stopped, 0 zombie
%Cpu0 : 6.1 us, 1.9 sy, 0.0 ni, 70.9 id, 17.5 wa, 0.0 hi, 3.6 si, 0.0 st
%Cpu1 : 5.6 us, 0.7 sy, 0.0 ni, 86.0 id, 7.0 wa, 0.0 hi, 0.7 si, 0.0 st
%Cpu2 : 7.2 us, 2.3 sy, 0.0 ni, 73.8 id, 16.4 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu3 : 6.3 us, 1.7 sy, 0.0 ni, 72.6 id, 16.8 wa, 0.0 hi, 2.6 si, 0.0 st
%Cpu4 : 7.6 us, 1.3 sy, 0.0 ni, 83.5 id, 7.3 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu5 : 6.6 us, 1.3 sy, 0.0 ni, 82.7 id, 9.0 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu6 : 6.0 us, 1.7 sy, 0.0 ni, 71.5 id, 20.8 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu7 : 6.3 us, 1.3 sy, 0.0 ni, 83.7 id, 7.3 wa, 0.0 hi, 1.3 si, 0.0 st
%Cpu8 : 7.7 us, 1.7 sy, 0.0 ni, 71.9 id, 18.7 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu9 : 6.7 us, 2.3 sy, 0.0 ni, 73.2 id, 15.4 wa, 0.0 hi, 2.3 si, 0.0 st
%Cpu10 : 7.6 us, 1.7 sy, 0.0 ni, 78.2 id, 12.5 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu11 : 6.0 us, 1.3 sy, 0.0 ni, 62.2 id, 30.4 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu12 : 6.1 us, 4.4 sy, 0.0 ni, 55.2 id, 34.3 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu13 : 5.7 us, 1.7 sy, 0.0 ni, 71.7 id, 20.7 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu14 : 5.6 us, 3.3 sy, 0.0 ni, 49.0 id, 41.1 wa, 0.0 hi, 1.0 si, 0.0 st
%Cpu15 : 6.0 us, 2.0 sy, 0.0 ni, 69.5 id, 22.2 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu16 : 5.7 us, 2.0 sy, 0.0 ni, 72.3 id, 17.0 wa, 0.0 hi, 3.0 si, 0.0 st
%Cpu17 : 7.8 us, 1.6 sy, 0.0 ni, 67.0 id, 23.5 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu18 : 7.3 us, 1.0 sy, 0.0 ni, 85.3 id, 4.7 wa, 0.0 hi, 1.7 si, 0.0 st
%Cpu19 : 5.7 us, 1.7 sy, 0.0 ni, 80.1 id, 12.5 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu20 : 7.6 us, 1.0 sy, 0.0 ni, 86.8 id, 4.6 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu21 : 5.7 us, 1.7 sy, 0.0 ni, 66.7 id, 23.3 wa, 0.0 hi, 2.7 si, 0.0 st
%Cpu22 : 5.0 us, 1.7 sy, 0.0 ni, 69.6 id, 23.4 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu23 : 6.7 us, 0.7 sy, 0.0 ni, 85.7 id, 6.7 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu24 : 8.3 us, 1.0 sy, 0.0 ni, 85.4 id, 5.3 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu25 : 7.3 us, 1.3 sy, 0.0 ni, 79.7 id, 11.7 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu26 : 6.7 us, 1.7 sy, 0.0 ni, 83.0 id, 7.7 wa, 0.0 hi, 1.0 si, 0.0 st
%Cpu27 : 6.1 us, 1.4 sy, 0.0 ni, 80.7 id, 11.8 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu28 : 7.7 us, 2.0 sy, 0.0 ni, 65.2 id, 23.1 wa, 0.0 hi, 2.0 si, 0.0 st
%Cpu29 : 6.3 us, 2.0 sy, 0.0 ni, 57.3 id, 34.3 wa, 0.0 hi, 0.0 si, 0.0 st
%Cpu30 : 5.7 us, 1.7 sy, 0.0 ni, 66.7 id, 25.7 wa, 0.0 hi, 0.3 si, 0.0 st
%Cpu31 : 6.0 us, 1.7 sy, 0.0 ni, 79.7 id, 12.7 wa, 0.0 hi, 0.0 si, 0.0 st
MiB Mem : 255670.7 total, 1264.8 free, 253165.8 used, 1240.0 buff/cache
MiB Swap: 0.0 total, 0.0 free, 0.0 used. 1070.4 avail Mem 

    PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND                                                                                                   
    895 netdata -76 0 255.2g 245.6g 1.0g S 262.5 98.4 55805:55 netdata                                                                                                   
   2088 netdata -76 0 135832 23272 2112 S 2.3 0.0 365:35.10 apps.plugin                                                                                                                     

```

Netdata is completely locked up and not responding to requests:

```auto
# time curl -I http://localhost:19999
^C

real	99m9.004s
user	0m0.126s
sys	0m0.126s

```

The volume IO isn’t saturated, but there does seem to be a lot of activity.

```auto
Total DISK READ: 6.10 M/s | Total DISK WRITE: 389.61 K/s
Current DISK READ: 6.12 M/s | Current DISK WRITE: 285.92 K/s
    TID PRIO USER DISK READ DISK WRITE SWAPIN IO> COMMAND                                                                                                         
   1598 rt/4 netdata 182.24 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1600 rt/4 netdata 257.64 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1601 rt/4 netdata 0.00 B/s 12.57 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1602 rt/4 netdata 37.70 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1603 rt/4 netdata 0.00 B/s 3.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1604 rt/4 netdata 18.85 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1607 rt/4 netdata 414.75 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1608 rt/4 netdata 0.00 B/s 34.56 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1609 rt/4 netdata 50.27 K/s 31.42 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1610 rt/4 netdata 53.41 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1611 rt/4 netdata 433.60 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1612 rt/4 netdata 185.38 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1613 rt/4 netdata 311.06 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1614 rt/4 netdata 0.00 B/s 40.85 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1615 rt/4 netdata 323.63 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1616 rt/4 netdata 153.96 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1617 rt/4 netdata 0.00 B/s 25.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1618 rt/4 netdata 0.00 B/s 25.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1620 rt/4 netdata 326.77 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1621 rt/4 netdata 270.21 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1622 rt/4 netdata 3.14 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1623 rt/4 netdata 0.00 B/s 28.28 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1624 rt/4 netdata 166.53 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1625 rt/4 netdata 0.00 B/s 3.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1627 rt/4 netdata 91.12 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1628 rt/4 netdata 543.57 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1630 rt/4 netdata 160.24 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1632 rt/4 netdata 9.43 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1633 rt/4 netdata 172.81 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1635 rt/4 netdata 216.80 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1636 rt/4 netdata 235.65 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1637 rt/4 netdata 216.80 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1638 rt/4 netdata 0.00 B/s 84.83 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1639 rt/4 netdata 0.00 B/s 3.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1641 rt/4 netdata 113.11 K/s 3.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1642 rt/4 netdata 204.23 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1643 rt/4 netdata 21.99 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1644 rt/4 netdata 194.80 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1645 rt/4 netdata 0.00 B/s 3.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1647 rt/4 netdata 113.11 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1648 rt/4 netdata 219.94 K/s 0.00 B/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1649 rt/4 netdata 0.00 B/s 3.14 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]
   1651 rt/4 netdata 0.00 B/s 12.57 K/s ?unavailable? netdata -D -P /var/run/netdata/netdata.pid [LIBUV_WORKER]

```

Underlying volume’s performance does not appear to be the problem either:

```auto
# dd if=/dev/zero of=/var/cache/netdata/testfile.bin bs=1M count=1k conv=fdatasync; rm -f /var/cache/netdata/testfile.bin
1024+0 records in
1024+0 records out
1073741824 bytes (1.1 GB, 1.0 GiB) copied, 7.39761 s, 145 MB/s

```

145 active children (all running 1.38.1) connected to it as of this moment

```auto
# ss -4tn src :19999 | awk '{print $5}' | awk -F':' '{print $1}' | sort | uniq | wc -l
147

```

Looks like it’s really just doing nothing?

```auto
# strace -p895
strace: Process 895 attached
pause(

```

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [March 6, 2023, 10:26pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/10 "2023-03-06T22:26:29Z")

</div>

Interestingly enough… alerts still seem to be functioning (thank goodness).

---

<div class="post-metadata">

**Author:** ![ilyam8](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/ilyam8/32/134_2.png) [@ilyam8](https://community.netdata.cloud/u/ilyam8)\
**Post date:** [March 13, 2023, 11:18am UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/11 "2023-03-13T11:18:41Z")

</div>

Hi, @tuaris. What do you mean by “The parent server is locking up”?

- Can you disable ML and see if that helps
  - open `netdata.conf`
  - find `[ml]`
  - uncomment `enabled` and set it to `no`
  - restart netdata service

* * *

> [@tuaris](#):
>
> stats from about 150 clients using the stream feature

Do you have K8s clients? Or any nodes with ephemeral containers (containers that frequently come and go)?

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [March 17, 2023, 5:06am UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/12 "2023-03-17T05:06:00Z")

</div>

> [@ilyam8](#):
>
> Hi, @tuaris. What do you mean by “The parent server is locking up”?

For example if I try to `curl` it, it sits there forever.

> [@ilyam8](#):
>
> Can you disable ML and see if that helps

Did that in late January 2023 shortly after creating this thread. There was a huge improvement, but still see the issues like I mentioned recently.

> [@ilyam8](#):
>
> Do you have K8s clients? Or any nodes with ephemeral containers (containers that frequently come and go)?

No, these are all EC2 instances. Although we provision new ones and de-provision older ones from time to time, they don’t come and go as often as containers would.

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [April 19, 2023, 6:01pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/13 "2023-04-19T18:01:00Z")

</div>

The pattern so far appears to be roughly 20-30 days, where the Netdata parent slows down to a crawl and stops collecting and responding. A restart of the process seems to solve things until the next time.

I think for now as a workaround I will just need to setup a cron to restart the Netdata daemon every night, just to avoid issues.

Perhaps I have hit the upper limits of what a single Netdata parent can handle.

---

<div class="post-metadata">

**Author:** ![tuaris](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/tuaris/32/1084_2.png) [@tuaris](https://community.netdata.cloud/u/tuaris)\
**Post date:** [June 28, 2023, 4:34pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/14 "2023-06-28T16:34:47Z")

</div>

I understand what the issue is. The amount of data that Prometheus needs to scrape from this one Netdata parent is massive. It takes about 140 seconds from query to end of response. I will probably need to move to Netdata pushing instead of having Prometheus pull from Netdata.

In the meantime I’ll increase my scrape interval to 5 minutes and the time out to 200 seconds. I’m not sure if there is actually anything Netdata devs can do or recommend to help other than that.

---

<div class="post-metadata">

**Author:** ![andrewm4894](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/andrewm4894/32/125_2.png) [@andrewm4894](https://community.netdata.cloud/u/andrewm4894)\
**Post date:** [June 30, 2023, 9:03pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/15 "2023-06-30T21:03:43Z")

</div>

Interesting.

I guess try this and see. [Export metrics to Prometheus remote write providers | Learn Netdata](https://learn.netdata.cloud/docs/exporting/prometheus/remote-write)

Seems like an interesting set up so keep us posted on how you get on.

---

<div class="post-metadata">

**Author:** ![andrewm4894](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/andrewm4894/32/125_2.png) [@andrewm4894](https://community.netdata.cloud/u/andrewm4894)\
**Post date:** [June 30, 2023, 9:07pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/16 "2023-06-30T21:07:28Z")

</div>

I wonder also, if you could pass some sort of filter to /allmetrics then maybe you could set up a few different Prometheus jobs to scrape a subset each time and just time or sequence them to occur not at same time then I wonder if that could help.

I don’t think there is a way to filter /allmetrics based on some sort of query etc but if was a viable solution it could maybe make a useful feature request.

Not sure myself, need to think a bit more and maybe see what some others on agent team think.

---

<div class="post-metadata">

**Author:** ![andrewm4894](https://yyz1.discourse-cdn.com/flex029/user_avatar/community.netdata.cloud/andrewm4894/32/125_2.png) [@andrewm4894](https://community.netdata.cloud/u/andrewm4894)\
**Post date:** [June 30, 2023, 9:09pm UTC](https://community.netdata.cloud/t/is-it-time-for-multiple-streaming-servers/3806/17 "2023-06-30T21:09:07Z")

</div>

Oh wait seems like you can pass a filter to allmetrics so I wonder if you could just split it into two separate Prometheus scrape jobs in some.way using a filter or something like that.

[https://editor.swagger.io/?url=https://raw.githubusercontent.com/netdata/netdata/master/web/api/netdata-swagger.yaml](https://editor.swagger.io/?url=https://raw.githubusercontent.com/netdata/netdata/master/web/api/netdata-swagger.yaml)
