# What's your monitoring setup?

**URL:** <https://community.ntppool.org/t/whats-your-monitoring-setup/3219>\
**Category:** Server operators\
**Created:** [February 1, 2024, 1:10pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219 "2024-02-01T13:10:07Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Badeand](https://avatars.discourse-cdn.com/v4/letter/b/c2a13f/32.png) [@Badeand](https://community.ntppool.org/u/Badeand)\
**Post date:** [February 1, 2024, 1:10pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/1 "2024-02-01T13:10:07Z")

</div>

I’ve seen some people here having graphs to show their server traffic and resource usage, presumably Grafana.  
I’m curious, what are you using, how have you set it up, and what exactly are you monitoring? Where is the data collected (Prometheus, InfluxDB, Graphite?), and by what method? Basically, what’s your setup?

Looking into it, I’m a little intimidated by the number of parts involved in such a setup. Might be a fun project though.

---

<div class="post-metadata">

**Author:** ![ronv42](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/ronv42/32/1466_2.png) [@ronv42](https://community.ntppool.org/u/ronv42)\
**Post date:** [February 2, 2024, 11:46am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/2 "2024-02-02T11:46:27Z")

</div>

I am using Telegraf on my time servers to send both NT and system metrics into Influxdb. Then I am using Grafana for dashboards.

---

<div class="post-metadata">

**Author:** ![bjh21](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/bjh21/32/1181_2.png) [@bjh21](https://community.ntppool.org/u/bjh21)\
**Post date:** [February 2, 2024, 3:24pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/3 "2024-02-02T15:24:20Z")

</div>

In my bit of the University of Cambridge, we generally run CollectD on each host and they all send metrics to a central Graphite server. Then we have a Grafana system that can display data from that Graphite server. I take advantage of this for my NTP server metrics.

On the CollectD side, there are four plugins that I use that are relevant for NTP:

- chrony, for timekeeping data from chrony
- [iptables](https://github.com/collectd/collectd/wiki/Plugin-IPTables), for traffic levels
- [processes](https://github.com/collectd/collectd/wiki/Plugin-Processes), for chrony’s CPU usage
- [ipmi](https://github.com/collectd/collectd/wiki/Plugin-IPMI), for hardware temperature sensors

The “iptables” plugin is combined with IPTables rules that specifically match NTP traffic so I can separate that from everything else the servers do. The temperature sensors are useful for telling when changes in clock frequency are caused by problems with the cooling system.

I then have a single Grafana dashboard that puts all this information together for the NTP servers under my control.

---

<div class="post-metadata">

**Author:** ![Knot3n](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/knot3n/32/407_2.png) [@Knot3n](https://community.ntppool.org/u/Knot3n)\
**Post date:** [February 3, 2024, 3:54pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/4 "2024-02-03T15:54:22Z")

</div>

Im using netdata and zabbix for monitoring stuff.

---

<div class="post-metadata">

**Author:** ![aaxvig](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@aaxvig](https://community.ntppool.org/u/aaxvig)\
**Post date:** [February 6, 2024, 2:35am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/5 "2024-02-06T02:35:01Z")

</div>

I use Zabbix to monitor standard CPU, memory, disk, and network traffic like any other host. For specific NTP metrics I started with [this Zabbix template](https://github.com/zabbix/community-templates/tree/main/Applications/NTP/template_chrony_accuracy/6.0). That is only chrony client stuff so I added server-type things.

- NTP packets received: `system.run[sudo chronyc serverstats | grep "NTP packets received" | awk '{print $5}']`
- Client count: `system.run[sudo chronyc -c -n clients | wc -l]`

A few other ones following that pattern. It makes fine graphs.

 ![chart](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/d/d4efef829d93af2c576108d5d2f7d252bf4a8149.png)

That is with my speed set to 25mbps, BTW.

---

<div class="post-metadata">

**Author:** ![kenyon](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/kenyon/32/45_2.png) [@kenyon](https://community.ntppool.org/u/kenyon)\
**Post date:** [February 7, 2024, 12:09am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/6 "2024-02-07T00:09:59Z")

</div>

Using [munin](https://munin-monitoring.org/) with the chrony\_drift and chrony\_status plugins from here:

> <https://github.com/munin-monitoring/contrib/tree/master/plugins/chrony>
>
> //github.com/munin-monitoring/contrib/tree/master/plugins/chrony

The resulting graphs 📈 can be seen here:

[https://beta.kenyonralph.com/munin/time-day.html](https://beta.kenyonralph.com/munin/time-day.html)

---

<div class="post-metadata">

**Author:** ![ask](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/ask/32/907_2.png) [@ask](https://community.ntppool.org/u/ask)\
**Post date:** [February 7, 2024, 7:55am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/7 "2024-02-07T07:55:49Z")

</div>

For the NTP Pool Project itself it’s a lot of Grafana stuff plus custom software for data from the DNS servers and monitors.

- Prometheus (with various exporters) for scraping / collecting data, mostly over HTTP though the monitors are “scraped” over MQTT. The DNS servers are scraped with mTLS (each DNS server gets a private certificate for this).
- Longer term metrics storage in Mimir (from Grafana).
- Application logs go via promtail into Loki (also Grafana).
- System logs from the DNS servers are sent via vector (with mTLS) into the main cluster; also going into Loki.
- DNS server (query) logs are serialized in Avro files and then sent to the cluster with a custom program, then to Kafka and from there into ClickHouse (and summarized; only about a day of queries are kept).
- Alerts are made with a combination of grafana and prometheus + alertmanager (and a few Loki rules)
- Tracing (OpenTelemetry) is collected via various otel-collector instances and sent to Tempo (another Grafana product).

I haven’t figured out to count queries from the (single) NTP server I run though, which I guess is what the thread is about… 😃

---

<div class="post-metadata">

**Author:** ![Sebhoster](https://avatars.discourse-cdn.com/v4/letter/s/8edcca/32.png) [@Sebhoster](https://community.ntppool.org/u/Sebhoster)\
**Post date:** [February 7, 2024, 10:25am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/8 "2024-02-07T10:25:33Z")

</div>

I use Prometheus and Grafana. Data is collected with [node exporter](https://github.com/prometheus/node_exporter) for general metrics like CPU load, [chrony exporter](https://github.com/SuperQ/chrony_exporter) (with the unfinished serverstats PR) for chrony metrics and a small python script that processes the chrony client statistics and also exposes them for Prometheus.

The chrony portion of my dashboard: The yellow lines indicate dropped packets and clients that had at least one of their requests dropped

 ![grafik](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/e/efe7e844d967f23bfaf46288d235c9a379b6dfad.png)

---

<div class="post-metadata">

**Author:** ![Badeand](https://avatars.discourse-cdn.com/v4/letter/b/c2a13f/32.png) [@Badeand](https://community.ntppool.org/u/Badeand)\
**Post date:** [February 7, 2024, 1:40pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/9 "2024-02-07T13:40:02Z")

</div>

> [@ask](#):
>
> I haven’t figured out to count queries from the (single) NTP server I run though, which I guess is what the thread is about… 😃

I think it’s both relevant and interesting either way, so thanks for sharing ^^

---

<div class="post-metadata">

**Author:** ![Bas](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/bas/32/465_2.png) [@Bas](https://community.ntppool.org/u/Bas)\
**Post date:** [February 7, 2024, 3:29pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/10 "2024-02-07T15:29:07Z")

</div>

I use nothing, but my servers have limited DNS-range because they are set to 512K.  
However, in case of trouble I use nethogs and check chronyc clients for abusive polling.  
If there is one, I will block the IP via the firewall, so far just a few, not worth counting 🙂

---

<div class="post-metadata">

**Author:** ![NTPman](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/ntpman/32/450_2.png) [@NTPman](https://community.ntppool.org/u/NTPman)\
**Post date:** [February 7, 2024, 7:03pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/11 "2024-02-07T19:03:56Z")

</div>

I use `MRTG`:

 ![image](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/1/195c6df0827b865328c478fbdeccd96f50744e8d.png)

---

<div class="post-metadata">

**Author:** ![paulgear](https://avatars.discourse-cdn.com/v4/letter/p/ad7895/32.png) [@paulgear](https://community.ntppool.org/u/paulgear)\
**Post date:** [April 4, 2024, 11:20pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/12 "2024-04-04T23:20:49Z")

</div>

Late to the party, but I use Grafana, InfluxDB, telegraf, and my own script [NTPmon](https://github.com/paulgear/ntpmon/). This handles chrony and ntpd (both traditional and NTPsec versions) with the same monitoring infrastructure.

One of my pool hosts runs [ntpd-rs](https://github.com/pendulum-project/ntpd-rs), which provides its own telemetry endpoint, so that is scraped periodically by telegraf and graphed using a different Grafana dashboard.

[One of my recent blog posts](https://www.libertysys.com.au/2024/04/aws-microsecond-accurate-time-first-look/) links to a public Grafana dashboard where you can see the metrics I collect, and [another](https://www.libertysys.com.au/2023/12/an-update-on-ntpmon/) explains why I chose InfluxDB & telegraf for NTP monitoring.

---

<div class="post-metadata">

**Author:** ![NTP-LINUX](https://avatars.discourse-cdn.com/v4/letter/n/dfb087/32.png) [@NTP-LINUX](https://community.ntppool.org/u/NTP-LINUX)\
**Post date:** [June 1, 2024, 4:33am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/13 "2024-06-01T04:33:32Z")

</div>

I use shell + rrdtool for embedded device LuckFox pico with 128Mb storage  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/1/1a1266237e992684b8b77ef5cb7eedaaa6d1b87f.png)

---

<div class="post-metadata">

**Author:** ![PoolMUC](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/poolmuc/32/1473_2.png) [@PoolMUC](https://community.ntppool.org/u/PoolMUC)\
**Post date:** [June 17, 2024, 6:30pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/14 "2024-06-17T18:30:09Z")

</div>

I don’t monitor the timekeeping performance of my various servers beyond manually using the tools shipped alongside the daemon. I.e., ntpq for my ntpd classic and NTPsec instances, and chronyc for chronyd. And the Pool’s monitoring system of course.

I have basic traffic monitoring in place for my current main server in the pool, using darkstat and RIPE Atlas.

 ![Screenshot_20240613-182026_Ecosia](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/3/357816f9ed498269e4990d45fb07cc3b73b7ce86.jpeg)

I have two darkstat instances, one for IPv4, one for IPv6. IPv6 is somewhat boring, just 616 kbit/s or so with a 3 Gbit setting. IPv4 is more interesting, with something in the tool or the system being too slow to handle the pcap-based packet capture so that the max bitrate shown is typically clipped at around 8 Mbit/s. The counters for “captured” and “dropped” packets are from pcap, which has them as u\_int, so they wrap around rather often.

 ![Screenshot_20240613-191556_Ecosia](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/2/2e91b08f14a96afc9131bff6fccada3d35a99208.jpeg)

This is the overall interface traffic, so not only NTP, but other stuff as well. But since this server is primarily intended for NTP service, any non-NTP traffic (ICMP pings as discussed in a recent separate thread, HTTP(S) and other probing/scanning, me viewing darkstat, …) is basically negligible in comparison to the 2 Mbit/s and upwards NTP traffic.

My hypothesis is that the rather large variations throughout the day are due to some other servers in the same zone phasing in and out of the pool rather often due to high load (there are two monitors in the same zone, and local traffic seems to have a typical RTT of 1 or 2 ms, so the typical packet loss seen in other cases is rather unlikely to cause servers to drop here). (Very short bursts could also be the occasional software or GeoDB update download.)

With a 6 Mbit setting for IPv4, I have seen the occasional peak of up to 18 Mbit/s (which is roughly where the cloud provider’s DDoS detection typically kicks in), but otherwise, throughput typically varies between about 2 Mbit/s and 12 Mbit/s.

---

<div class="post-metadata">

**Author:** ![marco.davids](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/marco.davids/32/767_2.png) [@marco.davids](https://community.ntppool.org/u/marco.davids)\
**Post date:** [June 17, 2024, 7:12pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/15 "2024-06-17T19:12:29Z")

</div>

> [@PoolMUC](#):
>
> IPv6 is somewhat boring, just 616 kbit/s or so

As an IPv6 evangelist, I feel an urgent need to say that it is not only due to the low number of IPv6 clients (in fact, there are quite a few nowadays), but also because @ask simply refuses to add more AAAA records besides the one at [2.pool.ntp.org](http://2.pool.ntp.org) that we have had since IPv6 Launch Day 2012 or so. 😐

---

<div class="post-metadata">

**Author:** ![PoolMUC](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/poolmuc/32/1473_2.png) [@PoolMUC](https://community.ntppool.org/u/PoolMUC)\
**Post date:** [June 17, 2024, 7:27pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/16 "2024-06-17T19:27:39Z")

</div>

> [@marco.davids](#):
>
> there are quite a few nowadays

Though I have the impression that for some reason, despite the huge number of clients, uptake of IPv6 is a bit lower in large parts of Asia (where this particular server is located) than, e.g., in Europe (with a few exceptions such as India and North Korea), according to [this source](https://stats.labs.apnic.net/ipv6/).

Actually, taking a second look, Europe overall isn’t that great, either, in some parts…

---

<div class="post-metadata">

**Author:** ![marco.davids](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/marco.davids/32/767_2.png) [@marco.davids](https://community.ntppool.org/u/marco.davids)\
**Post date:** [June 17, 2024, 8:04pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/17 "2024-06-17T20:04:30Z")

</div>

> [@PoolMUC](#):
>
> according to [this source](https://stats.labs.apnic.net/ipv6/)

If the traffic ratio between IPv4 and IPv6 that I see on my NTP servers would match these stats from APNIC, I would be a very happy man. But this can at best only happen if all [0123].pool.ntp.org FQDNs receive an AAAA-record.

But I’ll stop now - because this thread was not about IPv6.

---

<div class="post-metadata">

**Author:** ![gunnar](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/gunnar/32/689_2.png) [@gunnar](https://community.ntppool.org/u/gunnar)\
**Post date:** [June 24, 2024, 10:14am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/18 "2024-06-24T10:14:01Z")

</div>

> [@PoolMUC](#):
>
> Actually, taking a second look, Europe overall isn’t that great, either, in some parts…

From my vantage point in Germany, no…

```auto
Applied filter: 'udp and port 123'
IPv4 bytes: 410,072,543 (94.385 %), IPv6 bytes: 29,811,210 (5.615 %).
Starttime: Mon Jun 24 11:57:34 2024
  Endtime: Mon Jun 24 12:05:02 2024
Total number of packets: 4,827,024.
IPv4: 4,556,013 packets, IPv6: 271,011 packets.

```

The uptake in IPv6 could be higher in my opinion, most of the IPv6 requests are from the known DSlight / IPv6 only internet providers ☹ depite what the map says…

---

<div class="post-metadata">

**Author:** ![folkertvanheusden](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/folkertvanheusden/32/1420_2.png) [@folkertvanheusden](https://community.ntppool.org/u/folkertvanheusden)\
**Post date:** [July 21, 2024, 12:56pm UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/19 "2024-07-21T12:56:51Z")

</div>

F.w.i.w.: I’m creating my own monitoring dashboard for gpsd + ntpsec.

> **[GitHub - folkertvanheusden/timeweb: Show NTPSEC & GPSd statistics in a...](https://github.com/folkertvanheusden/timeweb)**
>
> Show NTPSEC & GPSd statistics in a web-interface. Contribute to folkertvanheusden/timeweb development by creating an account on GitHub.

and a demo at [http://gateway.vanheusden.com:5000/](http://gateway.vanheusden.com:5000/)  
It’s a fairly new project (2 days when I write this) so there may be bugs and a lot to wish for 🙂 Also at the time of writing, the gps is located behind a coated window (recently moved) so reception is spotty (for the demo).

---

<div class="post-metadata">

**Author:** ![ChrisJ](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/chrisj/32/1822_2.png) [@ChrisJ](https://community.ntppool.org/u/ChrisJ)\
**Post date:** [August 3, 2025, 10:15am UTC](https://community.ntppool.org/t/whats-your-monitoring-setup/3219/20 "2025-08-03T10:15:55Z")

</div>

My setup is currently fairly simple. I capture hourly stats on number of clients (IPv4 and IPv6 tracked separately) and traffic (again IPv4 and IPv6 tracked separately) from my server (via ntpq) and store this in a relational database. I can then query and extract the data in various ways including graphing it to help me spot problems and trends.

I also have separate monitoring to check that my server is alive and responding on all its IP addresses and to periodically check its scores.

[Next page](https://community.ntppool.org/t/whats-your-monitoring-setup/3219.md?page=2)
