# \[chrony\] PPS or system clock glitch/spike

**URL:** <https://community.ntppool.org/t/chrony-pps-or-system-clock-glitch-spike/4770>\
**Category:** Server operators\
**Created:** [October 2, 2026, 8:21pm UTC](https://community.ntppool.org/t/chrony-pps-or-system-clock-glitch-spike/4770 "2026-10-02T20:21:30Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![fox](https://avatars.discourse-cdn.com/v4/letter/f/c5a1d2/32.png) [@fox](https://community.ntppool.org/u/fox)\
**Post date:** [October 2, 2026, 8:21pm UTC](https://community.ntppool.org/t/chrony-pps-or-system-clock-glitch-spike/4770/1 "2026-10-02T20:21:31Z")

</div>

Hello,  
I’ve recently joined [ntppool.org](http://ntppool.org) with 1 IPv4 and 1IPv6 IP.

I’m using **chrony** (version 4.9), and I’ve to say I’m quite impressed by the quality of this software!  
The documentation also is of excellent quality.

I was also quite impressed by the ntppool infrastructure, especially the monitor system which really is a great piece of distributed system engineering! Congrats guys! 🙂

## The problem

I’m having regularly (approximately once per hour) **a big “glitch”** in a lot of metrics that I’m recording using chrony\_exporter with Prometheus. Those metrics concern the **PPS** source and the **system clock**.

I create this topic in case someone can help me solve this mystery.

**The problem also occurs when there is no network traffic and even before I joined the pool.**

In the 2 first graphs, NMEA (green) uses the scale on the left of the graph and PPS (yellow) uses the scale on the right.  
Look how the standard deviation of PPS from the sourcestats command (first graph) has sudden spikes.  
The last offset of PPS from the sources command (second graph) as well.

You can see below them the DOP of my GNSS during that time (I found no correlation between PPS spike and TDOP to be honest).

 ![pps stddev and pps last offset](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/1/1da61007357cadaae97a88ca4f7febb8de14d2f4.png)

Here, you can also see the tracking performance which also show glitches at the exact same moment :

 ![tracking frequency, skew and residual frequency](https://us1.discourse-cdn.com/flex016/uploads/ntppool/original/2X/d/d757209c80fdb486c5f42be14c080fe8d737ecd5.png)

## The setup

I have following setup:

- Raspberry Pi 2B
- Uputronics GPS/RTC Hat for Raspberry Pi (U-blox Neo M8 + RV3028 RTC with supercaps)

Note: I know that the 2B is far from the best hardware to run a stratum 1.  
(To cite a few examples of limitations: IRQs can not be moved and are all stuck on CPU0, the network card is connected to a USB hub, etc.)  
But no matter the hardware, I think every hardware can be tuned to try to extract its best performance. But obviously, with the rpi2, I know I won’t be able to go as far as other setups.

I’m using my own Linux distribution, with upstream/mainline kernel 7.2.x (from [kernel.org](http://kernel.org)):

- I compile a minimal kernel, flag by flag, and only turn on a feature or driver if I need it
- This machine only runs **gpsd** , **chrony** , and **node\_exporter** , that I compiled myself. Other processes are the kernel threads.
- The filesystem is in RAM. Every file in the “/” hierarchy is living in RAM. There is no sdcard driver. (The kernel has an initramfs whose goal is not to mount a root, but to be the OS root itself.)

=\> This simplifies the diagnostic and decreases uncontrolled noise compared to normal distributions: no logrotate, no crontab, no systemd, no sdcard I/O etc.  
Nothing that I’m unaware of can run on the system.

```auto
refclock PPS /dev/pps0 refid PPS lock NMEA poll 0 precision 3e-8
refclock SHM 0 refid NMEA poll 0 precision 1e-1 offset 0.035

rtcsync
makestep 60 1

dumpdir /run/chrony/dump
driftfile /run/chrony/drift interval 300
leapseclist /run/chrony/leap-seconds.list

# 4194304 clients
clientloglimit 536870912

ratelimit interval x burst y leak z

allow

bindcmdaddress 0.0.0.0
cmdallow 10.0.0.1
opencommands activity clients serverstats sources sourcestats tracking

sched_priority 99

```

**gpsd** is run like this: `gpsd -D 1 -G -n -N -p /dev/gnss0`  
**chrony** is run like this: `chronyd -n -r`

## Things I’ve tried

- Use cpufreq governor **performance** to force CPU to 900Mhz (instead of staying stuck at 600Mhz).
- Fine tune the GNSS to stop SBAS, enable stationary mode, and only enable Galileo and GPS
- Move the GNSS antenna from inside the garage to the roof outside (average Signal to Noise Ratio increased from 20dBHz to 36dBHz, average TDOP decreased from 1.23 to 0.72, used satellites increased from 13 to 16).
- Logged each ublox binary message during a glitch (the lock and time related flags were always stable according to the GNSS receiver)
- Stop the fridge in case the compressor start-up would create an electrical surge?
- Decrease the **precision** directive from 3e-8 to 1e-6 on my PPS source in case over-estimating it could cause chrony to “over-react”?
- Use **sched\_priority** directive to put the chrony process into SCHED\_FIFO mode.
- Lock chrony process on CPU1, 2 and 3 (exclude CPU0 which has all USB, NET\_RX, GPIO, I2C and UART interrupts).

## The only thing that has “worked”

The only thing that has worked, is use the chrony directive **maxupdateskew 1**.  
It then completely removes all big oscillations.  
But as I understand it, it’s just filtering a problem to prevent it from contaminating the system clock. It’s not really removing the original problem.  
If I understand it correctly, the directive asks chrony “please close your eyes when you become uncertain, don’t touch the system clock, and then open your eyes again when the storm has passed”.

## How to proceed now?

It’s very difficult for me to pinpoint which part of the pipeline exactly has a problem.  
It could be:

- The GNSS firmware being unhappy about signal quality
- The GNSS firmware becoming unstable (may be because of a bad configuration or because it is routinely doing something disturbing the PPS signal)
- The capture by the kernel (pps-gpio) of the PPS signal could be delayed due to interrupt latency
- Chrony itself could be woken up late
- The system clock itself could be of bad quality and have spikes

As far as I understand, in most metrics, chrony is always comparing something with the system clock. So it’s hard to know if it’s the source or the system clock which is faulty, of it both are correct and it’s chrony which is messing with the system clock.

Does anyone have an idea on how to debug this?

---

<div class="post-metadata">

**Author:** ![Bas](https://sea2.discourse-cdn.com/flex016/user_avatar/community.ntppool.org/bas/32/465_2.png) [@Bas](https://community.ntppool.org/u/Bas)\
**Post date:** [October 4, 2026, 2:00pm UTC](https://community.ntppool.org/t/chrony-pps-or-system-clock-glitch-spike/4770/2 "2026-10-04T14:00:11Z")

</div>

In my opinion you poll your GPS too often.  
You should not poll it that much.  
My RS232 GPS is set this way:

```auto
refclock SHM 0 refid GPS poll 3 offset 0.083 delay 0.2
refclock SHM 1 refid PPS poll 3 precision 1e-9 maxlockage 32 lock GPS prefer

```

I found SHM to work better then /dev/pps0, as both signals are generated by GPSD anyway.

Further reduce clientloglimit to about 8192000, yours contains probably way too much entries for such a small system. It will take to much cpu-cycles to maintain them.

Also, your makestep is far to agressive, typical is this: makestep 1.0 3

As for allow, I doubt that does anything as it should be: allow all

And I use these to make it less spiking:

```auto
# Make scheduler change
sched_priority 40

# Run only in RAM
lock_all

```

Hopefully it works better.
