Many (probably all) NTP servers in the Philippines don't work

Agreed.

There seems to be a bit of blocking that I am seeing as well. For example requests from AWS ip ranges always getting i/o timeouts but requests from a non AWS hosting companies have no problems to the same pool server.

Gathering some data over the next few days/weeks to see what I can find out.

With a randomly selected set of about 520 clients (RIPE Atlas probes), over 99% are being filtered, among them about 60 from the Philippines, of which none reach the targets.

It doesn’t matter what zones the servers are or aren’t in, I think that is totally unacceptable. Not even sure why the user would put the servers into the pool and go through the trouble of identifying and whitelisting a few monitors only so they can then block for pretty much anyone else, supposedly allowing only their own clients, when setting up their own pool would have been so much easier. And they’d be sure only their clients access their servers, as that is all they seem to care about.

I have to disagree. The pool is a global project, and while it strives to assign local (same country) or at least regional (same continent) clients to a server, for technical and other reasons, there will always be clients from outside the server’s own country zone. Certainly when the server is in the continent zone as well, which is the default. I think when a server is not willing to help out country zones from the same continent that are unfortunate enough to not have any servers in them, they perhaps shouldn’t be in the pool in the first place. Some filtering is acceptable to protect the server, especially if targeted at identified misbehaving clients, or more general rate limiting. But flat-out ignoring the majority of clients is not.

I concede that it not being more widely known that the pool admins have the ability to not only assign different country zones to a server when it has been misplaced, but to drop the continent or country zone as well limits a server operator’s options to handle certain circumstances themselves. But I would have hoped they’d reach out to the pool admins via the e-mail address mentioned in many places on the pool’s website (if an operator isn’t aware of the forum, or willing to engage publicly), or raise the issue here on the forum, and ask for support via either channel, rather than resort to flat-out blocking the majority of clients.

There’s enough servers in the pool, it’s just a matter of them being distributed unevenly across the country zones, and many clients’ strict lock-in to their own country zone. Ask has previously acknowledged this to be an issue, and used to have plans to loosen that strictness a bit. With the current setup of the project, with him being the only who effectively can do such work, it unfortunately is likely to be a few years still until we can expect to see some relief on this front.

I suggested that also in the past, yet it should still be open for the world to request even when not given via the DNS.
Don’t forget, people travel and never change their pool settings, so if they once set it to PH but goto Europe, they keep getting the PH servers but get no time.

I’m all in favor to set preferable zones/countries to serve, but the system should decide if my server is needed outside the pref.
E.g. mine are set for Belgium, but are used for Europe if needed, doesn’t happen a lot, but it should.
Same way other servers are helping Belgium, as we have too little servers and too much clients.

That is the purpose of a pool. Locking everybody out is hurtful, if everybody does that, the pool will collapse.

Personally @alica I think you should reconsider, take own ph.pool-only or simply abandon the pool.
What you do now is hurtful to the pool.

No, I guess you misread my words. Currently I do not impose any network level ban on my own server, as you can see in my server’s page. What I said is a thought regarding unfulfilled continental-wide load back in the 2010s, not the current situation happening in ph zone.
Also in the past, querying pool.ntp.org may get random servers found all around the globe, not nearby servers, which is not the present case. At lease we can really consider abolishing (by CNAMEing) cc zones by now.

I guess we’ve reached a consensus that a pool NTP server should be generally available for everyone. This is true for the vast majority of the current pool NTP servers.

Algorithms could be tuned, yes, but except for these 10 NTP servers this doesn’t really seem to be a widespread problem. This current problem could be solved by contacting the owner of that NTP server cluster and giving two options: a) Lift the restrictions, or b) remove those NTP servers from the pool. If A does not happen or there’s no response, the option B would be implemented by pool admins.

I’ll note that we don’t know if the filtering is done by the NTP server admin or by Globe, their ISP. Regardless, it’d be up to the NTP server admin to sort out the situation.

Do you remember when that was still the case, and about what time that changed? Back in the 2010s as well, or more recently?

Hmm, why? At least I find it useful to be able to pick servers from a specific zone for specific circumstances, to bypass the automatic server assignment by the pool when desired.

Or are you referring to the client lock-in to country zones for the “global” names? In that case, I agree, somewhat.

In the sense that it should be a simple configuration change in the geodns targeting parameters to have it pick servers from the continent zone, rather than from the country zone, for queries for non-geographic zone names. The cc zones could remain untouched.

Whether that is the right approach to address the load issues in many parts of the world (vs., e.g., back-filling country zones with servers from outside the country zone in some way) is another matter, and out of scope here.

Hmm, this seems to be mixing two things that should be kept separate I think, even though they interact in some ways.

I agree the 10 NTP servers that are the scope of this thread indeed seem to not be a widespread problem. But the “algorithms could be tuned” part is, for a large number of zones, and even more users. In the sense of “algorithms” referring to how load is distributed in the system.

So I personally don’t think this “tuning” is merely a nice-to-have “could”, but rather a widespread problem in large parts of the world that is causing severe issues to a large number of people. It affects both users of NTP clients, but even more so NTP server operators, and numerous would-be server operators who would like to provide resources to the pool by adding their servers, but cannot because their server infrastructure cannot weather the traffic load that some zones attract.

However, that has been discussed in the past, in different threads. It’s not in the scope of this thread (though its effects are exacerbating the problems discussed in this thread), so I’ll leave it at that here.

I see a few people in this thread with comments along the lines of the “vast majority of pool servers are OK” and that things are not a “widespread problem”. While I agree with both these I am not sure we are monitoring and/or talking about it so the perception may be it is not widespread.

pool.ntp.org: Statistics for 2a12:bec0:168:343::123 and pool.ntp.org: Statistics for 2001:df4:bac0::e35b:f223 are both in the pool for Hong Kong. My testing shows a 100% failure rate (i/o timeout) from AWS ip addresses but a 100% success rate from non AWS hosting companies in Hong Kong and Singapore. These are not the only examples of possible pool quality problems but I am still in the info gathering mode - is it a filter, or a routing issue or something else. That is yet to be decided and will need more testing.

I think there may be similar issues to what is happening in the Philippines in other locations but they are not in the public view so to speak. I am hoping to highlight these sort of things in the future.

If anyone has suggestions for good hosting companies in a number of countries then let me know in another thread and I will see if I can add them to my monitoring system.

I think there is a marked difference between the AWS example, and the issue discussed in this thread. While from an AWS client, there may be 100% failure for whatever reason, at least on my servers, AWS traffic is typically just a fraction of the overall traffic, and thus of potentially affected clients.

On the other hand, there is an entire country zone, PH, with as of right now about 17 servers, and 10 of those are essentially not responding to anyone. As I showed in my measurements above, out of more than 500 RIPE Atlas probes worldwide, only 2 got a response. And none of them was from the PH zone.

There used to be way fewer servers in the PH zone, and the 10 servers being discussed in this thread now have a low share of the overall bandwidth (seem set at 512 kbit), so the issue probably is a bit better than it used to be. But still, if more than 50% of servers in a zone do not respond at all, to anyone (as far as we can tell), that is an issue for the clients in the zone. For all clients in the zone.

I used to have a server on AWS, and as AWS are shipping images that have the pool as NTP source, rather than their own servers (like some other hosters have it), a lot of traffic was coming from fellow AWS-hosted machines. And I was paying AWS (for traffic) to provide those AWS-based clients with time. Sorry, I understand if people don’t like that. That example has been discussed a few times already in this forum.

I still had some credit left over with a hosting provider in the Philippines so I setup a test server to see what things looked like from an ip address in the country. The results from the first run are:

time_stamp, test_result, location_of_request, dns_server_used, hostname_asked_for, ip_address_returned, reported_root_dispersion, reported_root_delay, round_trip_time, clock_offset_seconds, absolute_clock_offset_msec

1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, pool.ntp.org, 58.71.12.13, 0.032073975, 0.039916992, 0.016008629, 0.002671997, 2
1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, pool.ntp.org, 162.159.200.123, 0.001174927, 0.087203979, 0.001439069, 0.003380387, 3
1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, pool.ntp.org, 162.159.200.1, 0.001174927, 0.087203979, 0.001417647, 0.003445211, 3
1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, pool.ntp.org, 45.115.225.48, 0.021896362, 0.025466919, 0.032760454, -0.015005017, 15
1787623390, Tested_ERR, ph-mnl-lightnode, 127.0.0.1, ph.pool.ntp.org, 222.127.1.19, error: read udp 38.54.81.6:47779->222.127.1.19:123: i/o timeout
1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, asia.pool.ntp.org, 103.147.22.149, 0.000778198, 0.003219604, 0.061885926, 0.014575456, 14
1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, asia.pool.ntp.org, 103.186.118.222, 0.000564575, 0.003143311, 0.058697416, 0.012880236, 12
1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, asia.pool.ntp.org, 103.186.118.216, 0.001251221, 0.002990723, 0.063337152, 0.014855637, 14
1787623390, Tested_OK, ph-mnl-lightnode, 127.0.0.1, asia.pool.ntp.org, 51.16.77.36, 0.000610352, 0.005584717, 0.224755134, 0.012045652, 12
1787623390, Tested_ERR, ph-mnl-lightnode, 1.1.1.1, pool.ntp.org, 222.127.1.27, error: read udp 38.54.81.6:46458->222.127.1.27:123: i/o timeout
1787623390, Tested_OK, ph-mnl-lightnode, 1.1.1.1, ph.pool.ntp.org, 222.127.4.114, 0.032226563, 0.039916992, 0.020990125, 0.002502883, 2
1787623390, Tested_OK, ph-mnl-lightnode, 1.1.1.1, asia.pool.ntp.org, 103.186.118.221, 0.001159668, 0.003112793, 0.059902829, 0.013643453, 13
1787623390, Tested_OK, ph-mnl-lightnode, 1.1.1.1, asia.pool.ntp.org, 103.186.118.217, 0.000366211, 0.000183105, 0.063499902, 0.014963388, 14
1787623390, Tested_OK, ph-mnl-lightnode, 1.1.1.1, asia.pool.ntp.org, 158.252.7.7, 0.061462402, 0.269287109, 0.299032964, -0.090141857, 90
1787623390, Tested_ERR, ph-mnl-lightnode, 1.1.1.1, asia.pool.ntp.org, 59.103.236.10, error: read udp 38.54.81.6:48504->59.103.236.10:123: i/o timeout
1787623390, Tested_ERR, ph-mnl-lightnode, 8.8.8.8, pool.ntp.org, 222.127.1.26, error: read udp 38.54.81.6:35227->222.127.1.26:123: i/o timeout
1787623390, Tested_ERR, ph-mnl-lightnode, 8.8.8.8, ph.pool.ntp.org, 222.127.1.19, error: read udp 38.54.81.6:50112->222.127.1.19:123: i/o timeout
1787623390, Tested_OK, ph-mnl-lightnode, 8.8.8.8, asia.pool.ntp.org, 172.104.182.184, 0.010253906, 0.001098633, 0.030061539, 0.001765314, 1
1787623390, Tested_OK, ph-mnl-lightnode, 8.8.8.8, asia.pool.ntp.org, 103.152.101.233, 0.030120850, 0.148361206, 0.124732364, 0.003693509, 3
1787623390, Tested_OK, ph-mnl-lightnode, 8.8.8.8, asia.pool.ntp.org, 95.216.144.226, 0.000732422, 0.006317139, 0.284387368, 0.036410593, 36
1787623390, Tested_OK, ph-mnl-lightnode, 8.8.8.8, asia.pool.ntp.org, 212.73.86.164, 0.037322998, 0.001556396, 0.250901944, -0.001114448, 1
1787623390, Tested_OK, ph-mnl-lightnode, 9.9.9.9, pool.ntp.org, 23.143.196.203, 0.000488281, 0.007965088, 0.223764702, 0.002044758, 2
1787623390, Tested_OK, ph-mnl-lightnode, 9.9.9.9, pool.ntp.org, 45.79.189.79, 0.025177002, 0.017044067, 0.277971537, 0.000408039, 0
1787623390, Tested_OK, ph-mnl-lightnode, 9.9.9.9, pool.ntp.org, 192.207.55.254, 0.000000000, 0.000000000, 0.210529515, 0.013045245, 13
1787623390, Tested_OK, ph-mnl-lightnode, 9.9.9.9, pool.ntp.org, 172.234.25.10, 0.007598877, 0.126647949, 0.239348244, 0.000854690, 0
1787623390, Tested_ERR, ph-mnl-lightnode, 9.9.9.9, ph.pool.ntp.org, 222.127.1.21, error: read udp 38.54.81.6:47481->222.127.1.21:123: i/o timeout
1787623390, Tested_OK, ph-mnl-lightnode, 9.9.9.9, asia.pool.ntp.org, 103.186.118.212, 0.000839233, 0.002960205, 0.061503414, 0.014139913, 14
1787623390, Tested_OK, ph-mnl-lightnode, 9.9.9.9, asia.pool.ntp.org, 119.28.183.184, 0.030548096, 0.028915405, 0.044702156, -0.010690659, 10
1787623390, Tested_OK, ph-mnl-lightnode, 9.9.9.9, asia.pool.ntp.org, 123.204.232.128, 0.045043945, 0.035446167, 0.088590858, 0.004743610, 4
1787623390, Tested_OK, ph-mnl-lightnode, 194.156.163.137, asia.pool.ntp.org, 51.16.235.6, 0.000991821, 0.006332397, 0.239445974, 0.016001804, 16
1787623390, Tested_OK, ph-mnl-lightnode, 194.156.163.137, asia.pool.ntp.org, 121.174.142.82, 0.056396484, 0.008285522, 0.132679412, -0.005133357, 5

The formatting of the above is not that good but in summary the results were

     29  Tested_CACHED
     25  Tested_OK
      6  Tested_ERR

The Tested_CACHED are just tests that have already produced results in this run so I removed them from the list printout above. That means for a total of 31 checks, 25 were OK and 6 errored. Not ideal but the current situation may not be as bad as thought - the pool is still “mostly” working in the Philippines. Note that this is for ipv4 addresses only as that is all the hosting company provides.

I will leave the system running for a few days and let people know what things look like.

As far as I can recall, the config changed somewhat recently, probably after the COVID19 pandemic, but I cannot tell the exact date of change.

In the past using cc zone is the only way for clients to config for domestic servers, to save client’s (and the country’s) international bandwidth, especially in sea-locked zones like tw and jp. If pool.ntp.org can properly return nearby servers then there is no reason for generic clients to configure cc zones. Debugging other cc zone is only useful for admins who appear in this forum… :sweat_smile:

There’s now more than just the offending servers in the pool, and with higher weights. So indeed not as bad anymore as in the past, when the servers at issue were pretty much the only ones in the zone, and clients only received those in responses, and got no service at all.

Hmm, ok, I don’t recall such a change as recently as the pandemic, but it is difficult to reconstruct the behavior in the past.

What has changed is the understanding of how the system works. I.e., many people, including myself, believed for a long time that the “global” zones (without geographic reference in the name) would provide addresses from all over the world. Until some research work showed that that was actually not the case, and that the “global” names rather locked clients into their local zone, at least if it had at least a single server in it.

I believe it does, at least for zones that have a minimum of a single server in them.

I can see various reasons why a client might do so. E.g., when there are no servers in the client’s zone, the user might pick a “nearby” country zone rather than getting servers assigned from the continent zone, and thus possibly “far away”. Or when there are servers in the zone, but few, and providing bad service, e.g., due to overload. Or because the zone covers a large geographic area, and/or otherwise servers in a nearby country zone are “closer” than servers in their own. Or, in few cases, because the zone that a client is considered to be in by the pool is wrong, and the easiest, or only way to get servers from the client’s own zone is via explicit configuration. Or just for redundancy, i.e., in addition to the majority of servers from the pool-determined zone get a few from a manually-picked remote zone.

But clients from outside the zone also can end up on a server without the client’s user explicitly picking a remote zone. Namely, when there are no servers in the client’s own zone, and the pool falls back to assigning servers from the enclosing zone, i.e., in all but the rarest cases from the continent zone.

Unfortunately, there are no statistics that I am aware of that show the number of clients per possible way how/why they reach a particular server. E.g., I have three active servers in the JP zone, and right now, slighly more than 50% of traffic is identified as coming from the CN zone (like many of my servers in Asia see a large, if not dominating portion of clients from the CN zone).

The CN zone does have servers in it, so it is unlikely that those clients reach my servers via the fallback to the continent zone. Rather, I suspect this is manual configuration out of desperation due to bad service from servers from within the CN zone due to the high load there. But that is pure speculation, absent a way to actually determine why a client was assigned a particular server.

There’s another server hosted within the Globe network, and it seems reachable without noticeable issues.

So while we still don’t know for sure what is going on, and a single other working server is not sufficient to definitely rule out blocking at the ISP level, it suggests that that might not be the reason for the other servers not being reachable.

Remember cn zone was twice collapsed (2016-2017 and 2019) with only several IPv4 servers facing billions of clients. cn.pool.ntp.org could not return usable servers during this scenario, so there were tutorials teaching user to manually change the zone from cn to asia or nearby countries. An example:

I guess this workaround could be automated, eg. if a cc zone has low server:user ratio then querying pool.ntp.org inside this country will return server from nearby countries (better determined by actual measurements, not only geographical distance) and/or continental zones.

I can not agree with this, as only a few monitors pass and all the rest is blocked.
Among it a German monitor passes, but when I test with my German server, the query is blocked.
This means a deliberate block of networks and only let pass a few monitors to get handed by the DNS-server.

This is not an .PH filtering issue, this is a deliberate firewall setting on those servers, while set to @global. Meaning the pool is hurt by this behavior as clients are blocked to if not met the pass-criteria.

In my opinion servers that do this should be removed from the pool.
Sure I block a few abusive IP’s too, but I asked them to stop hitting my server every few seconds, they denied. But those aren’t networks or countries.

Now that those 10 servers have been out of the pool for some days I had a look at my traffic graphs. There has always been some differences in traffic between days and it’s hard to see if there has been any increase. Most probably there is some traffic increase to the remaining ph pool servers, but I don’t see anything dramatic. Thanks for the changes!