Validation¶
What the test suite checks¶
Around 340 tests run in CI on Python 3.9, 3.11 and 3.13, together with ruff and mypy. None of them need network access.
- Estimators: literal implementations of the NIST SP 1065 sums (ADEV, OADEV, MDEV, TDEV,
HDEV, TOTDEV, MTOT, Theo1, MTIE, TIErms); analytic log-log slopes for the five power-law noise
types; and a frozen table of values from an independent implementation (
allantools2024.06, stored intests/data, not a dependency). - Published reference values (NIST SP 1065): the NBS Monograph 140 nine-point data (table 30)
and the 1000-point test suite (table 31) are regenerated in
tests/test_reference_nist.py. The values agree to the 7 printed digits:
| Statistic | NBS 9-point (m = 1, 2) | 1000-point (m = 1, 10, 100) |
|---|---|---|
| ADEV, OADEV, MDEV, TDEV | ✓ | ✓ |
| HDEV (overlapping) | ✓ | ✓ |
| TOTDEV | ✓ | ✓ |
| MTOT, TTOT (bias-corrected) | ✓ | ✓ |
| HTOT (bias-corrected) | within 0.3 % | within 0.3 % |
The published MTOT values include Stable32's noise-type bias correction (white FM: variance
÷ 0.73). ntpstats applies the same correction by default; bias_correction=False or
--raw-mtot gives the raw SP 1065 eq. (27) value. On simulated noise, the raw MTOT/MVAR ratio
measured here is 0.99, 0.85, 0.77, 0.72 and 0.68 for white PM to random-walk FM, which matches
the factors used (0.94, 0.83, 0.73, 0.70, 0.69).
- EDF and confidence intervals: closed forms (white FM at m = 1: EDF = 2n/3), Monte Carlo
EDF for every estimator and noise type, and CI coverage. The exact discrete EDF is also
compared with the approximate OADEV formulas of SP 1065 table 5 (Stable32's simple method):
| Noise | agreement (N = 1025, m = 1…64) |
|---|---|
| white PM, white FM | within 3 % |
| flicker FM, random-walk FM | within 7 % |
| flicker PM | within 16 % (SP 1065 notes this approximation is the roughest) |
| - Noise model, hat and holdover: Monte Carlo coverage of the h_α intervals, the N-cornered | |
| hat and the holdover envelope (Metrology). | |
| - Noise identification: every α from +2 to −2. | |
| - Parsers: line layouts from the ntpd, chrony and linuxptp documentation and sources. The | |
| same simulated peer must read identically from peerstats, rawstats, chrony measurements and a | |
| pcap capture; this now guards the chrony sign convention corrected in 2.2.0. | |
| - Clients: SNTP, NTPv5 and NTS against local test servers, including Kiss-o'-Death, spoofed | |
| replies, forged NTS responses, certificate mismatch, TAI timescale and the 2036 era rollover. | |
| - Web UI: API, CSRF and Host-header checks, header injection, path traversal. |
Live interoperability¶
The Live interop workflow queries public NTPv4, NTS, NTS-pool, NTPv5 and Roughtime
servers every week from GitHub-hosted runners, and records one run per month in the open dataset
data/interop/.
Not yet done¶
- Comparison with actual Stable32 output files on further datasets, including confidence intervals (#28); the published tables and formulas are covered above.
- Long-term real logs against an independent reference (for example a GNSS-disciplined host).
Contributions of reference datasets and Stable32 outputs are welcome.
Validating your own setup¶
# chrony's view against its PPS reference clock: bias, RMS, TDEV and MTIE of the error
ntpstats compare /var/log/chrony/tracking.log /var/log/chrony/refclocks.log --ref-peer PPS0
To check a synchronisation algorithm rather than a deployment, simulate the scenario with
ntpstats bench and compare it with the reference estimators.