Monitoring with Prometheus and Grafana¶
ntpstats monitor (servers) and ntpstats watch (the local chrony or ntpd) can serve
OpenMetrics for Prometheus. Besides the latest offset they export an error bound and
rolling stability statistics, which standard exporters do not provide.
# NTS-authenticated measurements of three public servers, metrics on :9123
ntpstats monitor time.cloudflare.com nts.netnod.se ptbtime1.ptb.de --nts -i 64 --metrics-port 9123
# the local chrony, every 16 s
ntpstats watch chrony --metrics-port 9124
/metrics binds to 127.0.0.1 unless --metrics-host says otherwise. --metrics-window N
(default 1024) sets how many samples per source feed the rolling statistics.
Metrics¶
| Metric | Meaning |
|---|---|
ntpstats_offset_seconds |
last offset, reference − local |
ntpstats_delay_seconds |
last round-trip delay (monitor) |
ntpstats_offset_bound_seconds |
|offset| + delay/2 + root delay/2 + root dispersion: how far the local clock can be from the source's reference, given this sample |
ntpstats_root_delay_seconds, ntpstats_root_dispersion_seconds, ntpstats_stratum |
as reported by the source |
ntpstats_frequency_ppm |
local frequency correction (watch) |
ntpstats_authenticated |
1 when the sample was NTS-authenticated |
ntpstats_last_sample_timestamp_seconds |
time of the last good sample |
ntpstats_offset_stddev_seconds |
offset standard deviation over the window |
ntpstats_tdev_seconds{tau} |
TDEV over the window at 1, 4 and 16 sample intervals |
ntpstats_samples_total, ntpstats_errors_total{kind} |
counters; kind is e.g. timeout, kod_rate |
All carry a source label (server name, or chrony/ntpd).
Dashboard, alerts and a ready-made stack¶
The repository's contrib/
directory has:
grafana/ntpstats-dashboard.json: a Grafana dashboard. It shows current offset and bound, time since the last sample, NTS status, and offset, bound, delay, rolling TDEV, stddev, root delay/dispersion, and sample and error rates. Import it and pick your Prometheus data source.prometheus/ntpstats-rules.yml: alert rules for a stale source, an error bound above a limit (1 ms as an example; set yours, e.g. 100 µs for MiFID II RTS 25 HFT), unstable TDEV, errors or Kiss-o'-Death, and unauthenticated samples.docker-compose.yml: ntpstats, Prometheus and Grafana, with the dashboard and rules provisioned.
cd contrib && docker compose up --build
# Grafana on http://localhost:3000 (admin/admin): dashboard "Time / ntpstats — network time"
The tests check that every metric the dashboard and rules use is exported, and that all their PromQL expressions parse.
OpenTelemetry¶
--otlp pushes the same metrics to an OpenTelemetry Collector, or any OTLP/HTTP receiver, as
OTLP JSON. It uses only the standard library, with no SDK needed.
ntpstats monitor time.cloudflare.com --nts --otlp http://collector:4318 --otlp-interval 30
OTEL_EXPORTER_OTLP_ENDPOINT=https://otel.example OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer%20TOKEN" \
ntpstats watch chrony --otlp
- Names: OpenTelemetry style,
ntpstats.offset(units),ntpstats.offset_bound,ntpstats.tdev{tau}, and so on. The countersntpstats.samplesandntpstats.errorsare cumulative monotonic sums. - Environment: the standard
OTEL_EXPORTER_OTLP_ENDPOINT,…_METRICS_ENDPOINT,…_HEADERS,OTEL_SERVICE_NAMEandOTEL_RESOURCE_ATTRIBUTESare honoured. - Collector down: measurement goes on when the collector is unreachable; failed pushes are counted and retried at the next interval.
The payload is checked against the official OpenTelemetry protobuf definitions.
For CI checks (a GitHub Action and pytest assertions), see Audit, events & CI checks.
Please keep polling intervals at 64 s or more for servers you do not operate.