Skip to content

Monitoring with Prometheus and Grafana

ntpstats monitor (servers) and ntpstats watch (the local chrony or ntpd) can serve OpenMetrics for Prometheus. Besides the latest offset they export an error bound and rolling stability statistics, which standard exporters do not provide.

# NTS-authenticated measurements of three public servers, metrics on :9123
ntpstats monitor time.cloudflare.com nts.netnod.se ptbtime1.ptb.de --nts -i 64 --metrics-port 9123

# the local chrony, every 16 s
ntpstats watch chrony --metrics-port 9124

/metrics binds to 127.0.0.1 unless --metrics-host says otherwise. --metrics-window N (default 1024) sets how many samples per source feed the rolling statistics.

Metrics

Metric Meaning
ntpstats_offset_seconds last offset, reference − local
ntpstats_delay_seconds last round-trip delay (monitor)
ntpstats_offset_bound_seconds |offset| + delay/2 + root delay/2 + root dispersion: how far the local clock can be from the source's reference, given this sample
ntpstats_root_delay_seconds, ntpstats_root_dispersion_seconds, ntpstats_stratum as reported by the source
ntpstats_frequency_ppm local frequency correction (watch)
ntpstats_authenticated 1 when the sample was NTS-authenticated
ntpstats_last_sample_timestamp_seconds time of the last good sample
ntpstats_offset_stddev_seconds offset standard deviation over the window
ntpstats_tdev_seconds{tau} TDEV over the window at 1, 4 and 16 sample intervals
ntpstats_samples_total, ntpstats_errors_total{kind} counters; kind is e.g. timeout, kod_rate

All carry a source label (server name, or chrony/ntpd).

Dashboard, alerts and a ready-made stack

The repository's contrib/ directory has:

  • grafana/ntpstats-dashboard.json: a Grafana dashboard. It shows current offset and bound, time since the last sample, NTS status, and offset, bound, delay, rolling TDEV, stddev, root delay/dispersion, and sample and error rates. Import it and pick your Prometheus data source.
  • prometheus/ntpstats-rules.yml: alert rules for a stale source, an error bound above a limit (1 ms as an example; set yours, e.g. 100 µs for MiFID II RTS 25 HFT), unstable TDEV, errors or Kiss-o'-Death, and unauthenticated samples.
  • docker-compose.yml: ntpstats, Prometheus and Grafana, with the dashboard and rules provisioned.
cd contrib && docker compose up --build
# Grafana on http://localhost:3000 (admin/admin): dashboard "Time / ntpstats — network time"

The tests check that every metric the dashboard and rules use is exported, and that all their PromQL expressions parse.

OpenTelemetry

--otlp pushes the same metrics to an OpenTelemetry Collector, or any OTLP/HTTP receiver, as OTLP JSON. It uses only the standard library, with no SDK needed.

ntpstats monitor time.cloudflare.com --nts --otlp http://collector:4318 --otlp-interval 30
OTEL_EXPORTER_OTLP_ENDPOINT=https://otel.example OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer%20TOKEN" \
    ntpstats watch chrony --otlp
  • Names: OpenTelemetry style, ntpstats.offset (unit s), ntpstats.offset_bound, ntpstats.tdev{tau}, and so on. The counters ntpstats.samples and ntpstats.errors are cumulative monotonic sums.
  • Environment: the standard OTEL_EXPORTER_OTLP_ENDPOINT, …_METRICS_ENDPOINT, …_HEADERS, OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES are honoured.
  • Collector down: measurement goes on when the collector is unreachable; failed pushes are counted and retried at the next interval.

The payload is checked against the official OpenTelemetry protobuf definitions.

For CI checks (a GitHub Action and pytest assertions), see Audit, events & CI checks.

Please keep polling intervals at 64 s or more for servers you do not operate.