Changelog¶
All notable changes to this project are documented in this file. The format follows Keep a Changelog, and the project uses Semantic Versioning.
Unreleased¶
0.3.0 - 2026-09-29¶
Added¶
rq.flush_breaker_thresholdandrq.flush_breaker_cooldown: during a Collector outage, RQ work-horses skip their flush of log records, or of spans, after 3 failed flushes of that signal in a row (a flush that hitsrq.flush_timeout, or exports that return an error after retrying for at least 1 second, or a quarter ofrq.flush_timeoutor of that signal'sexporter.timeoutwhen that is shorter, in both cases with no export of that signal succeeding; a flush that hitsrq.flush_timeoutafter some exports succeeded does not count, nor does an export the Collector or a proxy rejects at once, such as HTTP413), with one full flush attempt every 30 seconds until exports succeed again. A worker no longer slows to about one job perrq.flush_timeoutfor the length of the outage. Records buffered in a horse that skips its flush are dropped; one warning (from the work-horse, in the worker's output) is logged when skipping starts and one when it ends, with the number of skipped flushes and dropped log records and spans. The breaker covers only providers the plugin builds, not aLoggerProviderorTracerProviderconfigured outside it.rq.flush_breaker_threshold = 0restores the previous behaviour (#4).- Dev stack: the
uicompose profile runsgrafana/otel-lgtm(Grafana, Loki, Tempo, Prometheus) with Grafana on port 3000, andmake dev-uistarts it with a Collector overlay (dev/otelcol/ui.yaml, loaded bydev/docker-compose.ui.yml) that forwards logs, traces and metrics to it over OTLP. Other dev targets takeUI=1. The default stack is unchanged (#6).
Changed¶
service.instance.idis now<hostname>-<pid>-<6 hex>: a random suffix, generated once per process and again in every forked child, is appended to the hostname and PID. A restarted container that got the same hostname and PID used to report the same identity as the one it replaced, which the OpenTelemetry semantic conventions do not allow. Every process start now begins new metric series, so a backend holds more series over time; aggregate acrossservice.instance.idrather than select one (#5).- Dev stack: the
uwsgiprofile now runs uWSGI on a uwsgi-protocol socket behind nginx, as NetBox'scontrib/uwsgi.inidoes, instead of uWSGI's built-in HTTP router. The router closed every connection after the response without sendingConnection: close, so a client that reused the connection could have its next request dropped. The e2e worker test no longer retries dropped logins and expects exactly one login record per login. It skips a web server only when the connection is refused and fails when the server answers with an error. nginx resolves the uWSGI container on each request, so rebuilding the stack does not leave it pointing at a stale address (#2). - Tests: new fork tests with real OTLP/gRPC exporters that export to a gRPC receiver running in its own process, from a parent and from several children forked from it (the parent must make a new metrics export after the forks, and records it left unflushed in its batch queues across the forks must be delivered exactly once, by the parent, and each child checks that the batch processors it inherits from the parent hold none of them, since the SDK clears their queues at fork), and a worker respawn test (gunicorn and uWSGI
max-requests, with Python's at-fork hooks, uWSGI'spost_fork_hook, or both). The test that holds the SDK's batch worker lock and the gRPC test now name the private SDK attributes they read when they skip, and a guard test fails when the pinned OpenTelemetry SDK minor version changes until those tests are revisited. The steps are listed in the development docs under "Updating OpenTelemetry" (#8).
Fixed¶
dev/scripts/check_dist.pyreports a malformed wheel filename, an unreadable sdist or wheel, or a wheel withoutMETADATAas a one-line error and exits non-zero, instead of failing with a traceback (#9).
0.2.2 - 2026-09-29¶
Fixed¶
- Outbound HTTP calls from NetBox no longer forward a client's
baggageheader: inbound baggage is dropped and only the trace context is propagated, with traces on or metrics only.OTEL_PROPAGATORSstill selects the trace-context format (#1). django.requestrecords for 4xx and 5xx responses, and other records Django writes after the request span has ended, now carry the request's trace id and span id (#12).
0.2.1 - 2026-09-28¶
Fixed¶
- The NetBox plugin page now shows the plugin's author;
authorandauthor_emailare set on the plugin config (#13).
0.2.0 - 2026-09-28¶
Added¶
exporter.insecure_skip_verify: send OTLP over HTTPS without verifying the Collector's certificate, for Collectors behind a self-signed certificate. HTTP only; rejected over gRPC and together withexporter.certificate. Logs a warning once per process and endpoint when on (#10).
0.1.0 - 2026-09-27¶
Added¶
- First release.
- Logs: NetBox's Python logs exported through the OTel Logs API, with allowlisted attributes and a feedback-loop filter so the plugin's own records and the exporters' internal logging are never exported.
- Audit records: one OpenTelemetry log record per committed
ObjectChange, with optional, filtered field data. - Traces: spans for Django requests, psycopg, redis and outbound
requestscalls, plus RQ job spans, with trace context propagated into enqueued jobs and redaction of headers, query strings and exception text. - Metrics: HTTP server and client request durations, RQ job duration and count, RQ queue depth, an object change counter, optional process runtime metrics, all restricted to an allowlist of names and attributes.
- Fork and multi-process safety for gunicorn, uWSGI and Granian, so each worker process exports under its own identity.
- RQ work-horse flush, bounding how long a job's log lines, audit records and spans wait to be exported before the horse exits.
- Configuration through
PLUGINS_CONFIGand the standardOTEL_*environment variables. - Reuse of an OpenTelemetry SDK already configured outside the plugin (for example by
opentelemetry-instrument). - NetBox 4.7 support.