Limitations¶
Every known limitation of the plugin, in one place, grouped by signal. Each item links to the page that explains it in full; this page is the index, not the only copy of the detail.
Logs¶
exception.messageon a log record is not scrubbed for URL query strings. This is unlike spans, where a status description, an exception event and URL attributes all have query strings replaced; log record content is a separate code path that the plugin does not apply the same scrubbing to. See Logs, known limitations and Data safety, traces in detail.- A 4xx or 5xx response returned, without raising, by a middleware that runs before the plugin's own (for example another plugin's middleware listed earlier in
PLUGINS) produces adjango.requestrecord without a trace id. Stock NetBox 4.7 has no such path. See Logs, known limitations.
Audit¶
- If NetBox updates an M2M change record in a transaction later than the one that created it, the data sent with the first record does not include that later update. This only matters with
audit.include_data; a same-transaction M2M update is handled correctly. See Audit records, including field data. - The 20,000-record log queue is shared between audit records and application logs: a burst of one can crowd out the other in the same process. See Audit records, sizing.
- When a
LoggerProviderconfigured outside the plugin is reused instead of one the plugin builds, its queue is not resized for audit bursts; the 20,000-record figure applies only to aLoggerProviderthe plugin builds itself. See Audit records, known limitations and How it works, an SDK configured outside the plugin. - With
audit.include_dataon, a record can grow large for an object with big JSON fields, and most Collectors reject a request over their configured body size limit, dropping the whole batch that record was in, not just that one record. See Audit records, sizing and Collector, what a Collector adds that the plugin does not.
Traces¶
- A psycopg connection opened before the plugin's
ready(), for example by startup code, is not traced until Django closes and reopens it (CONN_MAX_AGE). See Traces, known limitations. - Only a job enqueued through rq's
Queue.enqueue_jobcarries the enqueuing trace's context forward.enqueue_at,enqueue_inandenqueue_manyreach rq internals that never callenqueue_job, so a job created through any of them starts its own trace, including one scheduled byenqueue_atfrom inside a request. See Traces, context propagation into jobs. - A retried job, a requeued job, and a job the rq scheduler moves from scheduled back onto its queue all reuse the same
Jobobject and itsmeta, rather than being enqueued throughenqueue_jobagain, so they keep whatever context was stored at the original enqueue: a retry can surface in the original request's trace, possibly much later. See Traces, context propagation into jobs. - With django-rq's
COMMIT_MODEset to"request_finished", a job is enqueued after the request's span has already ended, so it is not linked to that request's trace. See Traces, context propagation into jobs. - With a
TracerProviderconfigured outside the plugin and reused, the plugin's redaction and its filter for parentless CLIENT spans do not apply; that provider's own configuration decides what is exported. See Traces, known limitations and Data safety, what a provider configured outside the plugin changes. - If an rq exception handler registered before the plugin's own returns
False, the job span still gets status ERROR, but without an exception event, since the plugin's handler is never reached. See Traces, known limitations. - The plugin wraps the global propagator so that baggage is never extracted or injected. A propagator set with
set_global_textmapafter startup, in a process that does not fork afterwards, replaces that wrapper, and baggage then flows again in that process. See Traces, baggage.
Metrics¶
- Outbound HTTP calls and object changes are counted only in a process that has a meter provider (
webandrqworker; see How it works, processes and roles): a call or a change made inside a job (which runs in the forked work-horse, not the worker parent) or inside any management command other thanrqworkeritself is not counted, since neither kind of process ever gets one. See Metrics, known limitations. - With a
MeterProviderconfigured outside the plugin and reused, the plugin's allowlist does not apply to it, and its own reader keeps running, unused, in every forked work-horse. See Metrics, known limitations and Data safety, what a provider configured outside the plugin changes. - A job whose Redis hash is already gone by the time the worker parent looks it up (for example
result_ttl=0) is counted with outcomeunknown. See Metrics, job metrics. netbox.rq.queue.depthis reported by every RQ worker process and by the rq scheduler process each worker forks (an unannounced fork keeps therqworkerrole rather than becoming a work-horse), so a single worker host emits more than one series per queue; aggregate withmax, notsum. See Metrics, queue depth.
RQ¶
rq.enabled = Falsedisables the whole RQ integration in one step:RqModule.install()never runs at all, soQueue.enqueue_jobis never wrapped (no enqueue-time trace context propagation),netbox.rq.queue.depthis never registered, andWorker.fork_work_horseis never wrapped either. That last part matters beyond just "norq_horserole": with the fork never announced, a work-horse keeps therqworkerrole, and if metrics are enabled, this unlabelled fork is treated as a fullrqworkerprocess and gets its own metrics pipeline, exactly as described forrq.patch_worker = Falsebelow. No job span is created and nothing is flushed before a horse exits either, sinceperform_jobis never wrapped. See Configuration reference (rq.enabled).rq.patch_worker = Falsetakes an early exit insideRqModule.install()(if not cfg.patch_worker: return) that skips three separate wraps together:perform_job(no job span),fork_work_horse(the fork is never announced, so a work-horse keeps therqworkerrole instead of becomingrq_horse, and nothing buffered before it exits is ever flushed), andexecute_job(nonetbox.rq.job.durationornetbox.rq.jobs; these two metrics are recorded by theexecute_jobwraps specifically, not by anything to do withfork_work_horseitself, but the same early exit skips all three). Because the fork is unannounced, the work-horse is, from the plugin's point of view, indistinguishable from a secondrqworkerprocess: if metrics are enabled, it gets its own complete metrics pipeline, and should it run longer than onemetrics.export_intervalbefore exiting, for example a slow job, it can export metrics (such as the queue depth gauge) under its own freshly builtservice.instance.id. Enqueue-time context propagation and the queue-depth gauge are registered before this check runs and are unaffected bypatch_worker = False. See How it works, RQ work-horses and Configuration reference (rq.patch_worker).- Every path that ends a work-horse with
SIGKILLinstead of lettingperform_jobreturn loses whatever that horse had buffered and not yet flushed, with no exception: working time pastjob.timeout + 60seconds, a stop-job command, a kill-horse command, and a cold shutdown of the worker (a second SIGINT or SIGTERM) all end this way. See How it works, RQ work-horses. - While work-horses skip their flush after repeated export failures (
rq.flush_breaker_threshold), what they buffered is dropped. During an outage it could not have been exported anyway, but after the Collector comes back horses can keep skipping for up torq.flush_breaker_cooldownseconds before the next full flush attempt ends the skipping. The breaker state is per worker process and per signal, so each worker detects the outage and the recovery on its own. See Failure behaviour. SpawnWorker(not used by NetBox's own worker setup) starts a fresh interpreter for each job rather than forking, and is not detected as a work-horse by the same fork-announcement mechanism described in How it works, RQ work-horses.- A retried, requeued or rescheduled job keeps whatever trace context (if any) was stored at its original enqueue, rather than getting fresh context from wherever it was retried or rescheduled; see Traces above for the detail and the corresponding gap in
enqueue_at/enqueue_in/enqueue_many.
Deployment¶
- The plugin targets NetBox 4.7.x only (
min_version = "4.7.0",max_version = "4.7.99"); wider support is a future milestone, not something this release covers. See the compatibility table in the project README and in the docs home page. service.instance.idhas a random part generated per process, so every process start, including a container restart, begins new metric series; a backend sees one set of series per process that ever ran, and old series linger until the backend marks them stale. There is no setting for a stable identity. See How it works, instance identity across restarts.- Each worker process (Granian, gunicorn, uWSGI) exports under its own
service.instance.id, so the data arrives as one series per worker process, not one series per host or per container; a fleet running--max-requestsor similar recycling reports many identities over time even though the container itself never restarts. See How it works, forking web servers. - Most Collectors enforce a maximum request body size, which a large
audit.include_databatch can exceed, dropping the whole batch rather than just the oversized record; sizing the Collector'sbatchprocessor and any receiver or backend limit needs to account for this. See Collector, what a Collector adds that the plugin does not and Audit records, sizing.