Failure behaviour¶
A failing module disables only itself. The plugin never raises into NetBox: ready() wraps its own setup in a single try/except that logs one warning and leaves the process otherwise unaffected if setup fails outright, and every module is installed in its own try/except inside bootstrap.install, so one module's failure to build (a bad setting, a missing dependency, an exporter that cannot be constructed at all) does not stop any other module, and does not stop NetBox from serving requests or running jobs. Building an exporter never tries to reach the endpoint: an unreachable Collector is not a build failure at all, and is instead the "Collector unreachable" row below.
| Situation | Behaviour |
|---|---|
| OTel import error, or any other unexpected failure during setup | One warning (with the exception's traceback, since this is the plugin's own catch-all in PluginConfig.ready()); the plugin is inactive for that process; NetBox keeps running |
Invalid config in one section (logs, audit, traces, metrics or rq) |
One warning naming that section; that section is disabled; every other section resolves normally |
An invalid top-level value (enabled, service_name, resource_attributes or exporter not of the expected type) |
One warning; the whole plugin is disabled for that process |
An invalid exporter.* value (protocol, timeout, headers, insecure, certificate, insecure_skip_verify, including insecure_skip_verify combined with gRPC or with a certificate) |
Every enabled signal that resolves an exporter from it fails to resolve: logs and audit together (they share one exporter), traces, and metrics, each disabled with its own warning. A signal that is off to begin with is unaffected, since its exporter is never resolved |
No endpoint resolves for a module (config-time: none of <signal>.endpoint, exporter.endpoint, or the OTEL_* fallbacks) |
One warning, that module disabled; the rest of the plugin keeps running |
| The exporter's own construction fails at setup or after a fork (config resolved, but building the actual client fails, for example a gRPC CA certificate file that cannot be read) | One warning naming which export is disabled (log, trace or metric); the rest of the plugin keeps running. Read the note on gRPC certificates below |
| Collector unreachable | The exporter retries in the background, then drops what it could not send, bounded by its own queue; a web request never blocks on this. Each RQ work-horse waits up to rq.flush_timeout after its job before exiting, and the worker starts its next job only once that horse has exited. After rq.flush_breaker_threshold consecutive failed flushes (default 3), horses skip their flush and the worker returns to its normal pace; one horse every rq.flush_breaker_cooldown seconds (default 30) tries a full flush, and the first one whose exports succeed ends the skipping (see below). Metrics are exported from the plugin's own thread, which retries per the exporter's own behaviour and never blocks a request or a job |
Horse flush slow (Collector reachable but slow, or flush_timeout set low) |
Bounded by rq.flush_timeout, run on a helper thread; delays the worker's next job, never the running job's own completion. A flush that misses rq.flush_timeout counts as failed only when none of that horse's exports of the signal succeeded; one that exported something before the deadline (a large buffer sent to a reachable Collector) changes nothing. After rq.flush_breaker_threshold of them in a row (default 3), horses skip that signal's flush and drop what they buffered until a full flush attempt, one every rq.flush_breaker_cooldown seconds, finishes in time with its exports succeeding (see Collector outage and RQ throughput) |
| Error in the audit receiver, or in its commit callback | Caught and logged (one warning per process, naming only the exception type, since the message could contain object data); the save that triggered it proceeds regardless |
uWSGI running without thread support (enable-threads not set, classic uwsgi binary) |
One warning naming the fix (enable-threads = true); nothing is exported until it is set, since the exporter's background thread cannot run at all |
| An RQ wrap's target has an unexpected signature (an rq internal changed shape) | One warning naming the target; that wrap alone is skipped, every other wrap and module continues |
An SDK (TracerProvider, MeterProvider or LoggerProvider) already configured outside the plugin |
Reused instead of building a new one, with one logger.info line rather than a warning; with NetBox's default LOGGING = {}, Python's last-resort handler only shows WARNING and above, so this line is not shown unless LOGGING is configured to include it. An instrumentor that reports itself as already applied is left alone rather than instrumented a second time |
Collector outage and RQ throughput¶
A work-horse's flush (its log provider and its tracer provider, in parallel) is the only chance it gets to export what it buffered, since it exits with os._exit right after, which skips atexit. During a Collector outage the exporter cannot succeed, so a horse that flushes pays the full rq.flush_timeout (default 5 seconds) before it exits, and the worker only picks up its next job once that horse has exited. Without anything else, this would cap each worker at roughly one job per flush_timeout for as long as the outage lasts, regardless of how fast the jobs themselves run.
To avoid that, each worker keeps a small circuit breaker per signal, one for log records (including audit records) and one for spans, shared with the horses it forks through anonymous shared memory. Logs and traces can go to different endpoints (traces.endpoint), so an outage of one never makes horses drop the other:
- A signal's flush counts as failed when none of that signal's exports in the horse succeeded and either it did not finish within
rq.flush_timeout, or exports of that signal in the horse returned an error and the failed exports took at least 1 second in total, or a quarter ofrq.flush_timeoutor of that signal'sexporter.timeoutwhen that is shorter. The second case is for anexporter.timeoutshorter thanrq.flush_timeout: the flush returns in time after the exporter's retries, but the export failed. Withexporter.timeout = 1, for example, an export to a Collector that refuses connections gives up either at once (the exporter's first retry backoff, 0.8 to 1.2 seconds, does not fit in the time left) or after one backoff of about 0.8 to 1 second. The first costs the worker nothing and does not count; the second takes more than the 0.25 second threshold and counts. The time a horse pays is per export call:exporter.timeoutbounds one batch, so a horse with several batches, or with both log records and spans to flush, can pay more before its flush returns. It counts as successful when it finished in time and at least one export succeeded, even if another export of the same horse failed, however long that failure took. A flush that missedrq.flush_timeoutafter at least one of its exports succeeded, for example a bulk import withaudit.include_datawhose buffer does not fit inflush_timeout, shows that the Collector is reachable: it changes nothing, so it cannot open the breaker, nor close an open one. A flush with nothing to export, or whose exports were all rejected at once (see below), changes nothing: it neither adds to nor resets the count of failed flushes. - The breaker applies only to a provider the plugin builds. When NetBox reuses a
LoggerProviderorTracerProviderconfigured outside the plugin (for example byopentelemetry-instrument), the plugin cannot see that provider's export results, so horses always flush that signal in full. - After
rq.flush_breaker_thresholdconsecutive failed flushes of a signal (default 3), horses skip flushing that signal; when both are skipped, they exit right after their job. The horse that opens the breaker logs one warning, which appears in the worker's output. - Every
rq.flush_breaker_cooldownseconds (default 30), the next horse flushes that signal in full. If that flush succeeds, horses flush it normally again, and one warning reports how many flushes were skipped and how many buffered log records, or spans, they dropped. If it fails, horses skip their flush for anotherflush_breaker_cooldown.
During an outage, the records a skipping horse drops could not have been exported anyway. What the breaker gives up is the part of the outage that ends between two full flush attempts: after the Collector comes back, horses can keep skipping their flush for up to rq.flush_breaker_cooldown seconds, and what they buffered in that window is lost. Lower flush_breaker_cooldown to shorten that window, at the cost of one flush_timeout wait per worker per cooldown during an outage; set rq.flush_breaker_threshold = 0 to have every horse flush in full, as before this setting existed, and lower rq.flush_timeout instead.
The count of dropped records is the number of log records (including audit records) and spans waiting in the horse's batch queues when it skipped its flush. It relies on an internal of the OpenTelemetry SDK; if that internal is not available, the warning says in how many skipped flushes the records could not be counted. Records a horse's batch processor tried to export in the background during the job, and that failed, are not included, and nor is anything in a horse that was SIGKILLed.
A Collector that is reachable but so slow that rq.flush_breaker_threshold flushes in a row (default 3) hit rq.flush_timeout without a single export succeeding also triggers the breaker: horses then skip that signal's flush until one full flush finishes in time with its exports succeeding. A flush that hits rq.flush_timeout after some of its exports succeeded does not count, so a reachable Collector that is merely slower than flush_timeout for a large buffer never makes horses drop their records. See How it works, RQ work-horses for the paths that SIGKILL a horse outright (which neither the flush timeout nor the breaker has any effect on) and Limitations for the rest of what a low flush_timeout trades away.
An export that the Collector, or a proxy in front of it, rejects outright (for example HTTP 413 from a body size limit, 400, or 401 from bad credentials) returns at once: the exporter does not retry it, so the horse's flush costs the worker no time and skipping later flushes would gain nothing. Such a flush does not count as failed, so jobs that each produce one oversized batch (such as a large import split into chunked jobs, with audit.include_data) do not open the breaker, and horses running other jobs on the same worker keep flushing in full. The rejected batch itself is lost, as described under Body size rejections are not silent.
gRPC certificates and failure timing¶
Over gRPC, the CA certificate configured in exporter.certificate (or its signal-specific variant) is read from disk when the exporter is built, not on each connection: a missing or unreadable file fails that build immediately, producing the "could not build exporter" warning above, at process start or at the next fork. Over HTTP, the certificate path is instead handed to the underlying HTTP client and read when it opens a connection, so a missing file only surfaces as a failed export attempt later, not as a setup warning. See Configuration reference (exporter.certificate) for the same distinction from the configuration side.
Body size rejections are not silent¶
A batch over a receiver's or backend's configured body size limit is not a silent failure: the OTLP receiver answers with an HTTP 413 (or gRPC RESOURCE_EXHAUSTED), the exporter treats that as non-retryable, and the OTel SDK's own exporter code logs an error about it, in the NetBox process, on its own opentelemetry.* logger (excluded from every plugin export path, see Logs, feedback loop, and printed locally rather than sent anywhere). What is lost is the batch itself, not the record of losing it: nothing the plugin does surfaces this as one of its own warnings, since the failure is inside the SDK's exporter, not the plugin's own code, but it is not invisible either. See Collector, body size limits and Data safety for what this means for audit.include_data.
A genuinely silent failure¶
A Collector that accepts a batch (its receiver answers success) and only then fails to forward it to the actual backend, for example because the backend itself is unreachable or rejects it, never reports anything back to the process that sent it: the OTLP request already succeeded from the exporting process's point of view. This is the shape of failure that is actually invisible from inside NetBox; watch the Collector's own logs and metrics for it, not the plugin's.