Collector on Kubernetes¶
Why a Collector¶
The plugin can point exporter.endpoint (or OTEL_EXPORTER_OTLP_ENDPOINT) at any OTLP-compatible endpoint, including a backend's own native OTLP receiver directly; nothing in the plugin requires a Collector in between. Routing through an OpenTelemetry Collector instead is a recommendation, not a requirement, because it keeps a few things out of NetBox:
- Backend credentials live in the Collector's own configuration (and, on Kubernetes, a Secret), not in
PLUGINS_CONFIGor in a NetBox pod's environment. - Retries against a slow or unreachable backend are the Collector's problem, with its own queues and backoff, not something the plugin has to implement per signal.
- Routing (splitting signals across backends, duplicating to more than one, sampling with a processor such as
tail_samplingorprobabilistic_sampler, or filtering with thefilterprocessor) is configured once, in the Collector, without touching NetBox.
The NetBox side¶
Point the plugin at the Collector with exporter.endpoint in PLUGINS_CONFIG, or the equivalent OTEL_EXPORTER_OTLP_ENDPOINT environment variable, on both the NetBox web Deployment and the RQ worker Deployment:
PLUGINS_CONFIG = {
"netbox_opentelemetry_plugin": {
"exporter": {"endpoint": "http://otel-collector.<namespace>.svc:4318"},
},
}
or, as an environment variable on both Deployments:
env:
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: "http://otel-collector.<namespace>.svc:4318"
See Configuration for how each signal resolves its own endpoint from this base, and Installation for the full plugin setup.
Example files¶
Two files: a standalone Collector configuration, and a set of Kubernetes manifests that embed the same configuration in a ConfigMap. Both were validated once, manually, against the pinned Collector image and Kubernetes version: otelcol-contrib validate --config for the standalone configuration, kubeconform -strict for the manifests. Neither runs in CI (both need docker run against a pinned image, which the CI jobs in this repository do not do); rerun them by hand after changing either file. A test in this repository (tests/test_docs.py) does run in CI and asserts that the ConfigMap's embedded copy parses to the same YAML document as the standalone file (structural equality after yaml.safe_load, not a byte-for-byte comparison, so the two can differ in comments or formatting without failing the test), so the two cannot drift apart in substance.
docs/examples/kubernetes/collector-config.yaml:
# Example OpenTelemetry Collector configuration for the netbox-opentelemetry-plugin.
#
# This is a starting point, not a finished deployment. Replace:
# - the otlphttp exporter's endpoint (https://otlp.example.com) with your actual backend
# - the Authorization header, or the BACKEND_TOKEN environment variable it reads from, with
# whatever your backend expects (some backends use a different header name entirely)
# The debug exporter is useful while setting this up; drop it once the backend is confirmed to
# receive data, or keep it and rely on the Collector's own logging level to control its volume.
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
# Protects the Collector process itself from an unbounded queue if the backend is slow or down.
# Keep this first in every pipeline: it needs to see data before batch buffers more of it.
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 25
batch: {}
exporters:
debug:
verbosity: basic
# Replace this block with your real backend.
otlphttp:
endpoint: https://otlp.example.com
headers:
Authorization: "Bearer ${env:BACKEND_TOKEN}"
service:
pipelines:
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug, otlphttp]
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug, otlphttp]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug, otlphttp]
docs/examples/kubernetes/collector.yaml:
# Example Collector deployment for Kubernetes and OpenShift.
#
# Prerequisites this file does not create:
# - a Secret named otel-collector-backend with a key BACKEND_TOKEN, holding the credential the
# otlphttp exporter in collector-config.yaml sends to your backend
# - replacing the placeholder backend endpoint in that same file
#
# No runAsUser is set: on OpenShift, the restricted SCC assigns one from the namespace's allowed
# range at admission time. On plain Kubernetes this runs as whatever UID the image's own USER
# directive sets: UID 10001 for the upstream otel/opentelemetry-collector-contrib:0.161.0 image
# (docker image inspect shows this; it is not root, unlike this repository's own dev compose
# stack, which deliberately overrides it to root so its file exporter can write into a bind
# mount). runAsNonRoot is kept regardless, so the pod fails at admission instead of silently
# running as root if a future image change ever declares one; set an explicit runAsUser here if
# your own image needs a specific UID.
apiVersion: v1
kind: ConfigMap
metadata:
name: netbox-otel-collector
data:
# Keep this in sync with collector-config.yaml; a test asserts the two are identical.
config.yaml: |
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
# Protects the Collector process itself from an unbounded queue if the backend is slow or down.
# Keep this first in every pipeline: it needs to see data before batch buffers more of it.
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 25
batch: {}
exporters:
debug:
verbosity: basic
# Replace this block with your real backend.
otlphttp:
endpoint: https://otlp.example.com
headers:
Authorization: "Bearer ${env:BACKEND_TOKEN}"
service:
pipelines:
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug, otlphttp]
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug, otlphttp]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug, otlphttp]
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: otel-collector
labels:
app: otel-collector
spec:
replicas: 1
selector:
matchLabels:
app: otel-collector
template:
metadata:
labels:
app: otel-collector
spec:
containers:
- name: otel-collector
image: otel/opentelemetry-collector-contrib:0.161.0
args:
- --config=/conf/config.yaml
ports:
- name: otlp-grpc
containerPort: 4317
- name: otlp-http
containerPort: 4318
env:
- name: BACKEND_TOKEN
valueFrom:
secretKeyRef:
name: otel-collector-backend
key: BACKEND_TOKEN
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true
capabilities:
drop: [ALL]
seccompProfile:
type: RuntimeDefault
volumeMounts:
- name: config
mountPath: /conf
volumes:
- name: config
configMap:
name: netbox-otel-collector
---
# The NetBox web and worker Deployments point exporter.endpoint (or OTEL_EXPORTER_OTLP_ENDPOINT)
# at this Service, in-cluster only: http://otel-collector.<namespace>.svc:4318. A Route is not
# needed unless something outside the cluster must reach the Collector directly.
apiVersion: v1
kind: Service
metadata:
name: otel-collector
spec:
selector:
app: otel-collector
ports:
- name: otlp-grpc
port: 4317
targetPort: otlp-grpc
- name: otlp-http
port: 4318
targetPort: otlp-http
Replace the placeholder backend (https://otlp.example.com), its header, and the BACKEND_TOKEN environment variable it reads from, with whatever your actual backend needs; some backends use a different header name, or none at all. Create the otel-collector-backend Secret referenced by the Deployment separately, for example with kubectl create secret generic otel-collector-backend --from-literal=BACKEND_TOKEN=...; this example does not create it, since a Secret's value does not belong in a file checked into version control.
OpenShift notes¶
- The Deployment sets no
runAsUser. OpenShift's restricted SCC assigns a UID from the namespace's allowed range at admission time, which this example is written to accept rather than fight; setting an explicitrunAsUserwould need a SCC that allows it. On plain Kubernetes, with no SCC involved, the container runs as whatever UID the image itself declares:docker image inspect otel/opentelemetry-collector-contrib:0.161.0shows UID10001, not root (the dev compose stack in this repository overrides that to run as root for its own reasons; this example does not).runAsNonRoot: trueis still worth keeping regardless of which UID ends up in effect: it fails the pod at admission if a future image change (or a different image entirely) ever declares a root user, rather than silently running as root. - No
Routeis included, and none is needed for this setup: NetBox reaches the Collector through the in-clusterService(otel-collector.<namespace>.svc, ports 4317 and 4318), never from outside the cluster. Add aRouteonly if something outside the cluster, such as a second Collector forwarding to this one, needs to reach it directly.
Sidecar alternative¶
Instead of a standalone Deployment reached over the Service, the OpenTelemetry Operator can inject a Collector as a sidecar container into the NetBox pod itself, using an OpenTelemetryCollector resource with spec.mode: sidecar. The pod (the NetBox Deployment's pod template, not the Deployment's own metadata) is annotated sidecar.opentelemetry.io/inject: "true" (verified against the Operator's own sidecar injection documentation); the Operator then injects the Collector container alongside NetBox in every pod matching that annotation. With a sidecar, the endpoint is http://localhost:4318 (or 4317 for gRPC), since the Collector runs in the same pod network namespace as NetBox rather than behind a separate Service.
A sidecar gives every NetBox pod its own Collector instance rather than sharing one Deployment; consider the trade-off in resource overhead per pod against the isolation and per-pod queuing it buys, before choosing it over the standalone Deployment above.
What a Collector adds, that the plugin does not¶
- Host metrics. The plugin's own runtime metrics (
metrics.runtime) are process-level only (see Metrics); host-widesystem.*metrics (CPU, memory, disk, network of the node or container itself) are not something the plugin collects, by design, since every NetBox process on a host would report the same values. Add the Collector's ownhostmetricsreceiver (or, on Kubernetes, akubeletstatsor node-level Collector) if you need those. - Body size limits. The OTLP HTTP and gRPC receivers, and most backends behind them, enforce a maximum request body size. A batch that exceeds it is rejected whole, not partially accepted; this matters most for audit records with
audit.include_dataon, which can be large per record (see Audit records and Data safety). Size thebatchprocessor's batch size, and any receiver or backend limit, with that in mind, or keepinclude_dataoff and rely onexclude_fieldsinstead. include_datais a NetBox-side setting, not a Collector one. The Collector has no way to reconstruct data the plugin never sent; there is no processor that adds it back. Filtering happens once, in the plugin, before export (see Data safety).