Skip to content

Emit OpenTelemetry spans over OTLP - #123

Open
ross wants to merge 1 commit into
mainfrom
otlp-tracing
Open

ross wants to merge 1 commit into
mainfrom
otlp-tracing

Conversation

@ross

@ross ross commented Sep 4, 2026

Copy link
Copy Markdown
Owner

Adds distributed tracing to simone. Each gunicorn worker sets up an OpenTelemetry TracerProvider and exports spans over OTLP/HTTP to whatever OTEL_EXPORTER_OTLP_ENDPOINT names — Grafana Tempo, in the compose stack this runs in.

The point is a request showing up as one trace across tiers. nginx now mints a trace id per request and forwards it as a W3C traceparent; the default propagator picks it up here with no code of our own, so simone's request span becomes a child of the span nginx's access log line produces, instead of the two being unrelated halves.

Why post_fork

Not import time, and not the opentelemetry-instrument CLI wrapper. BatchSpanProcessor owns a background export thread, and a thread does not survive fork() — only the forking thread's state does. A provider built in the gunicorn master would leave the worker holding a processor whose export thread only ever existed in the parent, silently exporting nothing. post_fork runs inside each freshly forked worker, before it imports simone.wsgi.

This agrees with the reason already documented at the top of gunicorn.conf.py for not running with --preload; nothing about the Cron-thread hazard changes.

The DJANGO_SETTINGS_MODULE move

DjangoInstrumentor().instrument() reads Django settings itself (to find MIDDLEWARE), and gunicorn runs post_fork before the worker imports simone.wsgi — where the setdefault used to live. Without moving it to module scope, Django's LazySettings configures itself from global_settings first, and since it only consults the env var the first time it's forced to configure, wsgi.py's later get_wsgi_application() never loads simone.settings at all. Every request then 500s with AttributeError on ROOT_URLCONF. Hit this directly while testing, not by inference.

Notes

  • The hook is a no-op when OTEL_EXPORTER_OTLP_ENDPOINT is unset, so script/run and the dev docker-compose.yml need no collector listening.
  • requirements.txt is additive only. Regenerating it with script/update-requirements also re-resolved a dozen unrelated pins (django-debug-toolbar 7→8, gunicorn 26.0→26.2, …); those were dropped so this diff is just the new packages. Worth a separate maintenance pass.
  • wrapt resolves to 2.4.1rc1 — a release candidate, pulled in by opentelemetry-instrumentation. It's what resolves across the 3.10–3.14 range this project supports, but flagging it rather than burying it.

Verified

Ran the hook directly against a local OTLP sink: provider installed with service.name=simone/service.namespace=xormedia, Django reported instrumented, an inbound traceparent correctly extracted, and one real application/x-protobuf POST to /v1/traces carrying a span whose trace id and parent span id match the inbound header. Existing test suite (30 tests) and script/lint/script/format pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_013LW5eUDxvLcyY9JKbfqMux

Adds tracing to each gunicorn worker via a post_fork hook, exporting spans
over OTLP/HTTP to whatever OTEL_EXPORTER_OTLP_ENDPOINT names (Grafana Tempo,
in the compose stack this runs in). Requests arriving with a W3C traceparent
become children of the caller's span, so a request shows up as one trace
across nginx and simone rather than two unrelated halves.

post_fork rather than import time, or the opentelemetry-instrument wrapper:
BatchSpanProcessor owns a background export thread, and threads don't survive
fork(), so a provider built in the master would leave the worker exporting
nothing. That agrees with the existing reason this app doesn't run with
--preload.

DJANGO_SETTINGS_MODULE moves up to module scope in the same file.
DjangoInstrumentor().instrument() reads Django settings itself, and gunicorn
runs post_fork before the worker imports simone.wsgi -- where the setdefault
used to live. Without this, Django's lazy settings would configure themselves
from global_settings first, wsgi.py's later get_wsgi_application() would never
load simone.settings at all, and every request would 500 on ROOT_URLCONF.

The whole hook is a no-op when OTEL_EXPORTER_OTLP_ENDPOINT is unset, so
running outside the compose stack needs no collector listening.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LW5eUDxvLcyY9JKbfqMux
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant