Conversation
Adds tracing to each gunicorn worker via a post_fork hook, exporting spans over OTLP/HTTP to whatever OTEL_EXPORTER_OTLP_ENDPOINT names (Grafana Tempo, in the compose stack this runs in). Requests arriving with a W3C traceparent become children of the caller's span, so a request shows up as one trace across nginx and simone rather than two unrelated halves. post_fork rather than import time, or the opentelemetry-instrument wrapper: BatchSpanProcessor owns a background export thread, and threads don't survive fork(), so a provider built in the master would leave the worker exporting nothing. That agrees with the existing reason this app doesn't run with --preload. DJANGO_SETTINGS_MODULE moves up to module scope in the same file. DjangoInstrumentor().instrument() reads Django settings itself, and gunicorn runs post_fork before the worker imports simone.wsgi -- where the setdefault used to live. Without this, Django's lazy settings would configure themselves from global_settings first, wsgi.py's later get_wsgi_application() would never load simone.settings at all, and every request would 500 on ROOT_URLCONF. The whole hook is a no-op when OTEL_EXPORTER_OTLP_ENDPOINT is unset, so running outside the compose stack needs no collector listening. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013LW5eUDxvLcyY9JKbfqMux
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds distributed tracing to simone. Each gunicorn worker sets up an OpenTelemetry
TracerProviderand exports spans over OTLP/HTTP to whateverOTEL_EXPORTER_OTLP_ENDPOINTnames — Grafana Tempo, in the compose stack this runs in.The point is a request showing up as one trace across tiers. nginx now mints a trace id per request and forwards it as a W3C
traceparent; the default propagator picks it up here with no code of our own, so simone's request span becomes a child of the span nginx's access log line produces, instead of the two being unrelated halves.Why
post_forkNot import time, and not the
opentelemetry-instrumentCLI wrapper.BatchSpanProcessorowns a background export thread, and a thread does not survivefork()— only the forking thread's state does. A provider built in the gunicorn master would leave the worker holding a processor whose export thread only ever existed in the parent, silently exporting nothing.post_forkruns inside each freshly forked worker, before it importssimone.wsgi.This agrees with the reason already documented at the top of
gunicorn.conf.pyfor not running with--preload; nothing about the Cron-thread hazard changes.The
DJANGO_SETTINGS_MODULEmoveDjangoInstrumentor().instrument()reads Django settings itself (to findMIDDLEWARE), and gunicorn runspost_forkbefore the worker importssimone.wsgi— where thesetdefaultused to live. Without moving it to module scope, Django'sLazySettingsconfigures itself fromglobal_settingsfirst, and since it only consults the env var the first time it's forced to configure,wsgi.py's laterget_wsgi_application()never loadssimone.settingsat all. Every request then 500s withAttributeErroronROOT_URLCONF. Hit this directly while testing, not by inference.Notes
OTEL_EXPORTER_OTLP_ENDPOINTis unset, soscript/runand the devdocker-compose.ymlneed no collector listening.requirements.txtis additive only. Regenerating it withscript/update-requirementsalso re-resolved a dozen unrelated pins (django-debug-toolbar 7→8, gunicorn 26.0→26.2, …); those were dropped so this diff is just the new packages. Worth a separate maintenance pass.wraptresolves to2.4.1rc1— a release candidate, pulled in byopentelemetry-instrumentation. It's what resolves across the 3.10–3.14 range this project supports, but flagging it rather than burying it.Verified
Ran the hook directly against a local OTLP sink: provider installed with
service.name=simone/service.namespace=xormedia, Django reported instrumented, an inboundtraceparentcorrectly extracted, and one realapplication/x-protobufPOST to/v1/tracescarrying a span whose trace id and parent span id match the inbound header. Existing test suite (30 tests) andscript/lint/script/formatpass.🤖 Generated with Claude Code
https://claude.ai/code/session_013LW5eUDxvLcyY9JKbfqMux