feat(memtrack): capture allocation stacks in eBPF - #522
not-matthias wants to merge 21 commits into
Conversation
Merging this PR will degrade performance by 33.59%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| 🆕 | WallTime | encode_events_realistic[1] |
N/A | 733.4 ms | N/A |
| 🆕 | WallTime | encode_events_realistic[2] |
N/A | 432.6 ms | N/A |
| 🆕 | WallTime | write_stack_events |
N/A | 186.5 ms | N/A |
| 🆕 | Simulation | encode_events_realistic[1] |
N/A | 1.3 s | N/A |
| 🆕 | Simulation | encode_events_realistic[2] |
N/A | 1.3 s | N/A |
| 🆕 | Simulation | write_stack_events |
N/A | 405.5 ms | N/A |
| 🆕 | Memory | encode_events_realistic[1] |
N/A | 121.1 MB | N/A |
| 🆕 | Memory | encode_events_realistic[16] |
N/A | 131.5 MB | N/A |
| 🆕 | Memory | encode_events_realistic[2] |
N/A | 121.1 MB | N/A |
| 🆕 | Memory | encode_events_realistic[4] |
N/A | 121.1 MB | N/A |
| 🆕 | Memory | encode_events_realistic[8] |
N/A | 123.8 MB | N/A |
| 🆕 | Memory | write_events[10000] |
N/A | 2.1 MB | N/A |
| 🆕 | Memory | write_events[100000] |
N/A | 9.6 MB | N/A |
| 🆕 | Memory | write_stack_events |
N/A | 65.6 MB | N/A |
| 👁 | WallTime | encode_events_realistic[16] |
121.1 ms | 203.1 ms | -40.34% |
| 👁 | WallTime | encode_events_realistic[4] |
214.4 ms | 288.1 ms | -25.57% |
| 👁 | WallTime | encode_events_realistic[8] |
143.1 ms | 216.9 ms | -34.03% |
Comparing cod-3222-add-ebpf-based-dwarffp-unwinding (d0f3114) with main (6c953d6)
Footnotes
-
4 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
d99e3bf to
ed10235
Compare
|
de94fe1 to
2a7a839
Compare
GuillaumeLagrange
left a comment
There was a problem hiding this comment.
OLGTM, quite a big piece of work
4b4bbdd to
1085fd4
Compare
1085fd4 to
1b151c7
Compare
Copy the caller's user stack in chunks at allocator entry and fold an FNV-1a digest over it in the kernel. The digest rides on the allocation event as stack_hash; the copied bytes, a DWARF-numbered register snapshot and a frame-pointer walk are emitted once per distinct digest on a dedicated ring buffer, so unwinding and symbolication can happen offline. Capture stays off until userspace sets the rodata toggle, so allocator probes are unchanged by default. Refs COD-3222
Add the userspace half of allocation stack capture: env-driven configuration, stack-definition ring parsing, loss counters, per-pid module mapping tracking, a folding recorder that deduplicates definitions and counts occurrences, and the report it produces. Nothing constructs these yet; the tracker wiring follows. Refs COD-3222
Wire the capture rodata and map sizing into skeleton load, poll the stack-definition ring alongside the event ring, and expose the loss counters and frame-pointer chains. The attach worker snapshots module mappings while it holds a process stopped, which is the only point they are guaranteed readable. Guard the lifecycle: finishing with a live session would block forever on the recorder, and a second spawn would leave the capture rings undrained, so both now fail with a descriptive error. With capture disabled the ring buffer and frame-pointer map shrink to the allocator minimum rather than reserving tens of MiB. Refs COD-3222
Add a fixture with two non-inlinable malloc call paths and privileged tests over it: distinct call paths get distinct identities with module mappings for the binary and libc, repeated calls deduplicate, and the default-off path still reports allocations. Two cases guard failure modes the default budget cannot reach. The maximum copy budget is the only configuration that exercises the verifier's instruction limit, since the frozen rodata makes the copy and hash loops scale with the configured size. Shrinking the frame-pointer map to one slot proves exhaustion costs only the fallback chain, never an allocation event. Refs COD-3222
The symbol, unwind-data and debug-info extraction is not perf-specific: it turns a set of mapped ELF modules into the deduplicated keyed artifacts a profile references, whatever discovered the mappings. Memory mode needs the same pipeline, so it moves out of wall_time/profiler/perf into executor/shared/module_artifacts.
Memory mode needs the same per-pid module references walltime writes, so the five artifact fields move into a flattened `ModuleArtifacts` shared by both formats; walltime's JSON is unchanged, asserted against output captured from the flat struct. Flattening buffers those fields through serde's `Content`, which unlike serde_json's direct deserializer cannot parse a string JSON key into a pid, so pid-keyed maps get an explicit key-parsing helper.
Allocation stacks are raw addresses, so resolving them off-box needs the module geometry perf gets from PERF_RECORD_MMAP2. No single hook provides it: security_mmap_file has the file but runs before the VMA exists, and perf_event_mmap has the addresses but cannot resolve a path. So an LSM program caches the path once per inode and a perf_event_mmap fentry emits inode-keyed address records, joined in userspace while the maps are still live. Path resolution is only reachable from an LSM program at all, and only above 5.11 (bpf_d_path on the sleepable hook) or 6.12 (the bpf_path_d_path kfunc), with the bpf LSM active. MappingSupport probes both, and when neither holds stack capture is turned off rather than shipping stacks nothing can attribute.
Memory mode now turns the mappings memtrack recorded into the same keyed unwind_data/symbols.map files walltime writes, plus a memtrack.metadata referencing them per pid, so allocation stacks can be unwound off-box. Each mapping's inode is rechecked against the path before its ELF is read: BPF cannot produce a build id, so the recorded (dev, ino) is what proves the file on disk is still the one that was mapped rather than a rebuilt binary whose eh_frame would be bound to the wrong addresses.
Replace the BPF LSM path cache and mapping ring with inherited per-CPU PERF_RECORD_MMAP2 collectors. Store executable mappings as a terminal suffix in the main memtrack stream so existing timeline consumers remain compatible, then extract and order them in the runner before generating module artifacts.
Every frame's output buffer started empty and doubled its way to the compressed size, which for a 64k event frame is 8 realloc-and-copy steps over roughly 8 MB. Not measurable in wall clock at current frame sizes; it removes the copy traffic.
The memory instrument reports allocation counts and bytes per benchmark, which is what the encode path is bound by. It runs on a hosted runner like simulation does; the runner grants memtrack its capabilities during setup.
glibc exports cfree at the same file offset as free, so attaching both instrumented one function twice: every free() produced two Free events and two stack captures. Track (library, offset) pairs and skip symbols already covered by an alias.
…tation mimalloc lowers memtrack's own memory usage, fragmentation and allocation overhead compared to glibc's allocator. As a side effect, it also doesn't route through the exported malloc/free/calloc/realloc symbols, so it skips the allocator uprobes (attached system-wide with pid -1) that would otherwise fire for memtrack's own bookkeeping allocations. Pulled in via the ebpf feature, which the binary already requires.
Each stack record needs its frame-pointer chain looked up in the stack_traces map, which is a syscall per record. Doing that inside the ring-buffer parse callback made the poll thread pay it, so a burst of stacks could push it behind the producer and records were dropped. ResolvingPoller wraps a RingBufferPoller with a dedicated resolver thread: the poll thread only parses (event, stackid) and hands it over an internal channel, and the resolver does the map lookup and forwards the completed event. Drop order keeps the existing shutdown contract, the ring is dropped first so its poll thread joins and closes the internal sender, which lets the resolver drain what it already has before its join returns.
…y map A single-entry ARRAY map lookup is not actually checked at runtime: for a constant in-range key the verifier strips PTR_MAYBE_NULL and drops the `if (!enabled)` branch as dead code, while the inlined lookup still re-reads the key from the BPF stack. Uprobe programs run under migrate_disable() only, so a program preempting one on the same CPU shares its per-CPU private stack and can clobber that key slot, turning the lookup into an unchecked NULL deref. A global has no key to clobber. This supersedes 97d269c, whose fail-closed branch the verifier removes anyway. The two toggles became byte-identical apart from the written value, so they now delegate to a shared `set_tracking`.
Closing the last reference to a uprobe link fd waits for an RCU-tasks-trace grace period, and hangs indefinitely when the kernel is wedged. Doing that work in this process makes teardown unbounded no matter how many threads share the wait, which is what 705245d tried to solve. Fork holder children over disjoint fd chunks instead: they own the terminal close, so this process only drops duplicate references. Holders that do not exit within a shared 30s deadline are abandoned for init to reap, bounding teardown even when the grace period never completes.
1b151c7 to
d0f3114
Compare
Adds raw allocation-stack capture to memtrack and ships everything needed for off-box unwinding.
At every allocator entry the eBPF program copies the caller's user stack (chunked, budget-capped, FNV-1a-hashed) and walks frame pointers in-kernel. Definitions are deduplicated by hash in-kernel: one
StackDefinitionevent (raw bytes, registers, FP chain) per distinct stack, referenced by astack_hashon allocation events. Everything flows through the ordinaryMemtrackArtifactevent stream; unwinding and symbolication happen server-side later.After a capture-enabled run, the runner decodes the artifact and reuses the perf walltime machinery to dump per-module symbols, unwind data, and debug info into the profile folder, plus a
memory_metadata.jsonwith per-pid process names.Refs COD-3222