Skip to content

Stop the metrics feed asking for events it discards, and make allocation profiles usable - #30

Merged
kamilchodola merged 3 commits into
mainfrom
feature/dotnet-trace-clrevents-override
Sep 10, 2026
Merged

kamilchodola merged 3 commits into
mainfrom
feature/dotnet-trace-clrevents-override

Conversation

@kamilchodola

Copy link
Copy Markdown
Contributor

Three harness fixes that came out of investigating why superblocks K6 TTFB averages twice the processing time Nethermind reports over its SSE feed. They are independent; the first is the one with a measurement consequence.

Ask the metrics feed only for the event it reads

SseClient parses event: processed and discards everything else, but it subscribes to /data/events with no filter, so the client is asked for every event type. One of those is forkChoice, which makes Nethermind serialize the entire head block, every transaction, receipt and log, to JSON on each forkchoiceUpdated. On a 1 GGas superblock that is roughly 60 MB of JSON plus the writer's growth buffers, allocated on the large object heap, built on a thread-pool thread while the next block is being processed. Nethermind only does this while something is subscribed, so on a node with no dashboard open the cost does not exist. We are creating it by measuring.

The effect is not subtle. The same Nethermind image measured with and without that subscription:

superblocks TTFB avg
harness asks for all events 1515 ms (n=1)
harness asks for processed only 967-1075 ms (n=10)

That difference is entirely the serialization we asked for and threw away. It also inflates the client's large-object churn by about 125 MB per block, which drives extra background gen2 collections, which is what the TTFB gap is largely made of.

The change is a query string on the feed path. It is backward compatible in both directions: a node that does not understand events= ignores the query and streams everything as before, and the client filters as it always has. So this can merge independently of any client change.

Nethermind side: NethermindEth/nethermind#13334 adds the events= parameter and the per-type gating. Until this PR lands, PR-triggered benchmark runs use main and therefore keep measuring the client with the serialization included, which makes that change read as a no-op.

Stop the dotnet-trace collector before the client

The collector was stopped during scenario cleanup, after the client container had already been torn down. The runtime emits its method rundown when the trace session ends, so stopping it after the process is gone means the rundown never happens and no managed frame in the .nettrace resolves. Allocation events arrive with type names and sizes but no attribution.

Moving the stop ahead of the client teardown makes stacks resolve, which is what allowed the allocation above to be attributed to the data feed rather than guessed at from type names.

Let the sidecar's CLR event set be overridden

Alongside dotTrace the sidecar is hardcoded to gc+contention+threading+exception at informational, which does not include GCAllocationTick, so an allocation profile cannot be captured at all. EXPB_DOTNET_TRACE_CLREVENTS and EXPB_DOTNET_TRACE_CLREVENTLEVEL now override both, defaulting to today's values. Setting gc at verbose produces the allocation-by-type-and-stack profile used above.

Testing

Exercised end to end on the amd64 benchmark runner through run-expb-reproducible-benchmarks with expb_branch pointed at this branch: 5 multi-image superblocks runs, 3 fusaka runs, and 2 profiling runs with EXPB_DOTNET_TRACE_CLREVENTS=gc and EXPB_DOTNET_TRACE_CLREVENTLEVEL=verbose. The profiling runs produced .nettrace files with resolvable managed stacks, which the previous collector ordering did not.

If you would rather take the subscription fix on its own first, it is the last commit and stands alone.

@kamilchodola
kamilchodola merged commit 848dea7 into main Sep 10, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant