Skip to content

fix(pc-sampling): read samples into a buffer of their own - #164

Merged
CodingInAVan merged 1 commit into
mainfrom
pc-sampling-sampled-launch-threshold
Sep 29, 2026
Merged

CodingInAVan merged 1 commit into
mainfrom
pc-sampling-sampled-launch-threshold

Conversation

@CodingInAVan

Copy link
Copy Markdown
Contributor

Description

In KERNEL_SERIALIZED mode CUPTI moves every finished kernel's PC records into the SAMPLING_DATA_BUFFER given at configure time. gpufl passed that same buffer to cuptiPCSamplingGetData and zeroed its count first, which threw away everything held there. Only the per-PC overflow CUPTI creates once the 65,536-record buffer is full reached the output, so a session returned nothing until about 65,536 / (PC records per launch) launches had run, and then only later launches, under one correlation id.

Read into a separate buffer, as NVIDIA's samples do, and drain until CUPTI reports nothing pending. Emit one row per (launch, PC, stall reason) through Monitor::PushProfileSamples so per-launch attribution survives and large collects do not overrun the monitor ring.

Correct the comments, zero-sample warning and CLI help that blamed launch count or armed GetData, and fix the example the same way.

Type of Change

  • Bug fix
  • New feature
  • Documentation update

Testing

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas

In KERNEL_SERIALIZED mode CUPTI moves every finished kernel's PC records
into the SAMPLING_DATA_BUFFER given at configure time. gpufl passed that
same buffer to cuptiPCSamplingGetData and zeroed its count first, which
threw away everything held there. Only the per-PC overflow CUPTI creates
once the 65,536-record buffer is full reached the output, so a session
returned nothing until about 65,536 / (PC records per launch) launches
had run, and then only later launches, under one correlation id.

Read into a separate buffer, as NVIDIA's samples do, and drain until
CUPTI reports nothing pending. Emit one row per (launch, PC, stall
reason) through Monitor::PushProfileSamples so per-launch attribution
survives and large collects do not overrun the monitor ring.

Correct the comments, zero-sample warning and CLI help that blamed
launch count or armed GetData, and fix the example the same way.
@CodingInAVan
CodingInAVan merged commit 9fc1a77 into main Sep 29, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant