From bcfa2447fe77ca90456d46f41ff73d38bce6060a Mon Sep 17 00:00:00 2001 From: Mark Saroufim Date: Thu, 17 Sep 2026 14:50:45 -0700 Subject: [PATCH] Document NCU profiling support for new problems --- README.md | 3 + docs/ncu-profiling.md | 205 ++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 208 insertions(+) create mode 100644 docs/ncu-profiling.md diff --git a/README.md b/README.md index 1d25040aa..01c8e17d0 100644 --- a/README.md +++ b/README.md @@ -25,6 +25,9 @@ To add a new problem, create a new folder in the `problems/glory` directory wher - `task.yml` - This is the problem specification that will be used to generate test cases for different shapes - `task.py` - Specifies the schema of the inputs and outputs for the problem +For NVIDIA problems, follow [Adding Nsight Compute support](docs/ncu-profiling.md) +to implement the evaluator profile path and verify captured kernels on a real GPU. + You can evaluate problems with your own Modal account (they give you a free $30) by borrowing this [neat script from @gau-nernst](https://github.com/gpu-mode/reference-kernels/pull/96#issue-3850136894) ## License diff --git a/docs/ncu-profiling.md b/docs/ncu-profiling.md new file mode 100644 index 000000000..a4a676992 --- /dev/null +++ b/docs/ncu-profiling.md @@ -0,0 +1,205 @@ +# Adding Nsight Compute support to a problem + +Nsight Compute (NCU) collects NVIDIA GPU hardware counters. To support it, a +problem's evaluator must implement `profile` mode and launch the submitted +kernel inside an NVTX push/pop range named exactly `custom_kernel`. Merely +accepting `profile` and running `torch.profiler` is insufficient. + +This guide describes the Python evaluator contract used by KernelBot and +[popcorn-cli](https://github.com/gpu-mode/popcorn-cli). It applies to a single +NVIDIA GPU; AMD profiling requires a different profiler, and the current NCU +runner does not support multi-GPU tasks. + +## Choose an evaluator and benchmark shapes + +Start with a working example: + +- [QR v2 evaluator](../problems/linalg/qr_v2/eval.py): a dedicated NCU-compatible + profile path, with correctness checking outside the captured range. +- [Shared NVIDIA evaluator](../problems/nvidia/eval.py): separate NCU and + PyTorch profiler paths. Inspect the evaluator actually selected by your + `task.yml`; a task-specific evaluator can differ from the shared one. + +In `task.yml`, include the evaluator, input generator, submission, and their +imports in `files`, set `config.main` to the evaluator entry point, and provide +at least one representative entry under `benchmarks`. See the +[QR v2 task definition](../problems/linalg/qr_v2/task.yml). + +KernelBot profiles `benchmarks`, not `tests`: it launches a separate evaluator +run for each benchmark entry. With the Modal profiling command, `--benchmark-index N` +selects the zero-based `benchmarks[N]`; omitting it profiles every benchmark +entry. The evaluator reads the benchmark specification file supplied by the +runner. Do not hardcode a shape or import a separate shape list for profiling. + +Register the problem's leaderboard name, directory, and supported GPU names in +its competition YAML (for example, [linalg.yaml](../problems/linalg.yaml)). +The CLI resolves that mapping to find the task and rejects unsupported GPUs. + +## Mark the submitted work + +Adapt this worker to the task's existing `TestCase`, input cloning, and checking +helpers. `_clone_data` below is the task's cloning helper, not a library API; +use the same input-preservation semantics as your correctness/benchmark paths. +The `check_implementation` signature here follows the QR v2 evaluator. + +```python +import torch +from torch.cuda.nvtx import range as nvtx_range + +from reference import check_implementation, generate_input + + +def _run_single_profile_ncu(test): + from submission import custom_kernel + + data = generate_input(**test.args) + cloned = _clone_data(data) + # Finish setup and cloning before entering the captured range. + torch.cuda.synchronize() + + with nvtx_range("custom_kernel"): + output = custom_kernel(cloned) + torch.cuda.synchronize() + + # Keep reference work and correctness checking outside the captured range. + return check_implementation(data, output) +``` + +The runner filters with `--nvtx --nvtx-include 'custom_kernel/'`. The trailing +slash selects a push/pop range; the range name in Python has no slash. +`torch.profiler.record_function("custom_kernel")` alone does not satisfy this +contract. A native CUDA evaluator needs the equivalent NVTX push/pop range, +with the NVTX headers and linking appropriate to its build. + +Capture one call to the submitted implementation, rather than the benchmark's +repeated timing loop. That call may launch multiple CUDA kernels. NCU performs +its own replay to collect counters. If the implementation requires warmup or +JIT compilation before capture, do it outside the range, restore any mutated +inputs, and synchronize before entering the range. + +If the evaluator uses worker processes, define the worker at module scope and +use the existing pool. The runner must use `--target-processes all` to capture +kernels launched in those children; the Modal CLI runner already does this. + +## Dispatch and report profile results + +Retain the evaluator's normal argument parsing, seeding, worker setup, and +`POPCORN_FD` logging protocol. For an evaluator with the QR-style checker, the +profile dispatcher can be: + +```python +def run_profiling(logger, pool, tests): + logger.log("benchmark-count", len(tests)) + passed = True + for idx, test in enumerate(tests): + logger.log(f"benchmark.{idx}.spec", test.spec) + good, message = pool.apply(_run_single_profile_ncu, (test,)) + logger.log(f"benchmark.{idx}.status", "pass" if good else "fail") + if not good: + logger.log(f"benchmark.{idx}.error", message) + passed = False + logger.log("check", "pass" if passed else "fail") + return 0 if passed else 112 +``` + +In the existing `main`, handle `mode == "profile"` by returning +`run_profiling(logger, pool, tests)`. Preserve `sys.exit(main())` so the exit +status reaches the runner. Here `tests` is the evaluator's parsed input list; +for profile mode its contents come from the runner's benchmark file. + +Do not invoke `ncu` or write `.ncu-rep` files from the evaluator. The runner +wraps the evaluator in NCU, applies capture options, and packages the reports. + +If you keep a PyTorch profiler path, select the NCU worker explicitly: + +```python +if os.environ.get("POPCORN_NCU") == "1": + return pool.apply(_run_single_profile_ncu, (test,)) +return pool.apply(_run_single_profile_torch, (test,)) +``` + +This snippet assumes `import os` and both workers exist. Avoid +`bool(os.getenv("POPCORN_NCU", "0"))`: the string `"0"` is truthy. Do not run +`torch.profiler` inside the NCU path; both profilers can compete for profiling +resources. An NCU-only profile path, like QR v2's, can call its worker directly. + +## Validate on a real GPU + +Use a popcorn-cli build that exposes `submit --profile`; check `popcorn submit +--help` or see the [CLI profiling guide](https://github.com/gpu-mode/popcorn-cli/blob/main/docs/profiling.md). +`--profile` runs in your own Modal account. `--profile-brev` explicitly selects +the separate Brev service; there is no automatic provider switch on failure. + +Install and authenticate Modal, then profile a known-correct starter submission: + +```bash +pip install modal +modal setup +popcorn submit submission.py --leaderboard YOUR_LEADERBOARD --gpu B200 \ + --profile --benchmark-index 0 +``` + +Use a GPU declared by the problem. Check all of the following before claiming +the new problem supports NCU: + +1. The run resolves the intended problem directory and benchmark specification. +2. A nonempty `.ncu-rep` is saved and can be reopened by NCU. +3. `ncu-details.txt` or `ncu-details.csv` names the intended submitted kernel and + contains actual counter values. A successful process exit or an empty zip is + not proof of capture. +4. A second, different benchmark index captures the corresponding shape. Check + the remaining shapes as appropriate for the task; one small shape does not + establish coverage for all workloads. +5. Normal `test` and `benchmark` modes still work after evaluator changes. + +The default capture limit is 10 kernel launches per benchmark. To reach a late +kernel, increase `--ncu-launch-count` or select it explicitly: + +```bash +popcorn submit submission.py --leaderboard YOUR_LEADERBOARD --gpu B200 \ + --profile --benchmark-index 0 \ + --ncu-kernel-name 'regex:your_kernel' --ncu-kernel-name-base demangled \ + --ncu-launch-count 1 +``` + +A filter that matches no launched kernels cannot produce a useful profile. +`--set full` requests a broad metric set but does not guarantee that every +metric exists or has data on every GPU. The Modal runner leaves GPU clocks +unchanged, so use the normal benchmark path for latency comparisons. + +## Record source provenance + +The profiler runs the evaluator from its active reference-kernels checkout. +Editing a local `eval.py` does not change a remote run. The Modal CLI resolves +current repository revisions when launched and saves the problem directory, +selected shapes, capture options, and source refs in `manifest.json`. +To validate a public problem branch before merging, push it and select its +full commit SHA explicitly: + +```bash +export POPCORN_REFERENCE_KERNELS_REF=FULL_COMMIT_SHA +popcorn submit submission.py --leaderboard YOUR_LEADERBOARD --gpu B200 \ + --profile --benchmark-index 0 +``` + +This override selects a ref in `gpu-mode/reference-kernels`; it does not upload +a local working tree or select an arbitrary fork. Brev uses its own deployed +checkout, so the Modal override does not update that service. + +In the problem PR, record the resolved problem directory, reference-kernels +commit, GPU, benchmark indices tested, kernel filter/launch limit, and observed +artifacts. For example, the existing Modal integration was verified against +`problems/linalg/qr_v2`, benchmark 0, on B200 at reference-kernels +`51e22db671d36c1c76091c43c36a44546ba324a1`. That is evidence for that run, not a +claim that every problem in this repository already supports NCU. + +## Troubleshooting + +| Symptom | Check | +| --- | --- | +| `profile` mode is rejected | Add real profile dispatch in the evaluator selected by `task.yml`. | +| No kernels captured | Check the exact NVTX push/pop range name, matching kernel filter, and child-process capture. | +| Only setup kernels appear | Move generation, cloning, and reference work outside the range; select the intended kernel or expand the launch limit. | +| Profiling initialization fails | Ensure the NCU path does not also start PyTorch's profiler; inspect the NCU version and driver/runtime errors. | +| Wrong shape or missing recent evaluator changes | Inspect `manifest.json` and the remote source ref; local edits are not automatically uploaded. | +| Timeout | Start with one shape and a bounded kernel capture; NCU replay can be much slower than a normal benchmark. |