plan handle cache fix - #1070
Conversation
|
@luraess Would you be able to have a look? |
There was a problem hiding this comment.
AMDGPU.jl Benchmarks
Details
| Benchmark suite | Current: 32dfa70 | Previous: 4ff61e3 | Ratio |
|---|---|---|---|
amdgpu/synchronization/context/device |
550 ns |
567.5 ns |
0.97 |
amdgpu/synchronization/stream/blocking |
230 ns |
232.5 ns |
0.99 |
amdgpu/synchronization/stream/nonblocking |
307.5 ns |
312.5 ns |
0.98 |
applications/bitonic_sort |
1654649 ns |
1650734.75 ns |
1.00 |
applications/convolution |
114364 ns |
106826.75 ns |
1.07 |
applications/floyd_warshall |
11720482 ns |
11787759.25 ns |
0.99 |
applications/histogram |
786033.75 ns |
789913 ns |
1.00 |
applications/prefix_sum |
253781 ns |
251799.25 ns |
1.01 |
array/accumulate/Float32/1d |
78586.25 ns |
72711.25 ns |
1.08 |
array/accumulate/Float32/dims=1 |
267648.75 ns |
284589.75 ns |
0.94 |
array/accumulate/Float32/dims=1L |
88828.75 ns |
81133.75 ns |
1.09 |
array/accumulate/Float32/dims=2 |
85178.75 ns |
70688.75 ns |
1.20 |
array/accumulate/Float32/dims=2L |
2752692.5 ns |
2754571.75 ns |
1.00 |
array/accumulate/Int64/1d |
80371.25 ns |
77051.25 ns |
1.04 |
array/accumulate/Int64/dims=1 |
243926 ns |
242136.5 ns |
1.01 |
array/accumulate/Int64/dims=1L |
84058.75 ns |
83854 ns |
1.00 |
array/accumulate/Int64/dims=2 |
87288.75 ns |
70858.75 ns |
1.23 |
array/accumulate/Int64/dims=2L |
2893489.5 ns |
2892754.5 ns |
1.00 |
array/broadcast |
72758.5 ns |
72816.5 ns |
1.00 |
array/construct |
2235 ns |
2200 ns |
1.02 |
array/copy |
37165.5 ns |
36975.5 ns |
1.01 |
array/copyto!/cpu_to_gpu |
111399 ns |
110439.5 ns |
1.01 |
array/copyto!/gpu_to_cpu |
112054 ns |
110576.75 ns |
1.01 |
array/copyto!/gpu_to_gpu |
58845.75 ns |
58938.5 ns |
1.00 |
array/iteration/findall/bool |
138292 ns |
133184.25 ns |
1.04 |
array/iteration/findall/int |
148644.5 ns |
146924.25 ns |
1.01 |
array/iteration/findfirst/bool |
183707.75 ns |
183770 ns |
1.00 |
array/iteration/findfirst/int |
162694.75 ns |
145197 ns |
1.12 |
array/iteration/findmin/1d |
110244 ns |
109714 ns |
1.00 |
array/iteration/findmin/2d |
112094.25 ns |
106499 ns |
1.05 |
array/iteration/logical |
241226 ns |
240093.25 ns |
1.00 |
array/iteration/scalar |
287576.5 ns |
292796.75 ns |
0.98 |
array/permutedims/2d |
71396 ns |
70961.25 ns |
1.01 |
array/permutedims/3d |
70928.5 ns |
70623.75 ns |
1.00 |
array/permutedims/4d |
73811 ns |
73141.25 ns |
1.01 |
array/random/rand/Float32 |
38838 ns |
44920.75 ns |
0.86 |
array/random/rand/Int64 |
46345.5 ns |
53576 ns |
0.87 |
array/random/rand!/Float32 |
58715.75 ns |
64641 ns |
0.91 |
array/random/rand!/Int64 |
72351.25 ns |
71331.25 ns |
1.01 |
array/random/randn/Float32 |
80618.75 ns |
78961.25 ns |
1.02 |
array/random/randn!/Float32 |
80113.5 ns |
72186.25 ns |
1.11 |
array/reductions/mapreduce/Float32/1d |
101131.5 ns |
94511.5 ns |
1.07 |
array/reductions/mapreduce/Float32/dims=1 |
89928.75 ns |
75061.25 ns |
1.20 |
array/reductions/mapreduce/Float32/dims=1L |
831944.5 ns |
832329.5 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=2 |
87511.25 ns |
83001.5 ns |
1.05 |
array/reductions/mapreduce/Float32/dims=2L |
143689.75 ns |
143060 ns |
1.00 |
array/reductions/mapreduce/Int64/1d |
100709 ns |
93981.75 ns |
1.07 |
array/reductions/mapreduce/Int64/dims=1 |
90044 ns |
82693.75 ns |
1.09 |
array/reductions/mapreduce/Int64/dims=1L |
828552 ns |
833094.75 ns |
0.99 |
array/reductions/mapreduce/Int64/dims=2 |
90451.5 ns |
83931.5 ns |
1.08 |
array/reductions/mapreduce/Int64/dims=2L |
144284.75 ns |
143907.5 ns |
1.00 |
array/reductions/reduce/Float32/1d |
100948.75 ns |
94144.25 ns |
1.07 |
array/reductions/reduce/Float32/dims=1 |
90528.75 ns |
75216.5 ns |
1.20 |
array/reductions/reduce/Float32/dims=1L |
829094.75 ns |
832537.25 ns |
1.00 |
array/reductions/reduce/Float32/dims=2 |
87216.25 ns |
83219 ns |
1.05 |
array/reductions/reduce/Float32/dims=2L |
143742 ns |
143200 ns |
1.00 |
array/reductions/reduce/Int64/1d |
98919 ns |
91536.25 ns |
1.08 |
array/reductions/reduce/Int64/dims=1 |
90206.25 ns |
83593.5 ns |
1.08 |
array/reductions/reduce/Int64/dims=1L |
831364.25 ns |
829128.75 ns |
1.00 |
array/reductions/reduce/Int64/dims=2 |
87133.75 ns |
83766 ns |
1.04 |
array/reductions/reduce/Int64/dims=2L |
144259.5 ns |
143587 ns |
1.00 |
array/reverse/1d |
45030.75 ns |
45153 ns |
1.00 |
array/reverse/1dL |
75326.25 ns |
75516 ns |
1.00 |
array/reverse/1dL_inplace |
79653.75 ns |
79963.75 ns |
1.00 |
array/reverse/1d_inplace |
54255.75 ns |
60771 ns |
0.89 |
array/reverse/2d |
49568.25 ns |
43250.75 ns |
1.15 |
array/reverse/2dL |
60491 ns |
83061 ns |
0.73 |
array/reverse/2dL_inplace |
91433.75 ns |
89986.25 ns |
1.02 |
array/reverse/2d_inplace |
41990.5 ns |
55665.75 ns |
0.75 |
array/sorting/1d |
330699.75 ns |
333063 ns |
0.99 |
gemm/tiled |
1886265 ns |
1899696.5 ns |
0.99 |
gemm/tiled_unbounded |
1899457.5 ns |
1922439.5 ns |
0.99 |
integration/byval/reference |
40091 ns |
39001 ns |
1.03 |
integration/byval/slices=1 |
40650 ns |
39781 ns |
1.02 |
integration/byval/slices=2 |
146803 ns |
147233 ns |
1.00 |
integration/byval/slices=3 |
245973 ns |
245404 ns |
1.00 |
integration/volumerhs |
4892521 ns |
4962111 ns |
0.99 |
kernel/indexing |
57100.75 ns |
57025.75 ns |
1.00 |
kernel/indexing_checked |
58495.75 ns |
56630.75 ns |
1.03 |
kernel/launch |
1377.5 ns |
1365.25 ns |
1.01 |
kernel/rand |
102719 ns |
88876 ns |
1.16 |
latency/import |
1715513936 ns |
1752022784 ns |
0.98 |
latency/precompile |
40284399173 ns |
39860711432 ns |
1.01 |
latency/ttfp |
2332473179 ns |
2327930658 ns |
1.00 |
stencil/diffusion3d |
1617712.75 ns |
1624179.5 ns |
1.00 |
stencil/diffusion3d_checked |
1656408.25 ns |
1655412.5 ns |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
|
Thanks. I checked the proposed changes and iterated on it with the robot. The direction looks right to me (diagnosis matches what I tested, the fix works). I ran it on MI300 (gfx942, ROCm 7.2.3). The #1053 MWE now holds at 64 idle handles. Free memory goes flat from iteration 100 through 1200, with 195 MiB of drift that plateaus. Your tests pass as well: core 104/104, fft 7/7, A few things I noticed. On framing. Capturing the handle by value in 1. Eviction order in Your docstring already notes it is "a count budget, not an LRU". It may be a little sharper than arbitrary, though. Eviction restarts from the front each time and drains until the budget is met, so it tends to hit the same keys repeatedly. I measured this against the real On This might be one change together with Diff, if useful (
|
|
Thanks @luraess for the detailed feedback! I'll probably have all day on Tuesday next week to implement it.
I am happy to look at this in a separate PR. I can use LUMI for multi-GPU testing. |
Proposed fix #1053
HandleCachegot amax_idlebudget that bounds the total number of idle handles across all keys (not just per-key viamax_entries), with a newidle_dtorsdict holding each idle handle's destructor so that_evict_idle!can reclaim handles for any key once the budget is exceeded, running those destructors outside the lock. Also,release_plan!was changed to capture the plan handle in a local variable and close over that (() -> rocfft_plan_destroy(handle)) instead of closing over the plan and readingplan.handlelazily.With this implementation the MWE in #1053 gives the following output on LUMI:
➜ AMDGPU.jl git:(gd/plan-handle-cache-fix) ✗ julia --project=. mwe.jl Initial GPU memory: 63.896 GiB free / 63.984 GiB total Planning 6000 distinct FFT lengths (16, 18, 20, ... in steps of 2)... [50/6000] free=63.371 GiB cached_plan_shapes=64 idle_handles=64 used=0 bytes cached(pool reserved)=464 bytes [100/6000] free=63.293 GiB cached_plan_shapes=64 idle_handles=64 used=0 bytes cached(pool reserved)=864 bytes [150/6000] free=63.293 GiB cached_plan_shapes=64 idle_handles=64 used=0 bytes cached(pool reserved)=1.234 KiB [200/6000] free=63.262 GiB cached_plan_shapes=64 idle_handles=64 used=0 bytes cached(pool reserved)=1.625 KiB ...Is this PR going into the right direction?