Conversation
A parallel=True kernel called from a thread pool -- dask's threaded
scheduler, typically -- nests numba's pool under it. What that does
depends on numba's threading layer:
* tbb composes; nothing to do.
* omp starts a full team per calling thread: 179 OS threads from a
12-thread pool.
* workqueue aborts the process on concurrent entry. That is live today:
a chunked neighborhood reduction (the target="parallel" gufunc runs
inside apply_ufunc(dask="parallelized")) kills the interpreter. It
predates the nogil change; it aborts with the GIL held too.
numba_pool() in uxarray/utils/parallel.py guards each call by layer: a
lock under workqueue (on every thread, since a main-thread call racing a
dask task aborts too), numba.set_num_threads(1) off the main thread
under omp, restored afterwards because the setting is thread-local. The
14 parallel=True kernels move to @parallel_njit, which is njit plus that
guard; the neighborhood gufunc wraps its call in it.
The two remedies are per layer on purpose. Capping does not stop the
workqueue abort, and capping under the lock would run every pool-thread
call one at a time on one thread.
Chunked neighborhood percentile, 64 steps, best of 3:
layer before after
tbb 139ms 139ms
omp 179 OS threads 47 OS threads
workqueue abort ~1.1s
The workqueue time is the layer, not the lock: main takes 1176ms under
dask's sync scheduler.
Test suite: 983 passed, 1 skipped. test_plot_with_features fails
identically before and after (matplotlib figure size, unrelated).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflict in uxarray/grid/neighbors.py: UXARRAY#1768 replaced the target="parallel" gufunc with njit kernels, running the serial one on dask blocks and _reduce_rows_parallel only on in-memory arrays. Took that version, and put _reduce_rows_parallel under @parallel_njit like the other parallel kernels: an in-memory reduction called from a user's own thread pool still aborted under workqueue. test_parallel drops the neighborhood case, whose kernel no longer exists; that path now goes through @parallel_njit like the centroid kernel it keeps. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1798
Overview
A parallel=True kernel called from a thread pool -- dask's threaded scheduler, typically -- nests numba's pool under it. What that does depends on numba's threading layer:
numba_pool() in uxarray/utils/parallel.py guards each call by layer: a lock under workqueue (on every thread, since a main-thread call racing a dask task aborts too), numba.set_num_threads(1) off the main thread under omp, restored afterwards because the setting is thread-local. The 14 parallel=True kernels move to @parallel_njit, which is njit plus that guard; the neighborhood gufunc wraps its call in it.
The two remedies are per layer on purpose. Capping does not stop the workqueue abort, and capping under the lock would run every pool-thread call one at a time on one thread.
Chunked neighborhood percentile, 64 steps, best of 3:
The workqueue time is the layer, not the lock: main takes 1176ms under dask's sync scheduler.
Test suite: 983 passed, 1 skipped. test_plot_with_features fails identically before and after (matplotlib figure size, unrelated).
PR Checklist
General
Testing & Benchmarking
Documentation and Examples
docs/api.rst; internal (private) function names start with an underscore (_)AI Disclosure
AI Usage: Claude Opus 5.5