When a numba kernel with parallel=True is called from a dask worker thread, numba starts its own thread pool inside dask's. What happens next depends on numba's threading layer. Under workqueue the whole process aborts ("Concurrent access has been detected"); running a percentile over a chunked neighborhood reduction on dask's threaded scheduler reproduces it today. Under omp it doesn't crash, but every calling thread gets a full thread team, so a 12-thread pool grows to 179 OS threads. Only tbb handles it cleanly, and users get whichever layer happens to be installed.
When a numba kernel with parallel=True is called from a dask worker thread, numba starts its own thread pool inside dask's. What happens next depends on numba's threading layer. Under workqueue the whole process aborts ("Concurrent access has been detected"); running a percentile over a chunked neighborhood reduction on dask's threaded scheduler reproduces it today. Under omp it doesn't crash, but every calling thread gets a full thread team, so a 12-thread pool grows to 179 OS threads. Only tbb handles it cleanly, and users get whichever layer happens to be installed.