Skip to content

Commit bfea769

Browse files
authored
Merge pull request #184 from vchamarthi/asv-benchmarks
Add ASV benchmarks
2 parents b21a22c + 766de74 commit bfea769

19 files changed

Lines changed: 1151 additions & 0 deletions

‎.flake8‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -28,6 +28,7 @@ per-file-ignores =
2828
mkl_random/interfaces/__init__.py: F401
2929
mkl_random/tests/*.py: D102
3030
mkl_random/tests/**/*.py: D102
31+
benchmarks/**/*.py: D102
3132

3233
filename = *.py, *.pyx, *.pxi, *.pxd
3334
max_line_length = 80

‎.gitignore‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8,3 +8,6 @@ mkl_random/mklrand.cpp
88

99
# Byte-compiled / optimized / DLL files
1010
__pycache__/
11+
12+
# ASV benchmark artifacts
13+
.asv/

‎CHANGELOG.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
1010
* Added support for `array_like` (broadcastable) `low`/`high` bounds in `randint` [gh-168](https://github.com/IntelPython/mkl_random/pull/168)
1111
* Added support for free-threaded (GIL-disabled) CPython builds: the Cython extension is compiled with `freethreading_compatible=True`, so importing `mkl_random` no longer re-enables the GIL [gh-159](https://github.com/IntelPython/mkl_random/pull/159)
1212
* Added a thread-safety section to the how-to guide for free-threaded Python [gh-159](https://github.com/IntelPython/mkl_random/pull/159)
13+
* Added an [ASV](https://asv.readthedocs.io/en/stable/) benchmark suite under `benchmarks/` and a `benchmark` optional dependency group [gh-184](https://github.com/IntelPython/mkl_random/pull/184)
1314

1415
### Changed
1516
* Sped up `normal`, `uniform`, `exponential`, `laplace`, `gumbel`, `logistic`, `rayleigh` and `lognormal` for array-valued parameters. Seeded results change for these array paths; scalar paths are unchanged [gh-171](https://github.com/IntelPython/mkl_random/pull/171)

‎benchmarks/README.md‎

Lines changed: 107 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,107 @@
1+
# mkl_random ASV Benchmarks
2+
3+
Performance benchmarks for [mkl_random](https://github.com/IntelPython/mkl_random) using
4+
[Airspeed Velocity (ASV)](https://asv.readthedocs.io/en/stable/).
5+
6+
### Coverage
7+
8+
| File | API | Cases | Engines | Sizes |
9+
|------|-----|-------|---------|-------|
10+
| `bench_continuous.py` | `MKLRandomState` | Every continuous distribution; shape parameters chosen to hit each oneMKL method branch (e.g. gamma a>1, 0.6<a<1, a<0.6; beta Cheng/Johnk/Atkinson) | MT19937 | 1k, 1M |
11+
| `bench_discrete.py` | `MKLRandomState` | Every discrete distribution; binomial and hypergeometric in both table and acceptance/rejection regimes | MT19937 | 1k, 1M |
12+
| `bench_methods.py` | `MKLRandomState` | `method=` variants: gaussian, lognormal, multinormal Cholesky, poisson (POISNORM, PTPE at λ=10 and λ=100) | MT19937 | 1k, 1M |
13+
| `bench_engines.py` | `MKLRandomState(brng=...)` | uniform, normal and bounded `randint` per engine; full-range bits and `bytes` on engines that support UniformBits | all deterministic engines | 1M |
14+
| `bench_integers.py` | `randint` | Every dtype, each C code path (small range, wide range, power-of-two, full range), and array-valued bounds | MT19937 | 1k, 1M |
15+
| `bench_permutations.py` | `shuffle`, `permutation`, `choice` | 1-D, 2-D C/F order, lists; `choice` with and without replacement and `p` | MT19937 | 1k, 1M |
16+
| `bench_multivariate.py` | `multinomial`, `multivariate_normal`, `dirichlet` | small fixed dimensions | MT19937 | 1k, 1M |
17+
| `bench_array_params.py` | `MKLRandomState` | Distributions with one parameter value per output element | MT19937 | 1k, 100k |
18+
| `bench_interfaces.py` | `mkl_random.interfaces.numpy_random`, patched `numpy.random` | Common NumPy calls through the drop-in paths | MT19937 | 1k, 1M |
19+
| `bench_memory.py` | `MKLRandomState` | Peak RSS of 10M-element fills and of 10k repeated small calls | MT19937 | 10M |
20+
21+
## Threading
22+
23+
Set `MKL_NUM_THREADS` in the environment before running ASV to control the
24+
thread count used by MKL:
25+
26+
```bash
27+
MKL_NUM_THREADS=8 asv run --python=same --quick HEAD^!
28+
```
29+
30+
If `MKL_NUM_THREADS` is not set, `__init__.py` applies a default: **4** threads
31+
when the machine has 4 or more physical cores, or **1** (single-threaded)
32+
otherwise. This keeps results comparable across CI machines in the shared pool
33+
regardless of their total core count. Physical cores are detected via
34+
`psutil.cpu_count(logical=False)` — hyperthreads are excluded per MKL
35+
recommendation.
36+
37+
## Notes on Measurement
38+
39+
### Seeding and warmup
40+
41+
Every benchmark seeds its state with a fixed seed in `setup`. Timing
42+
benchmarks also make one untimed call per timed method, so stream
43+
initialization and first-call costs are not charged to the first measured
44+
iteration. Peak-memory benchmarks skip this, because ASV counts memory
45+
allocated in `setup` towards the peak.
46+
47+
### Engine restrictions
48+
49+
oneMKL does not implement `UniformBits32`/`UniformBits64` for `WH`, `MCG31`,
50+
`R250` and `MRG32K3A`, and on those engines mkl_random currently returns
51+
uninitialized memory instead of raising. Full-range integer draws and
52+
`bytes()` are therefore benchmarked only on the engines listed in
53+
`_utils._ENGINES_BITS`. The
54+
non-deterministic engine is excluded everywhere because its output, and its
55+
availability, depend on the hardware.
56+
57+
### Method names
58+
59+
An unrecognized `method=` string silently falls back to the default method.
60+
When adding a method benchmark, use a name listed in the docstring of the
61+
corresponding `MKLRandomState` method.
62+
63+
### Patched NumPy
64+
65+
`PatchedNumpyRandom` raises in `setup` if `numpy.random` is not served by
66+
mkl_random after `patch_numpy_random()`, so a broken patch shows up as a
67+
failed benchmark rather than as stock NumPy timings.
68+
69+
## Running Benchmarks
70+
71+
Prerequisites:
72+
73+
```bash
74+
pip install ".[benchmark]"
75+
```
76+
77+
Check that every benchmark imports and sets up:
78+
79+
```bash
80+
asv check --python=same
81+
```
82+
83+
Run benchmarks against the current environment:
84+
85+
```bash
86+
asv run --python=same --quick HEAD^!
87+
```
88+
89+
Compare two commits:
90+
91+
```bash
92+
asv continuous --python=same HEAD~1 HEAD
93+
```
94+
95+
View results in a browser:
96+
97+
```bash
98+
asv publish
99+
asv preview
100+
```
101+
102+
## CI
103+
104+
The benchmark pipeline installs `requirements.txt` with conda, so entries must
105+
be conda package names; keep it in step with the `benchmark` extra in
106+
`pyproject.toml`. Results are recorded only for branches listed under
107+
`branches` in `asv.conf.json`.

‎benchmarks/asv.conf.json‎

Lines changed: 25 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,25 @@
1+
{
2+
"version": 1,
3+
"project": "mkl_random",
4+
"project_url": "https://github.com/IntelPython/mkl_random",
5+
"show_commit_url": "https://github.com/IntelPython/mkl_random/commit/",
6+
"repo": "..",
7+
"branches": [
8+
"master",
9+
"dev-milestone"
10+
],
11+
"environment_type": "conda",
12+
"conda_channels": [
13+
"https://software.repos.intel.com/python/conda/",
14+
"conda-forge"
15+
],
16+
"benchmark_dir": "benchmarks",
17+
"env_dir": ".asv/env",
18+
"results_dir": ".asv/results",
19+
"html_dir": ".asv/html",
20+
"build_cache_size": 2,
21+
"default_benchmark_timeout": 500,
22+
"regressions_thresholds": {
23+
".*": 0.2
24+
}
25+
}

‎benchmarks/benchmarks/__init__.py‎

Lines changed: 46 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,46 @@
1+
# Copyright (c) 2026, Intel Corporation
2+
#
3+
# Redistribution and use in source and binary forms, with or without
4+
# modification, are permitted provided that the following conditions are met:
5+
#
6+
# * Redistributions of source code must retain the above copyright notice,
7+
# this list of conditions and the following disclaimer.
8+
# * Redistributions in binary form must reproduce the above copyright
9+
# notice, this list of conditions and the following disclaimer in the
10+
# documentation and/or other materials provided with the distribution.
11+
# * Neither the name of Intel Corporation nor the names of its contributors
12+
# may be used to endorse or promote products derived from this software
13+
# without specific prior written permission.
14+
#
15+
# THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
16+
# AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
17+
# IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
18+
# DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
19+
# FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
20+
# DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
21+
# SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
22+
# CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
23+
# OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
24+
# OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
25+
26+
"""ASV benchmarks for mkl_random"""
27+
28+
import os
29+
30+
import psutil
31+
32+
_MIN_THREADS = 4 # minimum physical cores required for multi-threaded mode
33+
34+
35+
def _physical_cores():
36+
"""Return physical core count; fall back to 1 (conservative)."""
37+
return psutil.cpu_count(logical=False) or 1
38+
39+
40+
def _thread_count():
41+
physical = _physical_cores()
42+
return str(_MIN_THREADS) if physical >= _MIN_THREADS else "1"
43+
44+
45+
_THREADS = os.environ.get("MKL_NUM_THREADS", _thread_count())
46+
os.environ["MKL_NUM_THREADS"] = _THREADS

‎benchmarks/benchmarks/_utils.py‎

Lines changed: 64 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,64 @@
1+
# Copyright (c) 2026, Intel Corporation
2+
#
3+
# Redistribution and use in source and binary forms, with or without
4+
# modification, are permitted provided that the following conditions are met:
5+
#
6+
# * Redistributions of source code must retain the above copyright notice,
7+
# this list of conditions and the following disclaimer.
8+
# * Redistributions in binary form must reproduce the above copyright
9+
# notice, this list of conditions and the following disclaimer in the
10+
# documentation and/or other materials provided with the distribution.
11+
# * Neither the name of Intel Corporation nor the names of its contributors
12+
# may be used to endorse or promote products derived from this software
13+
# without specific prior written permission.
14+
#
15+
# THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
16+
# AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
17+
# IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
18+
# DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
19+
# FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
20+
# DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
21+
# SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
22+
# CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
23+
# OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
24+
# OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
25+
26+
"""Shared helpers and parameter axes for mkl_random benchmarks."""
27+
28+
import mkl_random
29+
30+
_SEED = 42
31+
32+
# size axis shared across multiple files: per-call overhead and throughput
33+
_SIZES = [1_000, 1_000_000]
34+
35+
# Every deterministic basic generator. NONDETERM reads a hardware entropy
36+
# source, so its timings are neither reproducible nor available everywhere.
37+
_ENGINES = [
38+
"MT19937",
39+
"SFMT19937",
40+
"MT2203",
41+
"MCG31",
42+
"MCG59",
43+
"MRG32K3A",
44+
"R250",
45+
"WH",
46+
"PHILOX4X32X10",
47+
"ARS5",
48+
]
49+
50+
# Generators that implement viRngUniformBits32/64, which full-range integer
51+
# fills and bytes() call directly.
52+
_ENGINES_BITS = [
53+
"MT19937",
54+
"SFMT19937",
55+
"MT2203",
56+
"MCG59",
57+
"PHILOX4X32X10",
58+
"ARS5",
59+
]
60+
61+
62+
def _make_state(brng="MT19937"):
63+
"""Return a seeded MKLRandomState for the given basic generator."""
64+
return mkl_random.MKLRandomState(_SEED, brng=brng)
Lines changed: 81 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,81 @@
1+
# Copyright (c) 2026, Intel Corporation
2+
#
3+
# Redistribution and use in source and binary forms, with or without
4+
# modification, are permitted provided that the following conditions are met:
5+
#
6+
# * Redistributions of source code must retain the above copyright notice,
7+
# this list of conditions and the following disclaimer.
8+
# * Redistributions in binary form must reproduce the above copyright
9+
# notice, this list of conditions and the following disclaimer in the
10+
# documentation and/or other materials provided with the distribution.
11+
# * Neither the name of Intel Corporation nor the names of its contributors
12+
# may be used to endorse or promote products derived from this software
13+
# without specific prior written permission.
14+
#
15+
# THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
16+
# AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
17+
# IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
18+
# DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
19+
# FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
20+
# DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
21+
# SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
22+
# CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
23+
# OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
24+
# OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
25+
26+
"""Benchmarks for array-valued distribution parameters.
27+
28+
One parameter value per output element takes a separate code path from
29+
scalar parameters, so these are tracked apart from bench_continuous.py and
30+
bench_discrete.py.
31+
"""
32+
33+
import numpy as np
34+
35+
from ._utils import _SEED, _make_state
36+
37+
# Smaller than _utils._SIZES: the per-element path is much slower than a
38+
# scalar fill.
39+
_ARRAY_SIZES = [1_000, 100_000]
40+
41+
42+
class ArrayParams:
43+
"""Distributions called with one parameter value per output element."""
44+
45+
params = [_ARRAY_SIZES]
46+
param_names = ["size"]
47+
48+
def setup(self, size):
49+
self.rs = _make_state()
50+
rng = np.random.default_rng(_SEED)
51+
self.loc = rng.standard_normal(size)
52+
self.scale = rng.uniform(0.5, 2.0, size)
53+
self.high = self.loc + self.scale
54+
self.shape = rng.uniform(0.5, 5.0, size)
55+
self.lam = rng.uniform(1.0, 100.0, size)
56+
self.n = rng.integers(1, 100, size)
57+
self.p = rng.uniform(0.1, 0.9, size)
58+
self.rs.normal(self.loc, self.scale)
59+
self.rs.uniform(self.loc, self.high)
60+
self.rs.exponential(self.scale)
61+
self.rs.standard_gamma(self.shape)
62+
self.rs.poisson(self.lam)
63+
self.rs.binomial(self.n, self.p)
64+
65+
def time_normal(self, size):
66+
self.rs.normal(self.loc, self.scale)
67+
68+
def time_uniform(self, size):
69+
self.rs.uniform(self.loc, self.high)
70+
71+
def time_exponential(self, size):
72+
self.rs.exponential(self.scale)
73+
74+
def time_standard_gamma(self, size):
75+
self.rs.standard_gamma(self.shape)
76+
77+
def time_poisson(self, size):
78+
self.rs.poisson(self.lam)
79+
80+
def time_binomial(self, size):
81+
self.rs.binomial(self.n, self.p)

0 commit comments

Comments
 (0)