Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .flake8
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ per-file-ignores =
mkl_random/interfaces/__init__.py: F401
mkl_random/tests/*.py: D102
mkl_random/tests/**/*.py: D102
benchmarks/**/*.py: D102

filename = *.py, *.pyx, *.pxi, *.pxd
max_line_length = 80
Expand Down
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,6 @@ mkl_random/mklrand.cpp

# Byte-compiled / optimized / DLL files
__pycache__/

# ASV benchmark artifacts
.asv/
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
* Added support for `array_like` (broadcastable) `low`/`high` bounds in `randint` [gh-168](https://github.com/IntelPython/mkl_random/pull/168)
* Added support for free-threaded (GIL-disabled) CPython builds: the Cython extension is compiled with `freethreading_compatible=True`, so importing `mkl_random` no longer re-enables the GIL [gh-159](https://github.com/IntelPython/mkl_random/pull/159)
* Added a thread-safety section to the how-to guide for free-threaded Python [gh-159](https://github.com/IntelPython/mkl_random/pull/159)
* Added an [ASV](https://asv.readthedocs.io/en/stable/) benchmark suite under `benchmarks/` and a `benchmark` optional dependency group [gh-184](https://github.com/IntelPython/mkl_random/pull/184)

### Changed
* Sped up `normal`, `uniform`, `exponential`, `laplace`, `gumbel`, `logistic`, `rayleigh` and `lognormal` for array-valued parameters. Seeded results change for these array paths; scalar paths are unchanged [gh-171](https://github.com/IntelPython/mkl_random/pull/171)
Expand Down
107 changes: 107 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# mkl_random ASV Benchmarks

Performance benchmarks for [mkl_random](https://github.com/IntelPython/mkl_random) using
[Airspeed Velocity (ASV)](https://asv.readthedocs.io/en/stable/).

### Coverage

| File | API | Cases | Engines | Sizes |
|------|-----|-------|---------|-------|
| `bench_continuous.py` | `MKLRandomState` | Every continuous distribution; shape parameters chosen to hit each oneMKL method branch (e.g. gamma a>1, 0.6<a<1, a<0.6; beta Cheng/Johnk/Atkinson) | MT19937 | 1k, 1M |
| `bench_discrete.py` | `MKLRandomState` | Every discrete distribution; binomial and hypergeometric in both table and acceptance/rejection regimes | MT19937 | 1k, 1M |
| `bench_methods.py` | `MKLRandomState` | `method=` variants: gaussian, lognormal, multinormal Cholesky, poisson (POISNORM, PTPE at λ=10 and λ=100) | MT19937 | 1k, 1M |
| `bench_engines.py` | `MKLRandomState(brng=...)` | uniform, normal and bounded `randint` per engine; full-range bits and `bytes` on engines that support UniformBits | all deterministic engines | 1M |
| `bench_integers.py` | `randint` | Every dtype, each C code path (small range, wide range, power-of-two, full range), and array-valued bounds | MT19937 | 1k, 1M |
| `bench_permutations.py` | `shuffle`, `permutation`, `choice` | 1-D, 2-D C/F order, lists; `choice` with and without replacement and `p` | MT19937 | 1k, 1M |
| `bench_multivariate.py` | `multinomial`, `multivariate_normal`, `dirichlet` | small fixed dimensions | MT19937 | 1k, 1M |
| `bench_array_params.py` | `MKLRandomState` | Distributions with one parameter value per output element | MT19937 | 1k, 100k |
| `bench_interfaces.py` | `mkl_random.interfaces.numpy_random`, patched `numpy.random` | Common NumPy calls through the drop-in paths | MT19937 | 1k, 1M |
| `bench_memory.py` | `MKLRandomState` | Peak RSS of 10M-element fills and of 10k repeated small calls | MT19937 | 10M |

## Threading

Set `MKL_NUM_THREADS` in the environment before running ASV to control the
thread count used by MKL:

```bash
MKL_NUM_THREADS=8 asv run --python=same --quick HEAD^!
```

If `MKL_NUM_THREADS` is not set, `__init__.py` applies a default: **4** threads
when the machine has 4 or more physical cores, or **1** (single-threaded)
otherwise. This keeps results comparable across CI machines in the shared pool
regardless of their total core count. Physical cores are detected via
`psutil.cpu_count(logical=False)` — hyperthreads are excluded per MKL
recommendation.

## Notes on Measurement

### Seeding and warmup

Every benchmark seeds its state with a fixed seed in `setup`. Timing
benchmarks also make one untimed call per timed method, so stream
initialization and first-call costs are not charged to the first measured
iteration. Peak-memory benchmarks skip this, because ASV counts memory
allocated in `setup` towards the peak.

### Engine restrictions

oneMKL does not implement `UniformBits32`/`UniformBits64` for `WH`, `MCG31`,
`R250` and `MRG32K3A`, and on those engines mkl_random currently returns
uninitialized memory instead of raising. Full-range integer draws and
`bytes()` are therefore benchmarked only on the engines listed in
`_utils._ENGINES_BITS`. The
non-deterministic engine is excluded everywhere because its output, and its
availability, depend on the hardware.

### Method names

An unrecognized `method=` string silently falls back to the default method.
When adding a method benchmark, use a name listed in the docstring of the
corresponding `MKLRandomState` method.

### Patched NumPy

`PatchedNumpyRandom` raises in `setup` if `numpy.random` is not served by
mkl_random after `patch_numpy_random()`, so a broken patch shows up as a
failed benchmark rather than as stock NumPy timings.

## Running Benchmarks

Prerequisites:

```bash
pip install ".[benchmark]"
```

Check that every benchmark imports and sets up:

```bash
asv check --python=same
```

Run benchmarks against the current environment:

```bash
asv run --python=same --quick HEAD^!
```

Compare two commits:

```bash
asv continuous --python=same HEAD~1 HEAD
```

View results in a browser:

```bash
asv publish
asv preview
```

## CI

The benchmark pipeline installs `requirements.txt` with conda, so entries must
be conda package names; keep it in step with the `benchmark` extra in
`pyproject.toml`. Results are recorded only for branches listed under
`branches` in `asv.conf.json`.
25 changes: 25 additions & 0 deletions benchmarks/asv.conf.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
{
"version": 1,
"project": "mkl_random",
"project_url": "https://github.com/IntelPython/mkl_random",
"show_commit_url": "https://github.com/IntelPython/mkl_random/commit/",
"repo": "..",
"branches": [
"master",
"dev-milestone"
],
"environment_type": "conda",
"conda_channels": [
"https://software.repos.intel.com/python/conda/",
"conda-forge"
],
"benchmark_dir": "benchmarks",
"env_dir": ".asv/env",
"results_dir": ".asv/results",
"html_dir": ".asv/html",
"build_cache_size": 2,
"default_benchmark_timeout": 500,
"regressions_thresholds": {
".*": 0.2
}
}
46 changes: 46 additions & 0 deletions benchmarks/benchmarks/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Copyright (c) 2026, Intel Corporation
#
# Redistribution and use in source and binary forms, with or without
# modification, are permitted provided that the following conditions are met:
#
# * Redistributions of source code must retain the above copyright notice,
# this list of conditions and the following disclaimer.
# * Redistributions in binary form must reproduce the above copyright
# notice, this list of conditions and the following disclaimer in the
# documentation and/or other materials provided with the distribution.
# * Neither the name of Intel Corporation nor the names of its contributors
# may be used to endorse or promote products derived from this software
# without specific prior written permission.
#
# THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
# AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
# IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
# DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
# FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
# DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
# SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
# CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
# OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
# OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

"""ASV benchmarks for mkl_random"""

import os

import psutil

_MIN_THREADS = 4 # minimum physical cores required for multi-threaded mode


def _physical_cores():
"""Return physical core count; fall back to 1 (conservative)."""
return psutil.cpu_count(logical=False) or 1


def _thread_count():
physical = _physical_cores()
return str(_MIN_THREADS) if physical >= _MIN_THREADS else "1"


_THREADS = os.environ.get("MKL_NUM_THREADS", _thread_count())
os.environ["MKL_NUM_THREADS"] = _THREADS
64 changes: 64 additions & 0 deletions benchmarks/benchmarks/_utils.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Copyright (c) 2026, Intel Corporation
#
# Redistribution and use in source and binary forms, with or without
# modification, are permitted provided that the following conditions are met:
#
# * Redistributions of source code must retain the above copyright notice,
# this list of conditions and the following disclaimer.
# * Redistributions in binary form must reproduce the above copyright
# notice, this list of conditions and the following disclaimer in the
# documentation and/or other materials provided with the distribution.
# * Neither the name of Intel Corporation nor the names of its contributors
# may be used to endorse or promote products derived from this software
# without specific prior written permission.
#
# THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
# AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
# IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
# DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
# FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
# DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
# SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
# CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
# OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
# OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

"""Shared helpers and parameter axes for mkl_random benchmarks."""

import mkl_random

_SEED = 42

# size axis shared across multiple files: per-call overhead and throughput
_SIZES = [1_000, 1_000_000]

# Every deterministic basic generator. NONDETERM reads a hardware entropy
# source, so its timings are neither reproducible nor available everywhere.
_ENGINES = [
"MT19937",
"SFMT19937",
"MT2203",
"MCG31",
"MCG59",
"MRG32K3A",
"R250",
"WH",
"PHILOX4X32X10",
"ARS5",
]

# Generators that implement viRngUniformBits32/64, which full-range integer
# fills and bytes() call directly.
_ENGINES_BITS = [
"MT19937",
"SFMT19937",
"MT2203",
"MCG59",
"PHILOX4X32X10",
"ARS5",
]


def _make_state(brng="MT19937"):
"""Return a seeded MKLRandomState for the given basic generator."""
return mkl_random.MKLRandomState(_SEED, brng=brng)
81 changes: 81 additions & 0 deletions benchmarks/benchmarks/bench_array_params.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
# Copyright (c) 2026, Intel Corporation
#
# Redistribution and use in source and binary forms, with or without
# modification, are permitted provided that the following conditions are met:
#
# * Redistributions of source code must retain the above copyright notice,
# this list of conditions and the following disclaimer.
# * Redistributions in binary form must reproduce the above copyright
# notice, this list of conditions and the following disclaimer in the
# documentation and/or other materials provided with the distribution.
# * Neither the name of Intel Corporation nor the names of its contributors
# may be used to endorse or promote products derived from this software
# without specific prior written permission.
#
# THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
# AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
# IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
# DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
# FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
# DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
# SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
# CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
# OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
# OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

"""Benchmarks for array-valued distribution parameters.

One parameter value per output element takes a separate code path from
scalar parameters, so these are tracked apart from bench_continuous.py and
bench_discrete.py.
"""

import numpy as np

from ._utils import _SEED, _make_state

# Smaller than _utils._SIZES: the per-element path is much slower than a
# scalar fill.
_ARRAY_SIZES = [1_000, 100_000]


class ArrayParams:
"""Distributions called with one parameter value per output element."""

params = [_ARRAY_SIZES]
param_names = ["size"]

def setup(self, size):
self.rs = _make_state()
rng = np.random.default_rng(_SEED)
self.loc = rng.standard_normal(size)
self.scale = rng.uniform(0.5, 2.0, size)
self.high = self.loc + self.scale
self.shape = rng.uniform(0.5, 5.0, size)
self.lam = rng.uniform(1.0, 100.0, size)
self.n = rng.integers(1, 100, size)
self.p = rng.uniform(0.1, 0.9, size)
self.rs.normal(self.loc, self.scale)
self.rs.uniform(self.loc, self.high)
self.rs.exponential(self.scale)
self.rs.standard_gamma(self.shape)
self.rs.poisson(self.lam)
self.rs.binomial(self.n, self.p)

def time_normal(self, size):
self.rs.normal(self.loc, self.scale)

def time_uniform(self, size):
self.rs.uniform(self.loc, self.high)

def time_exponential(self, size):
self.rs.exponential(self.scale)

def time_standard_gamma(self, size):
self.rs.standard_gamma(self.shape)

def time_poisson(self, size):
self.rs.poisson(self.lam)

def time_binomial(self, size):
self.rs.binomial(self.n, self.p)
Loading
Loading