Skip to content

Reduce Boost dependency footprint - #3405

Open
andrjohns wants to merge 10 commits into
developfrom
boost-repl-stan
Open

Reduce Boost dependency footprint#3405
andrjohns wants to merge 10 commits into
developfrom
boost-repl-stan

Conversation

@andrjohns

Copy link
Copy Markdown
Contributor

Submission Checklist

  • Run unit tests: ./runTests.py src/test/unit
  • Run cpplint: make cpplint
  • Declare copyright holder and open-source license: see below

Summary

Related to this Math PR, this PR removes/replaces non-math Boost includes with std-equivalents.

The following includes are removed:

  • boost/random: Replaced with <random> equivalents, expected test values also updated
  • boost/algorithm: Only basic string manip was needed, these have been implemented under string_utils.hpp instead
  • boost/regex: Only used to check whether variable name was valid, pattern was very simple so just used std::isalpha & std::isalnum
  • boost/accumulators: Replaced with Eigen functions

Now the only non-math Boost include is boost/random/mixmax. The boost/accumulators removal in particular is a big win, since that pulled in both boost/mpl and boost/fusion - these two libraries alone are ~20mb!

Let me know if it would be easier to review if I split the changes into separate PRs

Intended Effect

Reduce amount of Boost libraries/headers needed by Stan

How to Verify

Side Effects

Changing the rng usage from boost -> std will return different values from same seed.

Documentation

N/A

Copyright and Licensing

Please list the copyright holder for the work you are submitting (this will be you or your assignee, such as a university or company): Andrew Johnson

By submitting this pull request, the copyright holder is agreeing to license the submitted work under the following licenses:

@stan-buildbot

Copy link
Copy Markdown
Contributor

Name Old Result New Result Ratio Performance change( 1 - new / old )
gp_regr/gp_regr.stan 0.14 0.11 1.26 20.69% faster
gp_regr/gen_gp_data.stan 0.03 0.02 1.11 9.99% faster
arK/arK.stan 1.78 1.66 1.07 6.88% faster
eight_schools/eight_schools.stan 0.06 0.05 1.29 22.39% faster
low_dim_gauss_mix_collapse/low_dim_gauss_mix_collapse.stan 9.65 10.61 0.91 -9.9% slower
pkpd/one_comp_mm_elim_abs.stan 21.89 18.2 1.2 16.88% faster
pkpd/sim_one_comp_mm_elim_abs.stan 0.3 0.29 1.05 4.82% faster
sir/sir.stan 78.79 74.77 1.05 5.1% faster
gp_pois_regr/gp_pois_regr.stan 2.57 2.47 1.04 3.92% faster
low_dim_gauss_mix/low_dim_gauss_mix.stan 3.1 3.14 0.99 -1.3% slower
irt_2pl/irt_2pl.stan 6.14 4.87 1.26 20.68% faster
arma/arma.stan 0.38 0.33 1.14 12.38% faster
garch/garch.stan 0.54 0.45 1.19 16.06% faster
low_dim_corr_gauss/low_dim_corr_gauss.stan 0.01 0.01 1.01 1.3% faster
performance.compilation 301.9 284.57 1.06 5.74% faster
Mean result: 1.1098053193097879

Jenkins Console Log
Blue Ocean
Commit hash: fb55e5ae8f992ecca9977e062d78369635037345


Machine information No LSB modules are available. Distributor ID: Ubuntu Description: Ubuntu 20.04.3 LTS Release: 20.04 Codename: focal

CPU:
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Byte Order: Little Endian
Address sizes: 46 bits physical, 48 bits virtual
CPU(s): 80
On-line CPU(s) list: 0-79
Thread(s) per core: 2
Core(s) per socket: 20
Socket(s): 2
NUMA node(s): 2
Vendor ID: GenuineIntel
CPU family: 6
Model: 85
Model name: Intel(R) Xeon(R) Gold 6148 CPU @ 2.40GHz
Stepping: 4
CPU MHz: 3351.069
CPU max MHz: 3700.0000
CPU min MHz: 1000.0000
BogoMIPS: 4800.00
Virtualization: VT-x
L1d cache: 1.3 MiB
L1i cache: 1.3 MiB
L2 cache: 40 MiB
L3 cache: 55 MiB
NUMA node0 CPU(s): 0,2,4,6,8,10,12,14,16,18,20,22,24,26,28,30,32,34,36,38,40,42,44,46,48,50,52,54,56,58,60,62,64,66,68,70,72,74,76,78
NUMA node1 CPU(s): 1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,31,33,35,37,39,41,43,45,47,49,51,53,55,57,59,61,63,65,67,69,71,73,75,77,79
Vulnerability Gather data sampling: Mitigation; Microcode
Vulnerability Itlb multihit: KVM: Mitigation: Split huge pages
Vulnerability L1tf: Mitigation; PTE Inversion; VMX conditional cache flushes, SMT vulnerable
Vulnerability Mds: Mitigation; Clear CPU buffers; SMT vulnerable
Vulnerability Meltdown: Mitigation; PTI
Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
Vulnerability Reg file data sampling: Not affected
Vulnerability Retbleed: Mitigation; IBRS
Vulnerability Spec rstack overflow: Not affected
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; IBRS; IBPB conditional; STIBP conditional; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected
Vulnerability Srbds: Not affected
Vulnerability Tsx async abort: Mitigation; Clear CPU buffers; SMT vulnerable
Vulnerability Vmscape: Mitigation; IBPB before exit to userspace
Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 invpcid_single pti intel_ppin ssbd mba ibrs ibpb stibp tpr_shadow vnmi flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req pku ospke md_clear flush_l1d arch_capabilities

G++:
g++ (Ubuntu 9.4.0-1ubuntu1~20.04) 9.4.0
Copyright (C) 2019 Free Software Foundation, Inc.
This is free software; see the source for copying conditions. There is NO
warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.

Clang:
clang version 10.0.0-4ubuntu1
Target: x86_64-pc-linux-gnu
Thread model: posix
InstalledDir: /usr/bin

@WardBrian

Copy link
Copy Markdown
Member

boost/random/mixmax

My coworker @DiamonDinoia has encouraged me to stop using this generator as well (can be a separate PR)

@andrjohns

Copy link
Copy Markdown
Contributor Author

boost/random/mixmax

My coworker @DiamonDinoia has encouraged me to stop using this generator as well (can be a separate PR)

Interesting! What's the recommended replacement?

If we're looking at changing the RNG it would probably be best if we did the boost/random -> std random and the RNG change in the same release, since both are going to change any fixed expectations for downstream users.

With the feature freeze coming up did you want to do the RNG change this release or hold off on the boost/random changes until the next release?

@DiamonDinoia

DiamonDinoia commented Aug 12, 2026

Copy link
Copy Markdown

@vigna has a great page to guide the choice: https://prng.di.unimi.it/ look for how to choose a PRNG.

If I were to implement the PRNG from scratch I would implement one of xoshiro256++/xoshiro256** which are 10 liners. For example in c++: https://github.com/DiamonDinoia/simdrng/blob/main/include/simdrng/xoshiro_scalar.hpp

If I can add a dependency I would use ChaCha20, AES in counter mode relies on having AES implemented in hw which not all platforms have. And I think Stan aims to support many no?

I personally implemented those into my package here https://github.com/DiamonDinoia/simdrng I am still developing and cleaning up a few things. But I wanted to take the change to make a product placement.

This package tests some generators and share some findings: https://github.com/vigna/modlin-rs

@DiamonDinoia

Copy link
Copy Markdown

If you implement any of those generators, I am happy to help review.

@SteveBronder SteveBronder left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Few changes and I need to look over the ring buffer impl still

Comment on lines +87 to -99
double variance = (y.array() - y.mean()).matrix().squaredNorm() / y.size();

accumulator_set<double, stats<variance>> acc;
for (int n = 0; n < y.size(); ++n) {
acc(y(n));
}

acov = acov.array() * boost::accumulators::variance(acc);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should use a little variance function like the below that uses welford's algorithm so we can be a bit more precise

https://godbolt.org/z/Porcq1z8W

* @return the tokens found in `s`
*/
inline std::vector<std::string> split(const std::string& s,
const std::string& delims,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bit of code golf, but I think you can shorten this a good bit

https://godbolt.org/z/hGTe59ber

Comment on lines +68 to +69
if (i > 0)
result += sep;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
if (i > 0)
result += sep;
if (i > 0) {
result += sep;
}

*/
inline std::string join(const std::vector<std::string>& parts,
const std::string& sep) {
std::string result;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You will be allocating at least parts.size() sep + parts[0]

Suggested change
std::string result;
std::string result;
result.reserve(parts[0].size() + parts.size());

Comment on lines +94 to +95
inline void replace_first(std::string& s, const std::string& target,
const std::string& replacement) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
inline void replace_first(std::string& s, const std::string& target,
const std::string& replacement) {
inline void replace_first(std::string_view& s, const std::string_view& target,
const std::string_view& replacement) {

Comment thread src/stan/mcmc/chains.hpp
Comment on lines +62 to +65
double mx = x.head(M).mean();
double my = y.head(M).mean();
return ((x.head(M).array() - mx) * (y.head(M).array() - my)).sum()
/ (M - 1.0);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Like the variance I think it would be nice to add a welford style estimator for covariance here

https://godbolt.org/z/E7v9dbj59

Comment thread src/stan/mcmc/chains.hpp
Comment on lines +86 to +88
std::vector<double> probs_vec(probs.data(), probs.data() + probs.size());
std::vector<double> q = stan::math::quantile(x, probs_vec);
return Eigen::Map<const Eigen::VectorXd>(q.data(), q.size());

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs to be cast to a vector to use quantile? Could we update that function to also accept Eigen vectors?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants