Skip to content

v3.0.0: Exact square roots and faster wide-type paths - #11

Merged
GeEom merged 4 commits into
mainfrom
wide_types
Sep 3, 2026
Merged

GeEom merged 4 commits into
mainfrom
wide_types

Conversation

@GeEom

@GeEom GeEom commented Sep 3, 2026

Copy link
Copy Markdown
Owner

Implements the findings from a downstream audit of the 128-bit types, plus the corrections they exposed.

  • Bit-identical speedups on every layout with at least 8 integer bits: raw-representation division in exp and sinh_cosh, reciprocal reductions in sin_cos and exp, a bit-scan normalisation in ln, wrapping Horner evaluation. I64F64: sin_cos 2x, exp 3.9x, sinh_cosh 3.3x. Also fixes exp above 2^31·ln 2 on wide types, sin_cos within π of MAX, and exp/sinh_cosh on types with fewer than 8 integer bits.
  • pow, fallible, as exp(exponent · ln(base)).
  • Accuracy: sqrt is now correctly rounded at every magnitude (it was unusable below 10^-11 on I64F64), CORDIC vectoring closes with one division (atan2 3.2x, ln 1.9x, hyperbolic kernel delivers 64 bits instead of 54), asinh/acosh use the logarithmic form above 1, and exp/sinh_cosh gain a series tier for 40+ fractional bits.

Accuracy gate against the 2.1.0 baseline: 30 improve, 10 unchanged, none regress. CordicNumber gained required methods and the kernels return the residual vector, hence the major version.

Each commit builds and passes tests, clippy and the no-panic check on its own.

A 128-bit fixed-point division costs about twelve multiplications, and
forming a divisor like 182 in T overflows types with few integer bits.
CordicNumber gains div_int and mul_int, which work on the raw
representation, plus wrapping_mul/add/sub and checked_int_log2; round
and to_i32 now saturate instead of wrapping at MAX.

exp and sinh_cosh divide by their Taylor divisors with div_int. sin_cos
and exp estimate their reductions with stored reciprocals when the
integer bits do not outnumber the fractional bits, and correct the
estimate once in either direction. ln normalises from the leading bit
instead of a shift loop. Horner evaluation wraps, since every
intermediate stays below 1 for the tables in use.

Results are bit-identical on every layout with at least 8 integer bits
except where the old code was wrong: exp above 2^31·ln 2 on 64-bit and
wider types wrapped in to_i32, and angles within π of MAX saturated the
reduction to sin_cos(0). I4F12, I4F60 and I3F125 gain working exp and
sinh_cosh. I64F64: sin_cos 52 -> 26 ns, exp 159 -> 41 ns, sinh_cosh
175 -> 54 ns, and ln normalisation is O(1).

Adds I64F64 benchmarks, instantiates the no-panic binary on three
layouts, and raises the no-panic floor to 0.1.37.
pow(base, exponent) = exp(exponent · ln(base)), fallible on a negative
base or a zero base with a negative exponent; 0^0 is 1 and exp's
saturation applies. ln is split into its domain check and ln_positive
so pow can skip the check.

The logarithm's rounding error is multiplied by the exponent, so the
relative error grows with |exponent|: about 1e-4 at I16F16 for
exponent 1.5. 219 ns on I64F64, 29 ns on I16F16.
@codecov-commenter

codecov-commenter commented Sep 3, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 98.62385% with 3 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/ops/exponential.rs 96.82% 2 Missing ⚠️
src/kernel/cordic.rs 95.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

sqrt refined a seed with an absolute 2^-32 convergence epsilon, so on
I64F64 it was exact above 100 and returned relative errors above 1
below 10^-13. It now returns the representable value nearest the true
root at every magnitude: the integer root of the raw bits, plus one
Newton step and a 256-bit remainder test on 128-bit types. Verified
against the fixed crate's rounded-down root on 14 layouts. I64F64 sqrt
114 -> 42 ns above 1 and 180 -> 24 ns below.

CORDIC vectoring stops after ⌈frac/3⌉ shift stages and closes the
residual with one division; the skipped stages only added table angles
of exactly 2^-i. atan2 212 -> 66 ns and ln 240 -> 127 ns on I64F64, and
the hyperbolic kernel delivers 64 bits instead of 54.

asinh and acosh use the logarithmic form above 1 and 1.5, where the
atanh argument reduction tripled the error at every step. exp and
sinh_cosh gain a third series tier for 40 or more fractional bits
(degree 19, and 21/22), taking their I64F64 relative error from 1e-13
to 2e-18; the previous comment understated the degree-12 truncation by
400x.

Accuracy gate against the 2.1.0 baseline: 30 improve, 10 unchanged,
none regress. asinh I16F16 mean 6.44e-4 -> 2.42e-5, sqrt I32F32
2.70e-12 -> 1.37e-12, atan I16F16 1.50e-5 -> 7.79e-6. CordicNumber
gained required methods and the kernels now return the residual
vector, so this is a major version.
@GeEom
GeEom force-pushed the wide_types branch 2 times, most recently from 32bdd1b to b16d680 Compare September 3, 2026 08:53
Codecov's ignore pattern did not match the bench file, which sits
directly under benches/, and the no-panic binary is compiled only under
its feature, so neither can be covered by tests; both are excluded by
folder.
@GeEom
GeEom merged commit a74e65e into main Sep 3, 2026
12 checks passed
@GeEom
GeEom deleted the wide_types branch September 3, 2026 09:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants