Feature or enhancement
Proposal:
Several functions in longobject.c (v_lshift, v_rshift, and x_divrem) narrowed a loop carry value to digit or sdigit, then widened it again on the next iteration.
On AArch64, that 64-to-32-to-64 conversion inserts an extra mov on the loop-carried critical path. Keeping the carry at two-digit width until the function returns drops that mov and shortens the carry chain. Results are unchanged for valid limbs.
Use v_lshift on AArch64 as an example. The code inside the loop is:
ldr w0, [x4, x2, lsl #2] ; a[i]
mov w3, w3 ; carry chain
lsl x0, x0, x24 ; a[i] << d
orr x0, x0, x3 ; carry chain
and w1, w0, #0x3fffffff
ubfx x3, x0, #30, #32 ; carry chain
str w1, [x26, x2, lsl #2]
add x2, x2, #1
cmp x25, x2
b.ne
removing the narrowing, the code will be optimized to:
ldr w0, [x4, x2, lsl #2] ; a[i]
lsl x0, x0, x24 ; a[i] << d
orr x0, x0, x3 ; carry chain
and w1, w0, #0x3fffffff
str w1, [x26, x2, lsl #2]
add x2, x2, #1
lsr x3, x0, #30 ; carry chain
cmp x25, x2
b.ne
The instruction count on the loop carried chain is reduced from 3 to 2.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Feature or enhancement
Proposal:
Several functions in
longobject.c(v_lshift,v_rshift, andx_divrem) narrowed a loop carry value todigitorsdigit, then widened it again on the next iteration.On AArch64, that 64-to-32-to-64 conversion inserts an extra
movon the loop-carried critical path. Keeping the carry at two-digit width until the function returns drops thatmovand shortens the carry chain. Results are unchanged for valid limbs.Use
v_lshifton AArch64 as an example. The code inside the loop is:removing the narrowing, the code will be optimized to:
The instruction count on the loop carried chain is reduced from 3 to 2.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response