Repository navigation
Conversation
FUTEX_REQUEUE moves a parked waiter into another bucket's chain, but the waiter went on sleeping, and testing waiter.woken, under the lock of the bucket it entered with. A wake at the new address takes only the new bucket's lock, so nothing ordered the two: - A waiter leaving on a timeout or a signal looked for itself in the chain after the wake had unlinked it, and looped forever holding the new bucket's lock. Every later futex call on that bucket parked behind it. - The waiter could return and destroy its condvar while the wake was still inside pthread_cond_signal on it. main took a host SIGBUS there, in _pthread_cond_updateval, in a 40000-round run. - A wake landing between the waiter's test of woken and its sleep was answered a polling quantum late. Requeue now records the bucket that holds the waiter in waiter.home and signals it, and the waiter re-parks under that bucket's lock before it tests woken or sleeps again. It takes the same lock before it dequeues, which is what makes the test of woken final and the condvar safe to destroy. The search loop that could not end is gone. tests/test-futex-requeue-race.c requeues a parked thread and then wakes it at the new word, alone, against a signal, and against the waiter's own timeout. FUTEX_WAIT_BITSET is the case this tree can reach: a requeue wakes plain FUTEX_WAIT in place until the next commit moves it to the bucket. On the parent the timeout case hangs in 2 of 5 runs. A build without the re-park fails the first case 3 of 3, and one without the lock before the dequeue fails the signal case. Refs sysprog21#378
Plain FUTEX_WAIT parked in os_sync_wait_on_address, and the only way to reach a thread parked there is to wake its address. That is also what FUTEX_WAKE does, so the waiter cannot tell a wake from anything else that needs it out of the wait. Ending the wait for a queued signal that way made a waiter a FUTEX_WAKE had already counted report EINTR, and made a signal the waiter could not claim, an ignored SIGCHLD among them, return 0 to it and to every sibling on the word. A bucket waiter has waiter.woken to tell the two apart, so plain FUTEX_WAIT now takes the path FUTEX_WAIT_BITSET always has, and a later commit adds the signal wake on top of it. FUTEX_REQUEUE moves a plain waiter instead of waking it in place, which is the path musl's pthread_cond_broadcast takes and why the previous commit comes first. tests/test-futex-requeue-race.c gains the plain case. The address-wait code stays, unreached, behind os_sync_wait_enabled. Cost, main against the tip of this branch, alternating, medians of 9: the bench-futex handoff reads 128.5 against 129.1 ns per round trip. bench-futex-contend with 4 threads reads 1.085 against 1.123 s, with ranges of 0.941 to 1.190 and 1.023 to 1.209, and with 8 threads 0.946 against 0.948 s. A relative timeout now overshoots about twice as far, 515 us rather than 264 us on a 1 ms wait, where Linux 6.12 under qemu overshoots 590 us. Refs sysprog21#378
A thread parked in FUTEX_WAIT or FUTEX_WAIT_BITSET is outside hv_vcpu_run and watches neither the wakeup pipe nor its condition variable, so it noticed a queued signal only when its 100 ms polling quantum ended: 43 ms from pthread_kill to the handler at the median, against 0.044 ms under Linux 6.12. Queueing a signal now kicks every other thread that has a pending signal it does not block. A parked thread publishes the mutex and condvar it sleeps on. The kick sets a flag and then signals under that mutex, and the thread tests the flag under the same mutex before it sleeps, so a kick that arrives ahead of the park is not lost. A waiter that FUTEX_REQUEUE moved publishes again under its new bucket's lock. A kicked thread that finds nothing to claim parks again, which is what an ignored signal or one a sibling took comes to. FUTEX_LOCK_PI and futex_waitv keep their quantum. tests/test-futex-signal-latency.c fails all four cases on the parent commit and passes here, and requires EINTR from each wait. Not in the tree: 40000 signals aimed at the entry of each wait were all answered within 0.19 ms, a FUTEX_WAKE racing a signal never left a counted waiter reporting EINTR in 40000 rounds, and an ignored SIGCHLD left both waiters on a word parked, as under Linux 6.12 in qemu. Refs sysprog21#378
Plain FUTEX_WAIT no longer parks in os_sync_wait_on_address, so nothing reaches futex_os_sync_wait, the per-bucket os_sync_waiters census that told a wake whether to drain the Darwin queue, or futex_wake_topup_osync, which every bucket-walking wake called to do it. They go, with the SDK and runtime probes that gated them. futex_wake, futex_requeue and futex_wake_op return what the bucket walk counted, and futex_should_block loses the out-parameter only the address wait used. No behavior changes: the gate these sat behind has been off since the commit that moved FUTEX_WAIT to the bucket. The note on why the address wait cannot serve FUTEX_WAIT moves to the FUTEX_WAIT case in sys_futex. FUTEX_OS_SYNC_POLL_CAP_NS keeps its name. It is the quantum of every futex wait and the futexdeadline proof names it. Refs sysprog21#378
Three tests wait for a thread to exit with
while (tid != 0)
raw_futex_wait_cleartid(&tid, tid);
which reads the volatile tid twice. A thread that exits between the two
reads clears it and sends its wake, the second read then passes 0 as
the value to wait on, and the wait parks on a word that already holds 0
with no wake left to come.
raw_wait_cleartid reads it once per pass. The five sites in
test-thread, test-stress and test-credentials use it.
This is in the tests, not in elfuse: test-thread's five cases run in a
loop stall on main and on Linux 6.12 under qemu as well, within 1000 to
140000 iterations. With the tid read once, six concurrent loops ran
200000 iterations each without a stall. It timed out one make check run
of this branch, which is why it is fixed here.
A round of test-futex-requeue-race sends SIGUSR1 with tgkill and then issues FUTEX_WAKE, and it read neither result as an error. A failed tgkill left the wake to return 1 with errno 0, so a signal round passed with no signal sent. A wake that returned -1 matched none of the three checks, which test only for 0 and 1, so the round passed with no wake. Both results now end the test with a message.
The shim answers FUTEX_WAKE at EL1 while the bucket's published waiter count is zero, so the concurrent wake row, which had no waiter, never reached futex_wake. It reported 36.6 ns per wake against 35.6 ns for the one-thread wake-nowaiter row and could not see the bucket lock it was written to measure. Each waker's word now carries one waiter parked with FUTEX_WAIT_BITSET on a bit the wakes do not ask for. The count is nonzero, so the call goes to the host, which takes the bucket lock, walks one entry and wakes nobody. The waiter sits on the waker's own word, so the bench needs neither the bucket hash nor the table size. A one-thread row on the same path replaces wake-nowaiter as the baseline, and the concurrent row is renamed wake-nomatch-concurrent. Measured over six runs: wake-nomatch 763 to 776 ns, at the HVC floor, and wake-nomatch-concurrent 300 to 319 ns per wake over 8 threads.
The thread_entry comment named three scans that walk the table without thread_lock. thread_kick_futex_waiters is a fourth, loading active, blocked and tpending.pending on the same terms, and a change to slot reuse has to account for it. The header of test-osync-requeue described the Darwin address-wait queue in a sentence that read as current behavior. It states the condition the test catches instead.
The thread_entry comment read as if each lock-free scan loads one field after active. thread_kick_futex_waiters loads two, tpending.pending and then blocked, so the sentence was wrong for the scan it had just added.
xalestar
marked this pull request as draft
October 9, 2026 18:00
jserv
reviewed
Oct 10, 2026
| * between the 100 ms os_sync poll cap (FUTEX_OS_SYNC_POLL_CAP_NS) and the | ||
| * pass threshold below, so a stranded waiter (regression: requeue no longer | ||
| * degrades to a wake at the source address) surfaces near the cap, far | ||
| * between the 100 ms polling quantum (FUTEX_OS_SYNC_POLL_CAP_NS) and the |
Contributor
There was a problem hiding this comment.
FUTEX_OS_SYNC_POLL_CAP_NS now names a path this PR deletes; it is the polling quantum every bucket sleep uses (futex_quantum_deadline, the spin guard in futex_wait_inner). Worth a follow-up renaming it to something like FUTEX_POLL_QUANTUM_NS, and test-osync-requeue with it, so the next reader does not go looking for an os_sync wait that no longer exists.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #378: the futex part. select and epoll_wait went in with #386.
A thread parked in FUTEX_WAIT or FUTEX_WAIT_BITSET is outside hv_vcpu_run and watches neither the wakeup pipe nor its condition variable, so it noticed a queued signal only when its 100 ms polling quantum ended. Queueing a signal now kicks the parked threads that can take it. Four commits, in the order they depend on each other, and a fifth that fixes a test.
1. Park a requeued futex waiter under its new lock. A bug that is already on
main, and that the next commit would put on musl'spthread_cond_broadcastpath. FUTEX_REQUEUE moved a waiter into another bucket's chain while the waiter went on sleeping, and testingwoken, under the lock of the bucket it entered with. A wake at the new address takes only the new bucket's lock, so a waiter leaving on a timeout or a signal could search the chain for itself after the wake had unlinked it and loop forever holding that bucket's lock, or return and destroy its condvar while the wake was still signalling it. Requeue now records the waiter's bucket inwaiter.homeand signals it, and the waiter re-parks under that lock before it testswokenor sleeps, and takes it before it dequeues.2. Park plain FUTEX_WAIT on the bucket condvar. Plain FUTEX_WAIT parked in
os_sync_wait_on_address, and the only way to reach a thread there is to wake its address, which it cannot tell from a FUTEX_WAKE. A first version of this change kicked it that way and was wrong twice over: a waiter a FUTEX_WAKE had already counted reported EINTR (2604 of 3000 rounds of a wake racing a signal), and a signal the waiter could not claim, a default-ignored SIGCHLD among them, returned 0 to it and to every sibling on the word. A bucket waiter haswaiter.wokento tell a wake from a kick, so plain FUTEX_WAIT takes the path FUTEX_WAIT_BITSET always has.3. Wake a futex wait on a queued signal. A parked thread publishes the mutex and condvar it sleeps on.
futex_kicksets a flag and then signals under that mutex, and the thread tests the flag under the same mutex before it sleeps, so a kick that arrives ahead of the park is not lost. A kicked thread that finds nothing to claim parks again. FUTEX_LOCK_PI and futex_waitv keep their quantum.4. Remove the unreached address-wait futex path. No behavior change: the code commit 2 left behind its gate.
5. Read a tid once before waiting for it to clear. Unrelated to #378, and here because it timed out one
make checkrun of this branch.test-thread,test-stressandtest-credentialswait for a thread to exit withwhile (tid != 0) raw_futex_wait_cleartid(&tid, tid);, which reads the volatile tid twice. A thread that exits between the two reads leaves the wait parked on a word that already holds 0. The five cases oftest-threadrun in a loop stall onmainand on Linux 6.12 under qemu as well, within 1000 to 140000 iterations. With the tid read once, six concurrent loops ran 200000 iterations each without a stall.Reproduction
tests/test-futex-signal-latency.cparks a sibling thread in each wait, sends SIGUSR1 with pthread_kill, and takes the median signal-to-handler gap over 8 rounds against a 20 ms bound. It also requires EINTR from the wait and under 5% CPU from the waiter.The Linux column is the figure from #378.
tests/test-futex-requeue-race.crequeues a parked thread to a word in another bucket and then wakes it there: alone, against a signal, and against the waiter's own timeout, for both waits. A wake that counted the waiter must find it returning 0 within 20 ms, and a wake that found nobody must find it reporting EINTR or ETIMEDOUT. Onmainthe FUTEX_WAIT_BITSET timeout case hangs in 2 of 5 runs. A build of this branch without the re-park at the top of the wait loop fails the first case 3 of 3, and one without the lock before the dequeue fails the signal case.Probes that are not in the tree:
main.Cost
mainagainst this branch, alternating, medians of 9 runs:bench-futexhandoff, per round tripbench-futex-contend 1000000 4bench-futex-contend 300000 8A relative FUTEX_WAIT timeout overshoots about twice as far as it did, 515 us where it was 264 us on a 1 ms wait. Linux 6.12 under qemu overshoots 590 us on the same wait. A requeue now wakes the host thread of each waiter it moves so that the waiter can change locks; the waiter does not return to the guest. No benchmark here issues a requeue, so that cost is unmeasured.
Validation
Rebased on
b57f367.make check,make check-contracts,make check-tsan,make verify,make lint,make check-format: exit 0.make check-contracts,make check-tsan,make verifyandmake lintran before commit 5, which touches only tests.make verify-mutants MUTANT_TARGET=futexdeadline: 4 mutations, 4 caught, all by exhausting the provertests/test-matrix.sh elfuse-aarch64: 299 passed, 0 failed, 10 skipped. Other runs failtest-sigioalone, which is test-sigio loses SIGURG on a TCP out-of-band byte about half the time #402 and fails 14 of 30 runs onmain.test-futex-requeue-raceand ten futex and thread tests.Found on the way and filed separately: #417, #418, #419.