Skip to content

Wake a futex wait on a queued signal - #420

Draft
xalestar wants to merge 9 commits into
sysprog21:mainfrom
xalestar:futex-signal-wake
Draft

xalestar wants to merge 9 commits into
sysprog21:mainfrom
xalestar:futex-signal-wake

Conversation

@xalestar

@xalestar xalestar commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #378: the futex part. select and epoll_wait went in with #386.

A thread parked in FUTEX_WAIT or FUTEX_WAIT_BITSET is outside hv_vcpu_run and watches neither the wakeup pipe nor its condition variable, so it noticed a queued signal only when its 100 ms polling quantum ended. Queueing a signal now kicks the parked threads that can take it. Four commits, in the order they depend on each other, and a fifth that fixes a test.

1. Park a requeued futex waiter under its new lock. A bug that is already on main, and that the next commit would put on musl's pthread_cond_broadcast path. FUTEX_REQUEUE moved a waiter into another bucket's chain while the waiter went on sleeping, and testing woken, under the lock of the bucket it entered with. A wake at the new address takes only the new bucket's lock, so a waiter leaving on a timeout or a signal could search the chain for itself after the wake had unlinked it and loop forever holding that bucket's lock, or return and destroy its condvar while the wake was still signalling it. Requeue now records the waiter's bucket in waiter.home and signals it, and the waiter re-parks under that lock before it tests woken or sleeps, and takes it before it dequeues.

2. Park plain FUTEX_WAIT on the bucket condvar. Plain FUTEX_WAIT parked in os_sync_wait_on_address, and the only way to reach a thread there is to wake its address, which it cannot tell from a FUTEX_WAKE. A first version of this change kicked it that way and was wrong twice over: a waiter a FUTEX_WAKE had already counted reported EINTR (2604 of 3000 rounds of a wake racing a signal), and a signal the waiter could not claim, a default-ignored SIGCHLD among them, returned 0 to it and to every sibling on the word. A bucket waiter has waiter.woken to tell a wake from a kick, so plain FUTEX_WAIT takes the path FUTEX_WAIT_BITSET always has.

3. Wake a futex wait on a queued signal. A parked thread publishes the mutex and condvar it sleeps on. futex_kick sets a flag and then signals under that mutex, and the thread tests the flag under the same mutex before it sleeps, so a kick that arrives ahead of the park is not lost. A kicked thread that finds nothing to claim parks again. FUTEX_LOCK_PI and futex_waitv keep their quantum.

4. Remove the unreached address-wait futex path. No behavior change: the code commit 2 left behind its gate.

5. Read a tid once before waiting for it to clear. Unrelated to #378, and here because it timed out one make check run of this branch. test-thread, test-stress and test-credentials wait for a thread to exit with while (tid != 0) raw_futex_wait_cleartid(&tid, tid);, which reads the volatile tid twice. A thread that exits between the two reads leaves the wait parked on a word that already holds 0. The five cases of test-thread run in a loop stall on main and on Linux 6.12 under qemu as well, within 1000 to 140000 iterations. With the tid read once, six concurrent loops ran 200000 iterations each without a stall.

Reproduction

tests/test-futex-signal-latency.c parks a sibling thread in each wait, sends SIGUSR1 with pthread_kill, and takes the median signal-to-handler gap over 8 rounds against a 20 ms bound. It also requires EINTR from the wait and under 5% CPU from the waiter.

signal to handler, median before after Linux 6.12
FUTEX_WAIT, timeout 42.6 ms 0.064 ms 0.044 ms
FUTEX_WAIT, no timeout 43.2 ms 0.076 ms 0.044 ms
FUTEX_WAIT_BITSET, timeout 32.6 ms 0.074 ms
FUTEX_WAIT_BITSET, no timeout 38.5 ms 0.081 ms

The Linux column is the figure from #378.

tests/test-futex-requeue-race.c requeues a parked thread to a word in another bucket and then wakes it there: alone, against a signal, and against the waiter's own timeout, for both waits. A wake that counted the waiter must find it returning 0 within 20 ms, and a wake that found nobody must find it reporting EINTR or ETIMEDOUT. On main the FUTEX_WAIT_BITSET timeout case hangs in 2 of 5 runs. A build of this branch without the re-park at the top of the wait loop fails the first case 3 of 3, and one without the lock before the dequeue fails the signal case.

Probes that are not in the tree:

  • 40000 signals aimed 0 to 40 us after a thread announces each wait: all answered within 0.19 ms.
  • A FUTEX_WAKE racing a signal on one word, 40000 rounds: never a counted waiter reporting EINTR, never a 0 without a wake.
  • A process-directed signal with two waiters on a word: exactly one returns EINTR. A signal blocked in the waiter, a SIG_IGN signal and a default SIGCHLD leave it parked. Linux 6.12 under qemu gives the same four answers.
  • Chained and repeated requeues among four words with signals, timeouts and wakes racing, about 1.4 million rounds: no hang, no wrong return value, no wake delayed past 20 ms. The signal and timeout variants hold on Linux 6.12 under qemu too, and hang on main.

Cost

main against this branch, alternating, medians of 9 runs:

main this branch
bench-futex handoff, per round trip 128.5 ns 129.1 ns
bench-futex-contend 1000000 4 1.085 s (0.941 to 1.190) 1.123 s (1.023 to 1.209)
bench-futex-contend 300000 8 0.946 s 0.948 s

A relative FUTEX_WAIT timeout overshoots about twice as far as it did, 515 us where it was 264 us on a 1 ms wait. Linux 6.12 under qemu overshoots 590 us on the same wait. A requeue now wakes the host thread of each waiter it moves so that the waiter can change locks; the waiter does not return to the guest. No benchmark here issues a requeue, so that cost is unmeasured.

Validation

Rebased on b57f367.

  • make check, make check-contracts, make check-tsan, make verify, make lint, make check-format: exit 0. make check-contracts, make check-tsan, make verify and make lint ran before commit 5, which touches only tests.
  • make verify-mutants MUTANT_TARGET=futexdeadline: 4 mutations, 4 caught, all by exhausting the prover
  • tests/test-matrix.sh elfuse-aarch64: 299 passed, 0 failed, 10 skipped. Other runs fail test-sigio alone, which is test-sigio loses SIGURG on a TCP out-of-band byte about half the time #402 and fails 14 of 30 runs on main.
  • Each of the first three commits builds on its own and passes test-futex-requeue-race and ten futex and thread tests.
  • The qemu matrix lane was not run. The two new tests have not been run under Linux, only the probes above.

Found on the way and filed separately: #417, #418, #419.

FUTEX_REQUEUE moves a parked waiter into another bucket's chain, but
the waiter went on sleeping, and testing waiter.woken, under the lock
of the bucket it entered with. A wake at the new address takes only the
new bucket's lock, so nothing ordered the two:

- A waiter leaving on a timeout or a signal looked for itself in the
  chain after the wake had unlinked it, and looped forever holding the
  new bucket's lock. Every later futex call on that bucket parked
  behind it.
- The waiter could return and destroy its condvar while the wake was
  still inside pthread_cond_signal on it. main took a host SIGBUS
  there, in _pthread_cond_updateval, in a 40000-round run.
- A wake landing between the waiter's test of woken and its sleep was
  answered a polling quantum late.

Requeue now records the bucket that holds the waiter in waiter.home and
signals it, and the waiter re-parks under that bucket's lock before it
tests woken or sleeps again. It takes the same lock before it
dequeues, which is what makes the test of woken final and the condvar
safe to destroy. The search loop that could not end is gone.

tests/test-futex-requeue-race.c requeues a parked thread and then
wakes it at the new word, alone, against a signal, and against the
waiter's own timeout. FUTEX_WAIT_BITSET is the case this tree can
reach: a requeue wakes plain FUTEX_WAIT in place until the next commit
moves it to the bucket. On the parent the timeout case hangs in 2 of 5
runs. A build without the re-park fails the first case 3 of 3, and one
without the lock before the dequeue fails the signal case.

Refs sysprog21#378
Plain FUTEX_WAIT parked in os_sync_wait_on_address, and the only way to
reach a thread parked there is to wake its address. That is also what
FUTEX_WAKE does, so the waiter cannot tell a wake from anything else
that needs it out of the wait. Ending the wait for a queued signal that
way made a waiter a FUTEX_WAKE had already counted report EINTR, and
made a signal the waiter could not claim, an ignored SIGCHLD among
them, return 0 to it and to every sibling on the word.

A bucket waiter has waiter.woken to tell the two apart, so plain
FUTEX_WAIT now takes the path FUTEX_WAIT_BITSET always has, and a later
commit adds the signal wake on top of it. FUTEX_REQUEUE moves a plain
waiter instead of waking it in place, which is the path musl's
pthread_cond_broadcast takes and why the previous commit comes first.
tests/test-futex-requeue-race.c gains the plain case.

The address-wait code stays, unreached, behind os_sync_wait_enabled.

Cost, main against the tip of this branch, alternating, medians of 9:
the bench-futex handoff reads 128.5 against 129.1 ns per round trip.
bench-futex-contend with 4 threads reads 1.085 against 1.123 s, with
ranges of 0.941 to 1.190 and 1.023 to 1.209, and with 8 threads 0.946
against 0.948 s. A relative timeout now overshoots about twice as far,
515 us rather than 264 us on a 1 ms wait, where Linux 6.12 under qemu
overshoots 590 us.

Refs sysprog21#378
A thread parked in FUTEX_WAIT or FUTEX_WAIT_BITSET is outside
hv_vcpu_run and watches neither the wakeup pipe nor its condition
variable, so it noticed a queued signal only when its 100 ms polling
quantum ended: 43 ms from pthread_kill to the handler at the median,
against 0.044 ms under Linux 6.12.

Queueing a signal now kicks every other thread that has a pending
signal it does not block. A parked thread publishes the mutex and
condvar it sleeps on. The kick sets a flag and then signals under that
mutex, and the thread tests the flag under the same mutex before it
sleeps, so a kick that arrives ahead of the park is not lost. A waiter
that FUTEX_REQUEUE moved publishes again under its new bucket's lock.
A kicked thread that finds nothing to claim parks again, which is what
an ignored signal or one a sibling took comes to. FUTEX_LOCK_PI and
futex_waitv keep their quantum.

tests/test-futex-signal-latency.c fails all four cases on the parent
commit and passes here, and requires EINTR from each wait. Not in the
tree: 40000 signals aimed at the entry of each wait were all answered
within 0.19 ms, a FUTEX_WAKE racing a signal never left a counted
waiter reporting EINTR in 40000 rounds, and an ignored SIGCHLD left
both waiters on a word parked, as under Linux 6.12 in qemu.

Refs sysprog21#378
Plain FUTEX_WAIT no longer parks in os_sync_wait_on_address, so nothing
reaches futex_os_sync_wait, the per-bucket os_sync_waiters census that
told a wake whether to drain the Darwin queue, or
futex_wake_topup_osync, which every bucket-walking wake called to do
it. They go, with the SDK and runtime probes that gated them.
futex_wake, futex_requeue and futex_wake_op return what the bucket walk
counted, and futex_should_block loses the out-parameter only the
address wait used.

No behavior changes: the gate these sat behind has been off since the
commit that moved FUTEX_WAIT to the bucket. The note on why the address
wait cannot serve FUTEX_WAIT moves to the FUTEX_WAIT case in sys_futex.

FUTEX_OS_SYNC_POLL_CAP_NS keeps its name. It is the quantum of every
futex wait and the futexdeadline proof names it.

Refs sysprog21#378
cubic-dev-ai[bot]

This comment was marked as resolved.

Three tests wait for a thread to exit with

    while (tid != 0)
        raw_futex_wait_cleartid(&tid, tid);

which reads the volatile tid twice. A thread that exits between the two
reads clears it and sends its wake, the second read then passes 0 as
the value to wait on, and the wait parks on a word that already holds 0
with no wake left to come.

raw_wait_cleartid reads it once per pass. The five sites in
test-thread, test-stress and test-credentials use it.

This is in the tests, not in elfuse: test-thread's five cases run in a
loop stall on main and on Linux 6.12 under qemu as well, within 1000 to
140000 iterations. With the tid read once, six concurrent loops ran
200000 iterations each without a stall. It timed out one make check run
of this branch, which is why it is fixed here.
A round of test-futex-requeue-race sends SIGUSR1 with tgkill and then
issues FUTEX_WAKE, and it read neither result as an error. A failed
tgkill left the wake to return 1 with errno 0, so a signal round passed
with no signal sent. A wake that returned -1 matched none of the three
checks, which test only for 0 and 1, so the round passed with no wake.

Both results now end the test with a message.
The shim answers FUTEX_WAKE at EL1 while the bucket's published waiter
count is zero, so the concurrent wake row, which had no waiter, never
reached futex_wake. It reported 36.6 ns per wake against 35.6 ns for
the one-thread wake-nowaiter row and could not see the bucket lock it
was written to measure.

Each waker's word now carries one waiter parked with FUTEX_WAIT_BITSET
on a bit the wakes do not ask for. The count is nonzero, so the call
goes to the host, which takes the bucket lock, walks one entry and
wakes nobody. The waiter sits on the waker's own word, so the bench
needs neither the bucket hash nor the table size.

A one-thread row on the same path replaces wake-nowaiter as the
baseline, and the concurrent row is renamed wake-nomatch-concurrent.

Measured over six runs: wake-nomatch 763 to 776 ns, at the HVC floor,
and wake-nomatch-concurrent 300 to 319 ns per wake over 8 threads.
The thread_entry comment named three scans that walk the table without
thread_lock. thread_kick_futex_waiters is a fourth, loading active,
blocked and tpending.pending on the same terms, and a change to slot
reuse has to account for it.

The header of test-osync-requeue described the Darwin address-wait
queue in a sentence that read as current behavior. It states the
condition the test catches instead.
cubic-dev-ai[bot]

This comment was marked as resolved.

The thread_entry comment read as if each lock-free scan loads one field
after active. thread_kick_futex_waiters loads two, tpending.pending and
then blocked, so the sentence was wrong for the scan it had just added.
@xalestar
xalestar marked this pull request as draft October 9, 2026 18:00
* between the 100 ms os_sync poll cap (FUTEX_OS_SYNC_POLL_CAP_NS) and the
* pass threshold below, so a stranded waiter (regression: requeue no longer
* degrades to a wake at the source address) surfaces near the cap, far
* between the 100 ms polling quantum (FUTEX_OS_SYNC_POLL_CAP_NS) and the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FUTEX_OS_SYNC_POLL_CAP_NS now names a path this PR deletes; it is the polling quantum every bucket sleep uses (futex_quantum_deadline, the spin guard in futex_wait_inner). Worth a follow-up renaming it to something like FUTEX_POLL_QUANTUM_NS, and test-osync-requeue with it, so the next reader does not go looking for an os_sync wait that no longer exists.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants