From eab435f97fbd1cb76c8bb34ccd54e72b4264ea56 Mon Sep 17 00:00:00 2001 From: Sergii Demianchuk Date: Sat, 3 Oct 2026 09:30:59 -0400 Subject: [PATCH 1/3] =?UTF-8?q?fix(test):=20authRetryCallers=20waits=20by?= =?UTF-8?q?=20the=20clock=20=E2=80=94=20its=20turn=20count=20ran=20out=20o?= =?UTF-8?q?n=20the=20Linux=20release=20gate?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit v1.3.0's release gate failed in authRetryCallers.test.js: "never happened: the renewal", in the first inline-completion case that polls. The suite had passed on every machine it had been run on, all of them Macs. until() gave a condition 400 turns of the event loop to come true. Inline completion waits out its debounce on a real timer, and a timer is at least a millisecond even when the setting is 0. How many turns fit into that millisecond is the machine's business: macOS a 0 ms timer fires after ~70 turns; 400 turns take ~5 ms Linux it fires after ~1,500 turns; 400 turns take ~0.3 ms So on the runner the count ran out before the timer had fired, and the request the case was waiting on had not been sent yet. Nothing is wrong with the code under test. The three cases that poll across the debounce are the ones exposed; which of them fails first varies from run to run. until() now waits by the clock: up to two seconds, the patience within() already has. A condition that never comes true still fails, with the same message. settle() stays a count of turns, and says what that is good for: only what no timer stands in the way of. Both places that use it are that. Reproduced before fixing, without the runner: in a Linux container the tag fails 3 runs of 3, and on a Mac with every timer made 15 ms slower it fails with the runner's exact message. With the fix: 60 of 60 in the container, all 48 suites there three times over, and all 48 on a Mac with and without the slowed timers. No other suite counts turns across a timer; the slowed-timer run would have shown one. Not run on the GitHub runner itself yet, and the container has Node 18 where the runner has 24. The next commit makes the pull request do that run. --- .../levelcode-ai/test/authRetryCallers.test.js | 18 +++++++++++++++--- 1 file changed, 15 insertions(+), 3 deletions(-) diff --git a/extensions/levelcode-ai/test/authRetryCallers.test.js b/extensions/levelcode-ai/test/authRetryCallers.test.js index 843a769..ea0dac0 100644 --- a/extensions/levelcode-ai/test/authRetryCallers.test.js +++ b/extensions/levelcode-ai/test/authRetryCallers.test.js @@ -210,11 +210,23 @@ const { registerInlineComplete } = require('../inlineComplete'); // ── the network: the only stand-in ─────────────────────────────────────────────────────────────── const tick = () => new Promise((resolve) => setImmediate(resolve)); +/** How long anything here is given to happen. Reached only by what never does. */ +const PATIENCE_MS = 2000; +/** + * Wait for `cond` — by the CLOCK, not by counting turns of the event loop. Inline completion waits + * out its debounce on a real timer, a millisecond even when the setting is 0, and how many turns fit + * into a millisecond is the machine's business: a few dozen on a Mac, well over a thousand on the + * Linux runner the release gate uses. A count that was plenty on one ran out on the other before + * the timer had fired. + */ async function until(cond, what) { - for (let i = 0; i < 400; i++) { if (cond()) { return; } await tick(); } - assert.fail('never happened: ' + what); + const deadline = Date.now() + PATIENCE_MS; + while (!cond()) { + if (Date.now() > deadline) { assert.fail('never happened: ' + what); } + await tick(); + } } -/** Let everything already in motion get as far as it can. */ +/** Let everything already in motion get as far as it can. Turns, not time — so only for what no timer stands in the way of. */ async function settle() { for (let i = 0; i < 40; i++) { await tick(); } } function deferred() { let resolve; const promise = new Promise((r) => { resolve = r; }); return { promise, resolve }; } function within(promise, ms, what) { From 7a91c637532fb5521c2566b0cc0498d02807e9d9 Mon Sep 17 00:00:00 2001 From: Sergii Demianchuk Date: Sat, 3 Oct 2026 09:30:59 -0400 Subject: [PATCH 2/3] =?UTF-8?q?ci:=20run=20the=20extension=20suites=20on?= =?UTF-8?q?=20pull=20requests=20and=20develop=20=E2=80=94=20the=20release?= =?UTF-8?q?=20gate's=20own=20script?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The suites ran in CI at one moment only: when a tag was pushed. A pull request got CodeQL and nothing else, so the first Linux run of any suite was the release gate, after the merge. That is how v1.3.0 was tagged on a suite that could not pass there (see the previous commit). The gate moves, unchanged, into scripts/test-extensions.sh: discover every extensions/*/test/*.test.js, stop at the first failing suite, and fail if the glob matches nothing. release.yml calls it where the inline loop was, and a new workflow, test.yml, calls it on every pull request and on pushes to develop — same runner image, same Node. One script, so the two cannot drift apart. test.yml asks for read access only and needs no secrets. A newer push to the same branch cancels the run before it. Checked: the script passes all 48 suites on a Mac and in a Linux container under the runner's shell flags; it exits with a failing suite's own code without running the next; it exits 1 when nothing matches. actionlint is clean on both workflow files. release.yml's gate step is the one thing here not exercised until a tag is pushed — but it is now a single line that test.yml runs on this very pull request. --- .github/workflows/release.yml | 27 +++------------------------ .github/workflows/test.yml | 33 +++++++++++++++++++++++++++++++++ scripts/test-extensions.sh | 30 ++++++++++++++++++++++++++++++ 3 files changed, 66 insertions(+), 24 deletions(-) create mode 100644 .github/workflows/test.yml create mode 100755 scripts/test-extensions.sh diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index d7d6dd9..0e9d75f 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -40,30 +40,9 @@ jobs: with: node-version: "24" - name: Extension unit tests - # `shopt` is a bash builtin, so pin the shell instead of relying on the runner default (bash on - # Linux/macOS, but pwsh on Windows — where this step would break if the job were ever copied). - # Explicit `shell: bash` also upgrades the default `bash -e` to - # `bash --noprofile --norc -eo pipefail`, so a stray profile file can't perturb the gate either. - shell: bash - # DISCOVERS suites — it used to `cd extensions/levelcode-ai`, so levelcode-updater's tests never - # ran here, including the one guarding the updater's Download button against serving a raw - # .app.zip. Globbing every extension means a new suite is gated the moment it is added, with no - # list here to keep in sync. Requires are file-relative, so running from the repo root is fine. - run: | - shopt -s nullglob - count=0 - for t in extensions/*/test/*.test.js; do - echo "── $t" - node "$t" # `-e` (from `shell: bash` above) aborts the job on the first failure - count=$((count + 1)) - done - # A zero-match glob would otherwise report success and gate nothing — the exact failure this - # step is fixing. Fail loudly instead. - if [ "$count" -eq 0 ]; then - echo "::error::No suites matched extensions/*/test/*.test.js — the gate would pass vacuously." - exit 1 - fi - echo "──────── $count test files passed ────────" + # The gate itself lives in a script, so the pull-request workflow (test.yml) runs exactly what + # this does. It discovers every extensions/*/test/*.test.js and fails if it finds none. + run: ./scripts/test-extensions.sh build: name: Build ${{ matrix.arch }} diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml new file mode 100644 index 0000000..8dd2e02 --- /dev/null +++ b/.github/workflows/test.yml @@ -0,0 +1,33 @@ +name: Extension tests + +# The release gate (scripts/test-extensions.sh), on every pull request and on develop. +# +# release.yml runs the same script, but only when a tag is pushed — which made a tag the first time +# the suites ran on Linux. v1.3.0's gate failed there on a suite that had only ever run on a Mac, +# after the pull requests that wrote it had merged. Here a suite that passes on one machine and not +# on the runner is found on the pull request that adds it. + +on: + pull_request: + push: + branches: [develop] + +permissions: + contents: read + +# A newer push to the same branch or pull request supersedes the run before it. +concurrency: + group: extension-tests-${{ github.ref }} + cancel-in-progress: true + +jobs: + test: + name: Extension unit tests + runs-on: ubuntu-latest # the runner the release gate uses + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-node@v7 + with: + node-version: "24" # as release.yml + - name: Extension unit tests + run: ./scripts/test-extensions.sh diff --git a/scripts/test-extensions.sh b/scripts/test-extensions.sh new file mode 100755 index 0000000..564da45 --- /dev/null +++ b/scripts/test-extensions.sh @@ -0,0 +1,30 @@ +#!/usr/bin/env bash +# +# Run every extension's unit suites. This is the gate: the release workflow runs it before spending +# an hour of macOS build minutes, and every pull request runs it too (.github/workflows/test.yml). +# +# It DISCOVERS suites — extensions/*/test/*.test.js — so a new one is gated the moment it is added, +# with no list to keep in sync. (The gate used to `cd extensions/levelcode-ai`; levelcode-updater's +# suites never ran, including the one guarding the Download button against serving a raw .app.zip.) +# Requires are file-relative, so the suites run from the repo root. +# +# Usage: +# ./scripts/test-extensions.sh +# +set -euo pipefail +shopt -s nullglob + +cd "$(dirname "$0")/.." + +count=0 +for t in extensions/*/test/*.test.js; do + echo "── $t" + node "$t" # `set -e` stops at the first suite that fails + count=$((count + 1)) +done +# A glob that matches nothing would otherwise report success and gate nothing. Fail loudly instead. +if [ "$count" -eq 0 ]; then + echo "::error::No suites matched extensions/*/test/*.test.js — the gate would pass vacuously." + exit 1 +fi +echo "──────── $count test files passed ────────" From 088317967773018a81c52a663ca2a49f37088d57 Mon Sep 17 00:00:00 2001 From: Sergii Demianchuk Date: Sat, 3 Oct 2026 09:32:54 -0400 Subject: [PATCH 3/3] =?UTF-8?q?test:=20say=20what=20was=20measured=20?= =?UTF-8?q?=E2=80=94=20the=20turn=20counts=20are=20Linux's,=20not=20the=20?= =?UTF-8?q?runner's=20own?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The comment on until() gave a figure for "the Linux runner the release gate uses". The figure was measured in a Linux container; on the runner itself only the outcome is known, that 400 turns ran out. Say Linux, where the gate runs. --- extensions/levelcode-ai/test/authRetryCallers.test.js | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/extensions/levelcode-ai/test/authRetryCallers.test.js b/extensions/levelcode-ai/test/authRetryCallers.test.js index ea0dac0..f41ccef 100644 --- a/extensions/levelcode-ai/test/authRetryCallers.test.js +++ b/extensions/levelcode-ai/test/authRetryCallers.test.js @@ -215,8 +215,8 @@ const PATIENCE_MS = 2000; /** * Wait for `cond` — by the CLOCK, not by counting turns of the event loop. Inline completion waits * out its debounce on a real timer, a millisecond even when the setting is 0, and how many turns fit - * into a millisecond is the machine's business: a few dozen on a Mac, well over a thousand on the - * Linux runner the release gate uses. A count that was plenty on one ran out on the other before + * into a millisecond is the machine's business: a few dozen on a Mac, well over a thousand on + * Linux, where the release gate runs. A count that was plenty on one ran out on the other before * the timer had fired. */ async function until(cond, what) {