Run each scheduled slot of a GitHub Actions workflow on time and exactly once,
with GitHub's own schedule trigger kept as the backstop.
GitHub starts a schedule run when its queue reaches it, which can be hours
after the cron instant, and it drops some runs outright when the queue is deep.
Moving the cron minute off the hour does not help once the queue is backed up.
The trigger GitHub starts on demand is an API dispatch, so a clock outside GitHub can deliver each slot on time. But then every slot arrives twice: once from the clock, and once, late, from GitHub's own schedule. For many scheduled jobs a second run is not harmless. A backup exports again, a staged rollout climbs another step, a report files another issue.
This project has three parts:
- The clock (
tick) runs on any host with a scheduler. Every few minutes it sends each enrolled workflow its newest due slot as arepository_dispatch. - The slot action makes every delivery of a slot agree on which slot a run belongs to. It reads a ledger of finished slots, so a delivery that arrives after the slot is done skips the work.
- The done action records a slot in the ledger once its work has succeeded.
A workflow whose repeated runs are harmless (a reconciler) needs only the clock.
- Slot: one firing of one
cron:line, meaning the line and the UTC instant it names. Every scheduled run belongs to exactly one slot, however late it starts. - Delivery: a slot reaches a workflow through GitHub's
scheduleevent, or through arepository_dispatchwhose event type isgithub-cron-trigger/<workflow file>and whoseclient_payloadnames the line ascronand the instant asslot(YYYY-MM-DDTHH:MMZ). - Ledger: one git ref per finished slot,
refs/github-cron-trigger/<workflow file>/<cron line>/<YYYYMMDDTHHMMZ>. A ref outsiderefs/headsandrefs/tagsis not a branch or a tag; nothing lists it in the UI or fetches it by default.slot_ledger.pysays why a ref is the record and how the line is written into it. - Mark: a clock's claim of a slot it sends, also a git ref, under
refs/github-cron-trigger-clock/. Marks are how several clocks share the work without sending a slot twice ("The clock"). - Part: a run whose legs are recorded separately (one per environment, say)
records
<slot>.<part>for each leg, so another delivery redoes only a leg that failed.
Keep the schedule: trigger, add the dispatch trigger, and split the workflow
into a slot job, the work, and a record job:
on:
schedule:
- cron: "30 19 * * *"
repository_dispatch:
types: [github-cron-trigger/backup.yml]
workflow_dispatch:
# Deliveries of one slot can be in flight at once. A shared group runs them
# one after the other, and `queue: max` keeps a third arrival from cancelling
# the one already waiting.
concurrency:
group: backup
queue: max
cancel-in-progress: false
permissions: {}
jobs:
slot:
runs-on: ubuntu-latest
permissions:
contents: read
outputs:
scheduled: ${{ steps.slot.outputs.scheduled }}
slot: ${{ steps.slot.outputs.slot }}
cron: ${{ steps.slot.outputs.cron }}
run: ${{ steps.slot.outputs.run }}
steps:
- id: slot
uses: curlewlabs-com/github-cron-trigger/slot@v0
work:
needs: slot
if: needs.slot.outputs.run == 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7.0.1
- run: ./back-up.sh
record:
needs: [slot, work]
if: needs.slot.outputs.scheduled == 'true' && needs.work.result == 'success'
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- uses: curlewlabs-com/github-cron-trigger/done@v0
with:
slot: ${{ needs.slot.outputs.slot }}
cron: ${{ needs.slot.outputs.cron }}A workflow is enrolled when it declares repository_dispatch with the type
github-cron-trigger/<its own file name>. A copied workflow keeps the old file
name in its type, so the clock refuses it rather than sending its slots under a
type it does not listen to.
The concurrency group makes a slot run once ("One run at a time" says why, and what other workflows declare).
Why a separate record job. Recording needs contents: write, a token that
can push to the repository. In its own job, holding nothing else, that grant is
never in reach of the work's steps or of anything they install.
scheduled:truefor ascheduleorrepository_dispatchdelivery, or aworkflow_runfrom one (see "Chained workflows"). Every other event (a manual dispatch, a pull request) givesfalse, and such a run behaves exactly as it did before enrollment.slotandcron: the instant and the line of the slot the run belongs to, empty when not scheduled.run:falseonly when the ledger already records the slot, so gating the work on it skips a second delivery's work and nothing else.
The action also writes one line to the job summary: the slot, which delivery brought it, and whether this run does the work.
For a schedule delivery the slot is the latest instant of the cron line that
fired at or before the run starts, because GitHub's event names the line but
not the instant. That is exact only while GitHub's delay is shorter than the
line's period, and GitHub's delay runs to hours. So each cron: line of a
workflow that uses the slot action must fire at most once a day: a single
minute and a single hour. A schedule delivery of a line that fires more often
is refused rather than attributed to whichever of its instants happens to be
latest, and the clock reports such a workflow instead of delivering it. A
dispatched slot is always exact.
A workflow can carry several lines, each with its own slots. A weekly line and a daily line that fire in the same minute are separate slots with separate records.
A job whose legs succeed or fail independently, such as a matrix over
environments, can record each leg as a part. Pass the same part to both
actions:
- id: slot
uses: curlewlabs-com/github-cron-trigger/slot@v0
with:
part: ${{ matrix.environment }}A later delivery then redoes only the legs whose part is missing.
A workflow that reads the current state on every run and changes only what is out of date is harmless to run again, so it needs no slot action, no ledger and no record job. Declaring the dispatch type and a concurrency group is the whole enrollment:
on:
schedule:
- cron: "17 * * * *"
repository_dispatch:
types: [github-cron-trigger/release-watch.yml]
concurrency:
group: release-watch
cancel-in-progress: false"One run at a time" says why a reconciler needs the group.
Without the slot action, a line may fire more than once a day. Each dispatch names its exact slot, so nothing has to be attributed by time.
A workflow that runs after a scheduled one through workflow_run, rather than
on a schedule of its own, uses the same slot action unchanged. A workflow_run
run whose parent was a scheduled delivery (a schedule or
repository_dispatch run) belongs to the parent's slot: the latest instant of
the parent workflow's cron line at or before the parent run was created. Each
delivery of the parent's slot then starts a chained run that resolves to the
same slot, so the chained workflow's own records make the later one skip. A
parent that was not a scheduled delivery leaves the chained run unscheduled.
The parent must carry exactly one cron: line, because a run does not say
which line fired it, and that line must fire at most once a day, for the same
reason as a schedule delivery.
Re-delivering a chained slot through its parent works only while that slot is still the parent line's newest, because the chained run takes the parent line's newest slot whatever the parent was sent.
A chained workflow declares a concurrency group as well, as any workflow running the slot action does ("One run at a time").
Every enrolled workflow, and every workflow that runs the slot action,
declares a concurrency group its scheduled runs share. Enrolling makes runs
overlap where they rarely did: when GitHub's queue is short, its schedule run
and the clock's dispatch of the same slot start minutes apart, and a slow run
can still be going when the next slot arrives.
- A workflow running the slot action needs the group to run a slot once.
The slot check and the record are separate calls, so deliveries of one slot
running side by side could each find the slot owed and each do the work.
Serialized, a later delivery starts only after the earlier one has recorded,
and it skips. Declare
queue: maxwithcancel-in-progress: false, as the example under "Enrolling a workflow" does: by default a group holds one waiting run and cancels it when another arrives, and a cancelled delivery leaves its slot to a later delivery, if one is coming. - A chained workflow declares a group as well; the parent's declaration does not reach it. The parent's group keeps the parent's deliveries apart, but each one starts a chained run as it finishes, and the first chained run can still be going when the second one starts.
- A reconciler needs one too. A reconciler is harmless to run again, not
necessarily to run twice at once: two runs that read the same state can both
act on it, as two that each open the same pull request would.
cancel-in-progress: falseon the default queue is enough here, withoutqueue: max. A waiting run that a newer one replaces loses nothing, because the newer one reconciles from later state.
A workflow that also runs on pull requests, where a PR run only verifies or dry-runs, can key that run by its run id so it never waits in the shared queue. To cancel superseded PR work, also give every job on the PR path a job-level group keyed on the PR number, with cancellation enabled only for PR runs. On other events, key each job's group on the run id so the workflow-level group alone orders deliveries. Add these concurrency blocks to the workflow and jobs in "Enrolling a workflow":
concurrency:
group:
${{ github.event_name == 'pull_request' && format('backup-pr-{0}',
github.run_id) || 'backup' }}
queue: max
cancel-in-progress: false
jobs:
slot:
concurrency:
group:
${{ github.event_name == 'pull_request' && format('backup-slot-pr-{0}',
github.event.pull_request.number) || format('backup-slot-run-{0}',
github.run_id) }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
work:
concurrency:
group:
${{ github.event_name == 'pull_request' && format('backup-work-pr-{0}',
github.event.pull_request.number) || format('backup-work-run-{0}',
github.run_id) }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}GitHub's concurrency rules
refuse queue: max beside cancel-in-progress: true in the same block, so
deliveries queue at the workflow level and PR jobs supersede at the job level.
Use distinct group names for each workflow and job. For a matrix job, include
the matrix leg's identity in the PR and non-PR group names too, so legs do not
cancel or serialize each other.
Group slot as well as the work it gates. GitHub orders a group by when a job
starts waiting on it, rather than when its workflow was dispatched. Grouping
only work lets an older run whose slot finishes late reach the work group
after the newer run and cancel the newer job. slot has no dependencies, so
grouping it starts supersession at the beginning of the chain; carry this
pattern through every job on the PR path. The enrollment example's record job
skips PR runs, so it needs no PR group.
A PR run that touches what the deliveries guard, such as a plan that reads a live state lock, stays in the shared workflow-level group instead.
tick is the delivering half, run on a short interval by any scheduler outside
GitHub, on one host or several:
python3 -m github_cron_trigger tick --repo owner/name --sendEach tick reads the workflow files of the repository's default branch through
the API, in a single GraphQL query, so the host needs no clone. It works out
every enrolled cron: line's latest due slot in UTC and sends it to the
workflow as a repository_dispatch: the newest due slot per line only, and
never a slot the ledger already records. A merged change to a workflow reaches
the clock at its next tick.
The clock keeps its memory in the repository it serves, not on its host. Before it sends a slot it claims it, by creating a git ref:
refs/github-cron-trigger-clock/<workflow file>/<cron line>/<YYYYMMDDTHHMMZ>
Creating a ref is atomic: when several clocks claim the same slot, exactly one create succeeds, and only that clock sends. So:
- Run it on as many hosts as you like. Each slot is still sent once. While one host is asleep or down, another delivers on time.
- A host keeps nothing. One can join, leave or be rebuilt with nothing to copy, and what has been sent can be read from the repository.
- A failed send releases its claim, as does a failed ledger read before it, so the next tick, on any host, sends the slot again. A claim that GitHub applied but answered as failed or refused is still recognized as the clock's own: each claim's ref points at a note unique to it. A clock stopped between its claim and its send leaves that slot to GitHub's backstop: late, but still delivered.
These marks are not the ledger. The ledger records slots whose work finished, and the slot action reads it; marks record slots a clock sent, for every enrolled workflow, reconcilers included, and only clocks read them. A tick keeps each line's two newest marks and removes the rest.
- Baselines. A line no clock has marked is baselined - its latest slot is marked without sending - because that slot may have run before anything could record it. A workflow that stops being enrolled has its marks removed, so enrolling it again baselines it again.
- Sleep. After a gap with no clock up, the next tick sends one delivery per line, for the newest due slot, never a burst. Slots no clock was up for are left to GitHub's backstop: late, but still exactly once.
- Dry run first. Without
--send, a tick reads everything and reports what it would do - baseline, send, or remove old marks - and writes nothing. Removing--sendlater is the off switch. - What an idle tick costs. Two small requests: the id of the default
branch's
.github/workflowstree, which changes only when a workflow file does, and the clock's marks. Workflow files are downloaded and parsed only when they change, and then only the changed ones; everything read before is reused from a cache under$XDG_CACHE_HOME/github-cron-trigger(or~/.cache/github-cron-trigger), keyed by content id. The cache is only a cache: deleting it costs one full read and changes nothing the clock decides. - Exit status. A tick that could not read a workflow, the marks or the ledger, or could not send a slot, logs the problem and exits 1, so the scheduler running it can report it.
Any host that can reach api.github.com will do, and more than one is better.
Install a release, which puts github-cron-trigger on PATH:
pip install "github-cron-trigger @ git+https://github.com/curlewlabs-com/github-cron-trigger@v0.2.1"Or check out a release tag and run it as PYTHONPATH=src python3 -m github_cron_trigger. Then run tick every few minutes, once per repository
served. A crontab line:
*/5 * * * * github-cron-trigger tick --repo owner/name --send >> "$HOME/github-cron-trigger.log" 2>&1On macOS, a launchd agent with StartInterval 300 does the same. On Linux, a
systemd timer with OnUnitActiveSec=5min does.
Ticking every few minutes, rather than scheduling one entry per slot, is deliberate. Slots are UTC instants, and a host scheduler's calendar entries are local times, which shift against UTC at every daylight saving change.
The clock calls the API through the gh CLI, so it uses whatever gh is
logged in as, or GH_TOKEN when that is set. A fine-grained personal access
token limited to the enrolled repositories needs:
- Contents: Read and write. Sending a
repository_dispatchrequires it, and so does creating and removing the clock's marks. Reading the workflow files, the marks and the ledger is the read half. - Metadata: Read, which every fine-grained token carries.
A dropped delivery leaves no run to go red, so a slot that neither the clock
nor GitHub's backstop completed is reported by nothing else. missed reads the
ledger against each line's due slots and names the gaps:
python3 -m github_cron_trigger missed --repo owner/nameIt checks each workflow that runs the slot action, on the lines it is enrolled
for and, when it chains through workflow_run, on each enrolled parent's line.
- A slot counts as due once
GRACE(12 hours) has passed since it, so one still in flight is not named. - A line is checked only from its oldest record on, so slots from before it enrolled are not misses, and a line that has never recorded is not checked yet.
- A record of any part of the slot counts as delivered.
- The window reaches back
LOOKBACK(7 days).
freshness.py's header says why each bound is where it is.
A workflow whose records, schedule or parent's line cannot be read is reported as not judged, never as missing every slot, and the run does not pass on it. The exit status is a bit set, so neither finding hides the other: 1 when a due slot has no record, 2 when anything could not be read, 0 when everything was read and no slot was missed. Run it daily, from the clock's host or from a scheduled workflow, and alert on a non-zero exit.
Every uncertain case fails the run rather than guessing, because a wrong guess either does owed work twice or skips it silently:
- The slot action cannot read the ledger or the workflow file, or a delivery's slot is missing or malformed. The slot step fails and the work does not run. The other delivery of the slot still gets its turn.
- A delivery's line is not one of the workflow's own
schedule:entries, or its entry is one this tool refuses (see "Accepted cron grammar"). The slot step fails for every delivery of that line until the workflow file changes. - A dispatched slot is more than ten minutes in the future, or older than the ledger keeps records. It is refused.
- The record step cannot write. The run fails and the slot stays unrecorded, so the other delivery does the work again. A visible red run and a repeat beat a slot silently marked done.
- Pruning old records fails. A warning only, because the record itself already landed.
The record step prunes its workflow's records older than 60 days
(RECORD_RETENTION in steps.py), a window far longer than any delivery delay
it has to absorb.
GitHub's own: minute, hour, day of month, month (1-12 or JAN-DEC) and day of
week (0-6 or SUN-SAT). Each field is a *, a value, a range a-b, or a comma
list of those, and any of them can be stepped with /n. A step on a single
value runs to the end of the field, so 20/15 is minutes 20, 35 and 50. A line
is resolved exactly or refused, never approximated. cron_slots.py says why
each refusal is one:
- A line restricting day of month and day of week together, where neither is a
bare
*. Cron implementations disagree on how such a line fires, and GitHub does not say which way it goes. - A line that can never fire, such as
0 0 30 2 *. - A
schedule:entry that setstimezone:. Slots are UTC instants, and GitHub does not document when a zoned schedule fires across a daylight saving fall-back. - Anything outside the grammar. That includes the
@daily-style macros GitHub itself does not support; a day of week 7, which some crons read as Sunday but GitHub does not document; and a month or weekday name in lower or mixed case, where GitHub documents only upper case.
The clock leaves out a workflow that is not enrolled and has a refused line,
and reports an enrolled one that has one. The slot action fails a delivery of a
refused line, GitHub's own schedule run included.
GitHub does not run a schedule more often than its documented minimum interval, and this tool deliberately does not hold a line to that floor. The clock sends a line's newest due slot on each tick, so how often a reconciler on an every-minute line runs is set by how often the clock ticks.
- Runners (the slot action): Linux or macOS, on x86_64 or arm64, with
python33.10 or newer, theghCLI, and mikefarah's yq v4.53.4 or newer on PATH. The done action needs onlypython3andgh. GitHub's hosted runners have all three, and neither action needs a checkout. - yq, on a self-hosted runner. Install it once, as you would any tool the
runner uses: Homebrew's
yqformula is mikefarah's, and the project's releases page has a binary for every platform. The apt package namedyqon Debian and Ubuntu is a different program, and the slot action refuses it by name.MINIMUM_VERSIONinyq.pysays why the floor sits where it does. - The clock's hosts:
python3,ghand yq as above, a scheduler, and a token as described under "The clock's token". Nothing else; a host keeps no state.
Releases are tagged vMAJOR.MINOR.PATCH, and those tags never move. The
vMAJOR tag follows the newest release of that major version. Pin an action to
a full commit SHA, or to a fixed tag, for exact reproducibility.
These are the project's contract:
- the dispatch event type and payload keys;
- the ledger's ref layout;
- the actions' inputs and outputs;
- the command-line interface.
From 1.0.0 on, they stay compatible within a major version, so a clock and the actions can be upgraded separately. Before 1.0.0, a minor release may change them, and its release notes say how; pin a fixed tag if you need them to hold still.
python3 -m unittest discover -s tests -t . # the unit tests
ruff format --check . && ruff check . && mypyCONTRIBUTING.md has the conventions. The live check against GitHub's API
(tests/live_check.py) runs in CI.
MIT. See LICENSE.