From 5c4831c6701c7c563aee24b292064f3398ba8a6d Mon Sep 17 00:00:00 2001 From: Vladimir nett00n Budylnikov Date: Tue, 22 Sep 2026 17:31:06 +0400 Subject: [PATCH 1/2] read only system tray for linux --- .github/workflows/ci.yml | 24 +- .gitignore | 3 + README.md | 37 +- console/Cargo.lock | 2 + console/Cargo.toml | 13 + console/src/bin/jobs.rs | 51 +- console/src/bin/setup.rs | 8 +- console/src/bin/tray.rs | 1052 +++++++++++----------- console/src/{jobs.rs => jobs/launchd.rs} | 474 +--------- console/src/jobs/mod.rs | 632 +++++++++++++ console/src/jobs/systemd.rs | 618 +++++++++++++ console/src/jobs/unsupported.rs | 31 + console/src/lib.rs | 1 + console/src/view.rs | 159 ++++ docs/DIAGNOSTICS.md | 32 +- docs/INSTALL.md | 38 +- 16 files changed, 2132 insertions(+), 1043 deletions(-) rename console/src/{jobs.rs => jobs/launchd.rs} (58%) create mode 100644 console/src/jobs/mod.rs create mode 100644 console/src/jobs/systemd.rs create mode 100644 console/src/jobs/unsupported.rs create mode 100644 console/src/view.rs diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 1e1a291..57c5016 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -69,15 +69,31 @@ jobs: # because it shares the toolchain, and a console that does not compile is a console nobody # notices is broken until they try to open it. # - # On this runner that means the data layer and the four command-line tools -- the tray and - # its two dependencies are declared for macOS only, since the schedules it reads are - # launchd's. Everything the tests cover is in the part that builds here. - - name: Build the console (data layer + CLI tools; the tray is macOS-only) + # On this runner that means the data layer and the four command-line tools with the default + # feature set: the tray's GUI dependencies are declared for macOS unconditionally, and for + # Linux only behind the `tray` feature (off by default), specifically so this step needs no + # GUI library at all. Everything the tests cover here is in the part that builds without it. + - name: Build the console (data layer + CLI tools; no GUI libraries needed) working-directory: console run: cargo build --release --locked - name: Console unit tests working-directory: console run: cargo test --locked + # The Linux tray, behind its feature: exercises the `tray-icon` + GTK path this runner does + # not otherwise touch. Its own dependencies need real headers, hence the apt step -- and + # that step is confined to here so the two steps above stay proof that nothing else in this + # crate needs them. + - name: Build the Linux tray (GTK + appindicator) + working-directory: console + run: | + sudo apt-get update + # libxdo-dev is easy to miss: tray-icon's menu-accelerator handling links -lxdo, and the + # link fails without it even though nothing above mentions xdo by name. + sudo apt-get install -y libgtk-3-dev libayatana-appindicator3-dev libxdo-dev + cargo build --release --locked --features tray + - name: Linux tray unit tests + working-directory: console + run: cargo test --locked --features tray schema: runs-on: ubuntu-latest diff --git a/.gitignore b/.gitignore index cee52d6..0831501 100644 --- a/.gitignore +++ b/.gitignore @@ -1,5 +1,8 @@ # Rust build artifacts mcp-server/target/ +# CARGO_TARGET_DIR for the console crate, when set to the repo root's build/ (keeps target output +# out of an immutable-OS host's toolbox checkout and in one place regardless of which crate built). +/build/ # Python __pycache__/ diff --git a/README.md b/README.md index 463ec7b..0220b06 100644 --- a/README.md +++ b/README.md @@ -270,10 +270,14 @@ holds, whether the scheduled passes are running, and what the tunables are set t writing a query. ```bash -cd console && cargo build --release # two dependencies, both only for the tray +cd console && cargo build --release # the CLI tools; zero dependencies ./target/release/hypermnesia-setup # the walkthrough: from nothing to a menu-bar icon ``` +The tray itself needs a GUI toolkit, so it is not part of that plain build: unconditional on +macOS, and on Linux behind `cargo build --release --features tray` (GTK3 + +libayatana-appindicator, see [docs/INSTALL.md](docs/INSTALL.md)). + It reaches the database exactly one way: a command that receives SQL on stdin. Direct psql, `docker exec`, `kubectl exec`, ssh to a machine that has kubectl — all of them are one string with different contents, which is why there is one setting and not five. The wizard tries the command @@ -282,7 +286,7 @@ cannot be told from an empty store. | Command | What | |---------|------| -| `hypermnesia` | the menu-bar tray (macOS) | +| `hypermnesia` | the menu-bar tray (macOS) / system tray (Linux, `--features tray`; jobs are read-only there so far) | | `hypermnesia-stats` | the same numbers on stdout | | `hypermnesia-jobs` | scheduled passes: what is configured, when each last worked, run one now, change a schedule | | `hypermnesia-settings` | the tunables, each with the value in force and where that value came from | @@ -295,9 +299,9 @@ cannot be told from an empty store.

A menu, not a window. Everything the console has to show is a dozen lines and a dozen buttons; a -window would mean a GUI framework for the same result. Two dependencies, both macOS-only, and the -data layer under them has none at all — a console that takes a minute to build is a console nobody -rebuilds. +window would mean a GUI framework for the same result. The tray's GUI dependencies are declared +per platform and behind a feature on Linux, and the data layer under them has none at all — a +console that takes a minute to build is a console nobody rebuilds. What is in the menu: @@ -308,20 +312,23 @@ What is in the menu: chunks have no embedding, how many embedding models are in the store, the review queue, the stale count, the database size. - **Run now** — every scheduled pass, with its schedule, when it last wrote to its log, and its - last exit code. Pressing one runs it through launchd and reports what launchd then did, not that - the request was accepted. -- **Schedule** — the common intervals and times, per job. It edits the plist, validates it, and - reloads the job, because launchd keeps its own copy from the moment it loaded it: writing the - file without the reload would show a new schedule while the old one is in force. + last exit code. On macOS, pressing one runs it through launchd and reports what launchd then + did, not that the request was accepted. On Linux this is read-only so far: the row shows the + same information, read from `systemctl --user show`, but pressing it reports "not implemented". +- **Schedule** — the common intervals and times, per job. On macOS it edits the plist, validates + it, and reloads the job, because launchd keeps its own copy from the moment it loaded it: writing + the file without the reload would show a new schedule while the old one is in force. Not + implemented on Linux yet. - **Refresh now** and **Quit**. -The menu-bar title carries one mark, `HM !`, and it is derived from the lines below rather than -computed beside them: anything the menu would show with a `!` puts the mark in the title. That is -the whole design in one detail — the icon is where a problem is noticed, so the icon must not be -able to disagree with the menu. +The tray carries one mark for "something here needs a look", derived from the lines below rather +than computed beside them: anything the menu would show with a `!` raises it. That is the whole +design in one detail — the mark is where a problem is noticed, so it must not be able to disagree +with the menu. On macOS the mark is in the menu-bar title, `HM !`; a system tray icon has no text +of its own, so on Linux it is the icon itself, calm green or warned red. `hypermnesia --install` puts it in launchd to start at login, restarted if it crashes and not if -you quit it. `--uninstall` takes it back out. +you quit it, on macOS. `--uninstall` takes it back out. Not implemented on Linux yet. **Try it without a database.** The console's only connection setting is a command that prints the query's answer, so a file works: diff --git a/console/Cargo.lock b/console/Cargo.lock index 50d2bdd..de79410 100644 --- a/console/Cargo.lock +++ b/console/Cargo.lock @@ -808,6 +808,8 @@ checksum = "e17592d60ebacc7d5e169f4663c5f84f9161cc90328abcfe8456f41e4dfcb284" name = "hypermnesia-console" version = "0.1.0" dependencies = [ + "glib", + "gtk", "tray-icon", "winit", ] diff --git a/console/Cargo.toml b/console/Cargo.toml index e765a9a..9dcd29d 100644 --- a/console/Cargo.toml +++ b/console/Cargo.toml @@ -19,6 +19,19 @@ license = "MIT" tray-icon = "0.25" winit = "0.30" +# The Linux tray, behind a feature and off by default -- for the same reason the macOS deps are +# per-target: the data layer and the four command-line tools must keep building on a machine with +# no GUI libraries at all, which is what CI's bare runner is. `tray-icon` drives the same crate as +# macOS, through `libayatana-appindicator`; on Linux it needs a GTK main loop rather than winit's, +# so `gtk` is a dependency here and `winit` is not. +[target.'cfg(target_os = "linux")'.dependencies] +tray-icon = { version = "0.25", optional = true } +gtk = { version = "0.18", optional = true } +glib = { version = "0.18", optional = true } + +[features] +tray = ["dep:tray-icon", "dep:gtk", "dep:glib"] + [[bin]] name = "hypermnesia-stats" path = "src/bin/stats.rs" diff --git a/console/src/bin/jobs.rs b/console/src/bin/jobs.rs index 629d5c5..108a65b 100644 --- a/console/src/bin/jobs.rs +++ b/console/src/bin/jobs.rs @@ -11,15 +11,15 @@ use hypermnesia_console::jobs::{self, Job}; fn main() { let args: Vec = std::env::args().skip(1).collect(); // Through jobs::job_prefix rather than the variable: only that path also consults the - // config file, and the tray -- started by launchd with a minimal environment -- has - // nothing else to read. + // config file, and the tray -- started by the service manager with a minimal environment -- + // has nothing else to read. let prefix = jobs::job_prefix(); let all = match jobs::list(&prefix) { Ok(j) => j, Err(e) => fail(&e), }; if all.is_empty() { - fail(&format!("no job with the prefix {prefix} found in ~/Library/LaunchAgents")); + fail(&format!("no job with the prefix {prefix} found in {}", jobs::UNITS_LOCATION)); } match args.first().map(String::as_str) { @@ -160,9 +160,10 @@ fn parse_weekday(s: &str) -> Option { fn find<'a>(all: &'a [Job], name: &str) -> &'a Job { match all.iter().find(|j| j.short() == name || j.label == name) { Some(j) if j.broken() => fail(&format!( - "{name}: this plist cannot be read ({}). launchd may still be running the copy it \ - loaded before the file broke -- what it will do at the next login is the question. \ - Fix the file first.", j.fault.clone().unwrap_or_default())), + "{name}: this {} cannot be read ({}). {} may still be running the copy it loaded \ + before the file broke -- what it will do at the next login is the question. Fix \ + the file first.", jobs::UNIT_NOUN, j.fault.clone().unwrap_or_default(), + jobs::BACKEND_NAME)), Some(j) => j, None => fail(&format!("no such job: {name}. There is: {}", all.iter().map(Job::short).collect::>().join(", "))), @@ -178,10 +179,11 @@ fn list(all: &[Job]) { println!("{:<13} {:<22} {:<14} {}", "JOB", "SCHEDULE", "LAST OUTPUT", "STATE"); for j in all { if let Some(why) = &j.fault { - // An unreadable plist is one launchd also refused at login: the job is not running. - // It used to be dropped from the list entirely, which is the one case the console had - // nothing at all to say about. - println!("{:<13} {:<22} {:<14} ! unreadable plist: {why}", j.short(), "—", "—"); + // An unreadable unit is one the service manager also refused: the job is not + // running. It used to be dropped from the list entirely, which is the one case the + // console had nothing at all to say about. + println!("{:<13} {:<22} {:<14} ! unreadable {}: {why}", j.short(), "—", "—", + jobs::UNIT_NOUN); continue; } let when = match j.since_last_output() { @@ -198,7 +200,7 @@ fn list(all: &[Job]) { None => "never ran".to_string(), } } else if j.last_exit.is_none() { - "not loaded into launchd".to_string() + format!("not loaded into {}", jobs::BACKEND_NAME) } else if j.runs == Some(0) { // Loaded, and launchd has not started it since this login. Its exit column says 0, // which it prints both for "finished successfully" and for "never finished at all" -- @@ -213,7 +215,7 @@ fn list(all: &[Job]) { // has nothing to tell them apart by. Some(0) => format!("last run ok{}", runs_note(j)), Some(c) => format!("last run: code {c}{}", runs_note(j)), - None => "not loaded into launchd".to_string(), + None => format!("not loaded into {}", jobs::BACKEND_NAME), } }; println!("{:<13} {:<22} {:<14} {}", j.short(), j.schedule.human(), when, state); @@ -227,7 +229,7 @@ fn list(all: &[Job]) { // job installed later than its own last slot honestly shows a zero and is not a // complaint: a false alarm here would teach one to scroll past this line. let what = if j.never_ran() { - "the period has already passed and launchd still never ran it" + "the period has already passed and it still never ran" } else { "it ran at some point, but the log has not moved for more than a period" }; @@ -254,6 +256,12 @@ fn ago(d: std::time::Duration) -> String { else { format!("{} d ago", s / 86400) } } +// The schedule-editing mechanism described below is real on macOS -- a plist is written, linted +// and reloaded into launchd, with a `.bak` for safety -- and is simply not implemented yet on +// Linux (`jobs::set_schedule` returns "not implemented" there). Two texts, not one interpolated +// with `jobs::BACKEND_NAME`, because the difference is not a noun, it is a paragraph that is +// false on one of the two platforms. +#[cfg(target_os = "macos")] const HELP: &str = "\ hypermnesia-jobs — the memory pipeline's scheduled jobs (launchd). @@ -273,3 +281,20 @@ The name is the tail of the label: extract, consolidate, reflect, freshness, rer The label prefix comes from HM_JOB_PREFIX, then the console config file, then the default com.hypermnesia. "; + +#[cfg(not(target_os = "macos"))] +const HELP: &str = "\ +hypermnesia-jobs — the memory pipeline's scheduled jobs (systemd user units). + + hypermnesia-jobs what is configured, when it worked, the last exit code + hypermnesia-jobs log [N] the last N lines of the log (40 by default), where one is \ +configured + +`run`, `every` and `at` are read on this platform, but not yet implemented: this build can list +systemd user timers and services, not edit or trigger them. Use `systemctl --user start +.service` and `systemctl --user edit .timer` directly for now. + +The name is the tail of the unit's file stem: extract, consolidate, reflect, freshness, rerank. +The unit prefix comes from HM_JOB_PREFIX, then the console config file, then the +default hypermnesia-. +"; diff --git a/console/src/bin/setup.rs b/console/src/bin/setup.rs index 4626817..5e912d7 100644 --- a/console/src/bin/setup.rs +++ b/console/src/bin/setup.rs @@ -113,12 +113,14 @@ fn walkthrough() { step(6, "The menu bar"); if yes("Add the tray to autostart?", true) { - // launchd starts the tray with a minimal environment: no login shell, none of your - // exports. A command that reads $DATABASE_URL was verified HERE, where you have it. + // The service manager starts the tray with a minimal environment: no login shell, none + // of your exports. A command that reads $DATABASE_URL was verified HERE, where you have + // it. let cmd = config().get("HM_PSQL_CMD").cloned().unwrap_or_default(); if let Some(var) = shell_variable_in(&cmd) { println!("! The command you configured uses ${var}, which this shell supplies and"); - println!(" launchd does not: it starts jobs with a minimal environment. In the tray"); + println!(" {} does not: it starts jobs with a minimal environment. In the tray", + jobs::BACKEND_NAME); println!(" that command will fail where it works here. Either write the value into"); println!(" the command (hypermnesia-setup --connect) or keep using the CLI tools."); if !yes("Install it anyway?", false) { diff --git a/console/src/bin/tray.rs b/console/src/bin/tray.rs index e66f435..6164c2e 100644 --- a/console/src/bin/tray.rs +++ b/console/src/bin/tray.rs @@ -1,4 +1,4 @@ -//! The memory console in the menu bar. +//! The memory console in the menu bar / system tray. //! //! A menu, not a window. Everything this console has to show is a dozen lines of state and a //! dozen buttons; a window would mean a GUI framework for the same result. @@ -6,611 +6,609 @@ //! Two rules everything else follows from: //! //! 1. The main thread never waits. Reading the store can take seconds (a connection, a query); -//! doing that on the menu thread freezes the menu bar, and on macOS the whole NSApplication -//! run loop with it. Everything slow lives on a worker thread and sends its result back. +//! doing that on the menu thread freezes the menu, and on macOS the whole NSApplication run +//! loop with it. Everything slow lives on a worker thread and sends its result back. //! 2. A failure is visible. The menu-bar title and the first menu line say when the data did not //! arrive, instead of letting yesterday's numbers look current. A console whose stale state //! is indistinguishable from its fresh state is worse than no console. +//! +//! `mod app` is everything above the event loop: state, the worker thread, and how the menu is +//! rendered from `hypermnesia_console::view`. It is shared, unchanged, by every platform this +//! binary supports, because none of it is platform-specific -- reading the store and reading the +//! job list are already portable, and `tray_icon`'s `Menu`/`MenuItem`/`TrayIcon` API is the same +//! crate on macOS and on Linux. Only *driving* that API differs: macOS needs a winit +//! `ApplicationHandler` and its NSApplication run loop; Linux needs a GTK main loop, since that is +//! what `tray-icon`'s Linux backend (`libayatana-appindicator`) is built on. `mod mac` and +//! `mod linux` hold exactly that seam and nothing else. + +#[cfg(any(target_os = "macos", all(target_os = "linux", feature = "tray")))] +mod app { + use std::sync::mpsc; + use std::time::{Duration, Instant, SystemTime}; + + use hypermnesia_console::jobs::{self, Job}; + use hypermnesia_console::view; + use hypermnesia_console::{fetch, Stats, Target}; + + use tray_icon::menu::{Menu, MenuEvent, MenuItem, PredefinedMenuItem, Submenu}; + use tray_icon::{Icon, TrayIcon, TrayIconBuilder}; + + /// How often to refresh on its own. A minute: the numbers move slowly and every reading is a + /// round trip to the store. + pub const REFRESH: Duration = Duration::from_secs(60); + + /// How long the worker may be silent after being asked something before the menu says so. The + /// store's own timeout is 30 s by default, so this is well past any normal answer: the point + /// is to notice a worker that will never answer at all, not to hurry a slow one. + const WORKER_PATIENCE: Duration = Duration::from_secs(90); + + /// Ready-made schedules offered in the menu. Exactly what is usually wanted and nothing more: + /// the rare case belongs on the command line -- where, on Linux today, it is the only place: + /// `jobs::set_schedule` refuses with "not implemented on Linux yet", and picking one of these + /// items just shows that refusal in the note line rather than silently doing nothing. + const PRESETS: &[(&str, jobs::Schedule)] = &[ + ("hourly", jobs::Schedule::Every(3600)), + ("every 4 hours", jobs::Schedule::Every(14_400)), + ("every 12 hours", jobs::Schedule::Every(43_200)), + ("daily at 05:30", jobs::Schedule::Calendar(jobs::Cal { hour: Some(5), minute: Some(30), day: None, weekday: None, month: None })), + ("weekly, Mon 06:10", jobs::Schedule::Calendar(jobs::Cal { hour: Some(6), minute: Some(10), day: None, weekday: Some(1), month: None })), + ]; -// The tray is macOS-only, and not by accident: the schedules it shows and edits are launchd's, -// and the icon lives in the system menu bar. The four command-line tools work anywhere psql -// does, so the crate still builds without a single GUI library present. -#[cfg(target_os = "macos")] -mod mac { -use std::sync::mpsc; -use std::time::{Duration, Instant, SystemTime}; - -use hypermnesia_console::jobs::{self, Job}; -use hypermnesia_console::{fetch, Stats, Target}; - -use tray_icon::menu::{Menu, MenuEvent, MenuItem, PredefinedMenuItem, Submenu}; -use tray_icon::{TrayIcon, TrayIconBuilder}; -use winit::application::ApplicationHandler; -use winit::event_loop::{ActiveEventLoop, ControlFlow, EventLoop}; - -/// How often to refresh on its own. A minute: the numbers move slowly and every reading is a -/// round trip to the store. -const REFRESH: Duration = Duration::from_secs(60); - -/// How long the worker may be silent after being asked something before the menu says so. The -/// store's own timeout is 30 s by default, so this is well past any normal answer: the point is -/// to notice a worker that will never answer at all, not to hurry a slow one. -const WORKER_PATIENCE: Duration = Duration::from_secs(90); - -/// Ready-made schedules offered in the menu. Exactly what is usually wanted and nothing more: -/// the rare case belongs on the command line. -const PRESETS: &[(&str, jobs::Schedule)] = &[ - ("hourly", jobs::Schedule::Every(3600)), - ("every 4 hours", jobs::Schedule::Every(14_400)), - ("every 12 hours", jobs::Schedule::Every(43_200)), - ("daily at 05:30", jobs::Schedule::Calendar(jobs::Cal { hour: Some(5), minute: Some(30), day: None, weekday: None, month: None })), - ("weekly, Mon 06:10", jobs::Schedule::Calendar(jobs::Cal { hour: Some(6), minute: Some(10), day: None, weekday: Some(1), month: None })), -]; - -pub fn main() { - // Install and uninstall run before the event loop exists: both finish immediately and ask - // for no window. - let args: Vec = std::env::args().skip(1).collect(); - match args.first().map(String::as_str) { - Some("--install") => return finish(jobs::install_self(jobs::DEFAULT_TRAY_LABEL)), - Some("--uninstall") => return finish(jobs::uninstall_self(jobs::DEFAULT_TRAY_LABEL)), - Some("-h") | Some("--help") => { - print!("{HELP}"); - return; - } - Some(other) => { - eprintln!("hypermnesia: unknown argument: {other}"); - std::process::exit(1); - } - None => {} - } - - let event_loop = EventLoop::builder().build().expect("event loop"); - event_loop.set_control_flow(ControlFlow::WaitUntil(Instant::now() + Duration::from_millis(200))); - - let (tx, rx) = mpsc::channel::(); - let (cmd_tx, cmd_rx) = mpsc::channel::(); - spawn_worker(tx, cmd_rx); - // Read immediately: an empty menu at startup looks like a broken console. - let _ = cmd_tx.send(Command::Refresh); - - let mut app = App { - tray: None, - items: Items::default(), - rx, - cmd: cmd_tx, - last: None, - status: "loading…".to_string(), - note: None, - awaiting: Some(("the first reading".into(), Instant::now())), - worker_dead: false, - last_refresh: Instant::now(), - }; - if let Err(e) = event_loop.run_app(&mut app) { - eprintln!("hypermnesia: the event loop ended: {e}"); - } -} - -fn finish(r: Result) { - match r { - Ok(msg) => println!("{msg}"), - Err(e) => { - eprintln!("hypermnesia: {e}"); - std::process::exit(1); - } - } -} - -const HELP: &str = "\ -hypermnesia — the memory console in the menu bar. + pub const HELP: &str = "\ +hypermnesia — the memory console in the menu bar / system tray. hypermnesia run the tray - hypermnesia --install add to autostart (launchd: at login, restarted if it crashes) - hypermnesia --uninstall remove from autostart + hypermnesia --install add to autostart, where supported + hypermnesia --uninstall remove from autostart, where supported The numbers and the buttons are the same ones hypermnesia-stats and hypermnesia-jobs give: the -tray calls the same functions. Connection settings come from hypermnesia-setup. +tray calls the same functions. Connection settings come from hypermnesia-setup. Autostart and +schedule editing are launchd-only for now; on Linux those actions say so rather than doing +nothing silently. "; -/// What the worker thread sends back. -enum Update { - Data(Box), - Failed(String), - /// The outcome of a button press: a line to show at the top of the menu. - Note(String), -} + /// Handle `--install` / `--uninstall` / `--help` before any window or GTK loop exists: all + /// three finish immediately and ask for no window. Returns `true` if one of them was handled + /// (the caller should return without building the tray), `false` to proceed. + pub fn handle_early_args(args: &[String]) -> bool { + match args.first().map(String::as_str) { + Some("--install") => { finish(jobs::install_self(jobs::DEFAULT_TRAY_LABEL)); true } + Some("--uninstall") => { finish(jobs::uninstall_self(jobs::DEFAULT_TRAY_LABEL)); true } + Some("-h") | Some("--help") => { print!("{HELP}"); true } + Some(other) => { + eprintln!("hypermnesia: unknown argument: {other}"); + std::process::exit(1); + } + None => false, + } + } -struct Snapshot { - stats: Stats, - jobs: Vec, - /// Why the job list is missing, if it is. `unwrap_or_default()` used to turn "could not read - /// LaunchAgents" into an empty list, which rendered as empty Run-now and Schedule submenus - /// and a calm title -- "no jobs readable" and "every job healthy" looked identical. - jobs_error: Option, - /// Wall clock, not `Instant`: the age of the data has to survive the Mac going to sleep, and - /// a monotonic clock stops while it does. That made a two-hour nap look like a fresh reading. - at: SystemTime, -} + fn finish(r: Result) { + match r { + Ok(msg) => println!("{msg}"), + Err(e) => { + eprintln!("hypermnesia: {e}"); + std::process::exit(1); + } + } + } -enum Command { - Refresh, - RunJob(String), - SetSchedule(String, jobs::Schedule), -} + /// What the worker thread sends back. + pub enum Update { + Data(Box), + Failed(String), + /// The outcome of a button press: a line to show at the top of the menu. + Note(String), + } -/// The worker: every slow thing lives here. The main thread only posts commands to it. -fn spawn_worker(tx: mpsc::Sender, rx: mpsc::Receiver) { - std::thread::spawn(move || { - let read_all = |tx: &mpsc::Sender| { - // Re-read the settings on every pass rather than once at startup. The tray runs for - // weeks; a connection fixed with the setup wizard in the meantime has to reach the - // running tray, or the fix looks like it did not work. - let target = Target::default(); - let prefix = jobs::job_prefix(); - // Jobs are read locally and fast; the store is read over whatever transport was - // configured and may be slow. If the store does not answer, the jobs are still worth - // showing. - let (job_list, jobs_error) = match jobs::list(&prefix) { - Ok(j) => (j, None), - Err(e) => (Vec::new(), Some(e)), - }; - match fetch(&target) { - Ok(stats) => { - let _ = tx.send(Update::Data(Box::new(Snapshot { - stats, jobs: job_list, jobs_error, at: SystemTime::now(), - }))); - } - Err(e) => { let _ = tx.send(Update::Failed(e)); } - } - }; - while let Ok(cmd) = rx.recv() { - match cmd { - Command::Refresh => read_all(&tx), - Command::RunJob(label) => { - let msg = match jobs::run_now(&label) { - Ok(m) => m, - // Marked, like every other failure: the title reads the leading "!". - Err(e) => format!("! did not start: {e}"), - }; - let _ = tx.send(Update::Note(msg)); - read_all(&tx); + pub struct Snapshot { + stats: Stats, + jobs: Vec, + /// Why the job list is missing, if it is. `unwrap_or_default()` used to turn "could not + /// read the job directory" into an empty list, which rendered as empty Run-now and + /// Schedule submenus and a calm title -- "no jobs readable" and "every job healthy" + /// looked identical. + jobs_error: Option, + /// Wall clock, not `Instant`: the age of the data has to survive the machine sleeping, + /// and a monotonic clock stops while it does. That made a two-hour nap look like a fresh + /// reading. + at: SystemTime, + } + + pub enum Command { + Refresh, + RunJob(String), + SetSchedule(String, jobs::Schedule), + } + + /// The worker: every slow thing lives here. The main thread only posts commands to it. + pub fn spawn_worker(tx: mpsc::Sender, rx: mpsc::Receiver) { + std::thread::spawn(move || { + let read_all = |tx: &mpsc::Sender| { + // Re-read the settings on every pass rather than once at startup. The tray runs + // for weeks; a connection fixed with the setup wizard in the meantime has to + // reach the running tray, or the fix looks like it did not work. + let target = Target::default(); + let prefix = jobs::job_prefix(); + // Jobs are read locally and fast; the store is read over whatever transport was + // configured and may be slow. If the store does not answer, the jobs are still + // worth showing. + let (job_list, jobs_error) = match jobs::list(&prefix) { + Ok(j) => (j, None), + Err(e) => (Vec::new(), Some(e)), + }; + match fetch(&target) { + Ok(stats) => { + let _ = tx.send(Update::Data(Box::new(Snapshot { + stats, jobs: job_list, jobs_error, at: SystemTime::now(), + }))); + } + Err(e) => { let _ = tx.send(Update::Failed(e)); } } - Command::SetSchedule(label, sched) => { - // Look the job up again: the list the menu was drawn from may be a minute - // old, and editing a schedule from a stale record means editing something - // other than what the person saw. - let msg = match jobs::list(&jobs::job_prefix()).ok() - .and_then(|all| all.into_iter().find(|j| j.label == label)) { - Some(j) => match jobs::set_schedule(&j, &sched) { + }; + while let Ok(cmd) = rx.recv() { + match cmd { + Command::Refresh => read_all(&tx), + Command::RunJob(label) => { + let msg = match jobs::run_now(&label) { Ok(m) => m, - Err(e) => format!("! {e}"), - }, - None => format!("! the job {label} is gone"), - }; - let _ = tx.send(Update::Note(msg)); - read_all(&tx); + // Marked, like every other failure: the title reads the leading "!". + Err(e) => format!("! did not start: {e}"), + }; + let _ = tx.send(Update::Note(msg)); + read_all(&tx); + } + Command::SetSchedule(label, sched) => { + // Look the job up again: the list the menu was drawn from may be a + // minute old, and editing a schedule from a stale record means editing + // something other than what the person saw. + let msg = match jobs::list(&jobs::job_prefix()).ok() + .and_then(|all| all.into_iter().find(|j| j.label == label)) { + Some(j) => match jobs::set_schedule(&j, &sched) { + Ok(m) => m, + Err(e) => format!("! {e}"), + }, + None => format!("! the job {label} is gone"), + }; + let _ = tx.send(Update::Note(msg)); + read_all(&tx); + } } } - } - }); -} + }); + } -/// The menu items we later act on. The menu is rebuilt whole on every change: there are dozens -/// of items, not thousands, and rebuilding wholesale rules out a label disagreeing with its data -/// -- a partly updated menu is the same class of quiet lie as everything else this fixes. -#[derive(Default)] -struct Items { - refresh: Option, - quit: Option, - run: Vec<(MenuItem, String)>, - sched: Vec<(MenuItem, String, jobs::Schedule)>, -} + /// The menu items we later act on. The menu is rebuilt whole on every change: there are + /// dozens of items, not thousands, and rebuilding wholesale rules out a label disagreeing + /// with its data -- a partly updated menu is the same class of quiet lie as everything else + /// this fixes. + #[derive(Default)] + struct Items { + refresh: Option, + quit: Option, + run: Vec<(MenuItem, String)>, + sched: Vec<(MenuItem, String, jobs::Schedule)>, + } -struct App { - tray: Option, - items: Items, - rx: mpsc::Receiver, - cmd: mpsc::Sender, - last: Option, - status: String, - /// The outcome of the last button press, on its own line and with its own age. - /// - /// It used to share one field with the status line, and the refresh that every button starts - /// overwrote it a moment later -- so a "Run now" that failed ended up reading "updated just - /// now". The outcome of something a person did is the last thing that should be overwritten. - note: Option<(String, SystemTime)>, - /// What the worker was last asked, and when. Cleared by any answer. - awaiting: Option<(String, Instant)>, - /// The worker thread is gone: nothing will ever be read again. - worker_dead: bool, - last_refresh: Instant, -} + pub struct App { + tray: Option, + items: Items, + rx: mpsc::Receiver, + cmd: mpsc::Sender, + last: Option, + status: String, + /// The outcome of the last button press, on its own line and with its own age. + /// + /// It used to share one field with the status line, and the refresh that every button + /// starts overwrote it a moment later -- so a "Run now" that failed ended up reading + /// "updated just now". The outcome of something a person did is the last thing that + /// should be overwritten. + note: Option<(String, SystemTime)>, + /// What the worker was last asked, and when. Cleared by any answer. + awaiting: Option<(String, Instant)>, + /// The worker thread is gone: nothing will ever be read again. + worker_dead: bool, + last_refresh: Instant, + } -impl ApplicationHandler for App { - fn resumed(&mut self, _: &ActiveEventLoop) { - if self.tray.is_none() { - self.rebuild(); - } + /// What a platform's event loop should do after a tick. + pub enum Tick { + /// Nothing to do until the next scheduled wake-up. + Continue, + /// The person chose Quit. + Quit, } - fn window_event(&mut self, _: &ActiveEventLoop, _: winit::window::WindowId, - _: winit::event::WindowEvent) {} - - fn about_to_wait(&mut self, el: &ActiveEventLoop) { - // Not one blocking call here: a non-blocking channel read and a timer. - let mut dirty = false; - loop { - match self.rx.try_recv() { - Ok(u) => { - self.awaiting = None; - match u { - Update::Data(s) => { - // Not a sentence: the sentence is rendered at rebuild time from - // `last.at`. Frozen, it said "updated just now" for the whole - // refresh cycle -- and after the Mac slept, for as long as the nap - // lasted, because the refresh timer is monotonic and stops with it. - self.status = String::new(); - self.last = Some(*s); - } - Update::Failed(e) => { - // The old numbers stay -- they beat an empty screen -- but they are - // labelled with the reason and with HOW OLD they are. Without the age - // you cannot tell "could not reach it a second ago" from "stuck since - // yesterday". - let age = self.last.as_ref() - .map(|s| format!(", showing state from {} ago", ago(age_of(s.at)))) - .unwrap_or_default(); - self.status = format!("! not updated: {e}{age}"); + impl App { + pub fn new(rx: mpsc::Receiver, cmd: mpsc::Sender) -> Self { + App { + tray: None, + items: Items::default(), + rx, + cmd, + last: None, + status: "loading…".to_string(), + note: None, + awaiting: Some(("the first reading".into(), Instant::now())), + worker_dead: false, + last_refresh: Instant::now(), + } + } + + /// Post a command and remember that an answer is owed. Every path to the worker goes + /// through here so that nothing can be asked without the watchdog knowing about it. + fn ask(&mut self, cmd: Command, status: &str, what: &str) { + self.status = status.to_string(); + self.awaiting = Some((what.to_string(), Instant::now())); + let _ = self.cmd.send(cmd); + } + + /// Not one blocking call in here: a non-blocking channel read and a timer. Called on a + /// ~200ms tick by whichever platform loop is driving this `App`. + pub fn tick(&mut self) -> Tick { + if self.tray.is_none() { + self.rebuild(); + } + + let mut dirty = false; + loop { + match self.rx.try_recv() { + Ok(u) => { + self.awaiting = None; + match u { + Update::Data(s) => { + // Not a sentence: the sentence is rendered at rebuild time from + // `last.at`. Frozen, it said "updated just now" for the whole + // refresh cycle -- and after the machine slept, for as long as + // the nap lasted, because the refresh timer is monotonic and + // stops with it. + self.status = String::new(); + self.last = Some(*s); + } + Update::Failed(e) => { + // The old numbers stay -- they beat an empty screen -- but they + // are labelled with the reason and with HOW OLD they are. + // Without the age you cannot tell "could not reach it a second + // ago" from "stuck since yesterday". + let age = self.last.as_ref() + .map(|s| format!(", showing state from {} ago", + view::ago(view::age_of(s.at)))) + .unwrap_or_default(); + self.status = format!("! not updated: {e}{age}"); + } + Update::Note(n) => self.note = Some((n, SystemTime::now())), } - Update::Note(n) => self.note = Some((n, SystemTime::now())), - } - dirty = true; - } - Err(mpsc::TryRecvError::Empty) => break, - Err(mpsc::TryRecvError::Disconnected) => { - // The worker thread is gone. Nothing will be read again, and the alternative - // to saying so is a menu that keeps showing a reading from the moment it died. - if !self.worker_dead { - self.worker_dead = true; - self.status = "! the reader thread has died -- nothing here will update \ - again; quit and start the tray anew".into(); dirty = true; } - break; + Err(mpsc::TryRecvError::Empty) => break, + Err(mpsc::TryRecvError::Disconnected) => { + // The worker thread is gone. Nothing will be read again, and the + // alternative to saying so is a menu that keeps showing a reading from + // the moment it died. + if !self.worker_dead { + self.worker_dead = true; + self.status = "! the reader thread has died -- nothing here will \ + update again; quit and start the tray anew".into(); + dirty = true; + } + break; + } } } - } - // A worker that is alive but has stopped answering. Without this the menu keeps asserting - // "updated just now" over a reading that never arrived. - if let Some((what, since)) = self.awaiting.clone() { - if since.elapsed() > WORKER_PATIENCE && !self.status.starts_with("! no answer") { - self.status = format!("! no answer about {what} for {} -- the reader is stuck", - ago(since.elapsed())); - dirty = true; + // A worker that is alive but has stopped answering. Without this the menu keeps + // asserting "updated just now" over a reading that never arrived. + if let Some((what, since)) = self.awaiting.clone() { + if since.elapsed() > WORKER_PATIENCE && !self.status.starts_with("! no answer") { + self.status = format!("! no answer about {what} for {} -- the reader is \ + stuck", view::ago(since.elapsed())); + dirty = true; + } } - } - while let Ok(ev) = MenuEvent::receiver().try_recv() { - if Some(&ev.id) == self.items.quit.as_ref().map(|i| i.id()) { - el.exit(); - return; - } - if Some(&ev.id) == self.items.refresh.as_ref().map(|i| i.id()) { - self.ask(Command::Refresh, "refreshing…", "a reading"); + while let Ok(ev) = MenuEvent::receiver().try_recv() { + if Some(&ev.id) == self.items.quit.as_ref().map(|i| i.id()) { + return Tick::Quit; + } + if Some(&ev.id) == self.items.refresh.as_ref().map(|i| i.id()) { + self.ask(Command::Refresh, "refreshing…", "a reading"); + dirty = true; + continue; + } + if let Some((_, label)) = self.items.run.iter().find(|(i, _)| i.id() == &ev.id) { + let (label, what) = (label.clone(), format!("running {label}")); + self.ask(Command::RunJob(label), &format!("{what}…"), &what); + dirty = true; + continue; + } + if let Some((_, label, sched)) = + self.items.sched.iter().find(|(i, _, _)| i.id() == &ev.id) { + let (label, sched) = (label.clone(), sched.clone()); + let what = format!("the schedule of {label}"); + self.ask(Command::SetSchedule(label, sched), &format!("changing {what}…"), + &what); + dirty = true; + continue; + } + // A press on an item from a menu that has since been rebuilt: the ids are new, + // so nothing above matched. Silently dropping it left a person pressing a button + // that did nothing, with no way to tell that from a button that did nothing + // visible. + self.note = Some(("that button was from an older copy of the menu -- press it \ + again".into(), SystemTime::now())); dirty = true; - continue; } - if let Some((_, label)) = self.items.run.iter().find(|(i, _)| i.id() == &ev.id) { - let (label, what) = (label.clone(), format!("running {label}")); - self.ask(Command::RunJob(label), &format!("{what}…"), &what); - dirty = true; - continue; + + if self.last_refresh.elapsed() >= REFRESH { + self.last_refresh = Instant::now(); + self.ask(Command::Refresh, &self.status.clone(), "a reading"); } - if let Some((_, label, sched)) = - self.items.sched.iter().find(|(i, _, _)| i.id() == &ev.id) { - let (label, sched) = (label.clone(), sched.clone()); - let what = format!("the schedule of {label}"); - self.ask(Command::SetSchedule(label, sched), &format!("changing {what}…"), &what); - dirty = true; - continue; + if dirty { + self.rebuild(); } - // A press on an item from a menu that has since been rebuilt: the ids are new, so - // nothing above matched. Silently dropping it left a person pressing a button that - // did nothing, with no way to tell that from a button that did nothing visible. - self.note = Some(("that button was from an older copy of the menu -- press it again" - .into(), SystemTime::now())); - dirty = true; + Tick::Continue } - if self.last_refresh.elapsed() >= REFRESH { - self.last_refresh = Instant::now(); - self.ask(Command::Refresh, &self.status.clone(), "a reading"); - } - if dirty { - self.rebuild(); - } - el.set_control_flow(ControlFlow::WaitUntil(Instant::now() + Duration::from_millis(200))); - } -} + fn rebuild(&mut self) { + let menu = Menu::new(); + let mut items = Items::default(); -impl App { - /// Post a command and remember that an answer is owed. Every path to the worker goes through - /// here so that nothing can be asked without the watchdog knowing about it. - fn ask(&mut self, cmd: Command, status: &str, what: &str) { - self.status = status.to_string(); - self.awaiting = Some((what.to_string(), Instant::now())); - let _ = self.cmd.send(cmd); - } + let _ = menu.append(&MenuItem::new(self.status_line(), false, None)); + // The outcome of the last button press keeps its own line, with its age, until + // another press replaces it. + if let Some((note, when)) = &self.note { + let _ = menu.append(&MenuItem::new( + format!("{note} ({} ago)", view::ago(view::age_of(*when))), false, None)); + } + if let Some(why) = jobs::autostart_fault(jobs::DEFAULT_TRAY_LABEL) { + let _ = menu.append(&MenuItem::new(format!("! {why}"), false, None)); + } + let _ = menu.append(&PredefinedMenuItem::separator()); - fn rebuild(&mut self) { - let menu = Menu::new(); - let mut items = Items::default(); + let mut warned = false; + if let Some(s) = &self.last { + for line in view::summary(&s.stats) { + let _ = menu.append(&MenuItem::new(line, false, None)); + } + let _ = menu.append(&PredefinedMenuItem::separator()); - let _ = menu.append(&MenuItem::new(self.status_line(), false, None)); - // The outcome of the last button press keeps its own line, with its age, until another - // press replaces it. - if let Some((note, when)) = &self.note { - let _ = menu.append(&MenuItem::new(format!("{note} ({} ago)", ago(age_of(*when))), - false, None)); - } - if let Some(why) = jobs::autostart_fault(jobs::DEFAULT_TRAY_LABEL) { - let _ = menu.append(&MenuItem::new(format!("! {why}"), false, None)); - } - let _ = menu.append(&PredefinedMenuItem::separator()); + if let Some(why) = &s.jobs_error { + let _ = menu.append(&MenuItem::new(format!("! jobs unreadable: {why}"), + false, None)); + } + let run = Submenu::new("Run now", true); + for j in &s.jobs { + let item = MenuItem::new(view::job_line(j), true, None); + let _ = run.append(&item); + items.run.push((item, j.label.clone())); + } + let _ = menu.append(&run); - if let Some(s) = &self.last { - for line in summary(&s.stats) { - let _ = menu.append(&MenuItem::new(line, false, None)); - } - let _ = menu.append(&PredefinedMenuItem::separator()); + let sched = Submenu::new("Schedule", true); + for j in &s.jobs { + if matches!(j.schedule, jobs::Schedule::None) || j.broken() { + continue; // nothing to change, or nothing readable to change + } + let sub = Submenu::new(format!("{} ({})", j.short(), j.schedule.human()), + true); + for (label, preset) in PRESETS { + let item = MenuItem::new(*label, true, None); + let _ = sub.append(&item); + items.sched.push((item, j.label.clone(), preset.clone())); + } + let _ = sched.append(&sub); + } + let _ = menu.append(&sched); - if let Some(why) = &s.jobs_error { - let _ = menu.append(&MenuItem::new(format!("! jobs unreadable: {why}"), - false, None)); + warned = view::warned(&s.stats, &s.jobs, &s.jobs_error); } - let run = Submenu::new("Run now", true); - for j in &s.jobs { - let item = MenuItem::new(job_line(j), true, None); - let _ = run.append(&item); - items.run.push((item, j.label.clone())); - } - let _ = menu.append(&run); + // The status line and the last note are as much a part of "does this need a mark" as + // the store's own numbers: a failed refresh has no `Snapshot` to derive a warning + // from, and without this a dead worker or a failed refresh would sit under a calm + // icon forever. + warned = warned + || self.status_line().starts_with('!') + || self.note.as_ref().is_some_and(|(n, _)| n.starts_with('!')); - let sched = Submenu::new("Schedule", true); - for j in &s.jobs { - if matches!(j.schedule, jobs::Schedule::None) || j.broken() { - continue; // nothing to change, or nothing readable to change + let _ = menu.append(&PredefinedMenuItem::separator()); + let refresh = MenuItem::new("Refresh now", true, None); + let _ = menu.append(&refresh); + items.refresh = Some(refresh); + let quit = MenuItem::new("Quit", true, None); + let _ = menu.append(&quit); + items.quit = Some(quit); + + self.items = items; + let title = if self.last.is_none() && !warned { "HM …".to_string() } + else { format!("HM{}", if warned { " !" } else { "" }) }; + let icon = mark_icon(warned); + match self.tray.as_ref() { + Some(t) => { + t.set_menu(Some(Box::new(menu))); + // `set_title` only draws anything on macOS; harmless, and cheaper than a + // second cfg-gated code path, to call it everywhere and let the icon carry + // the mark where there is no menu-bar text to put it in. + let _ = t.set_title(Some(&title)); + let _ = t.set_icon(Some(icon)); } - let sub = Submenu::new(format!("{} ({})", j.short(), j.schedule.human()), true); - for (label, preset) in PRESETS { - let item = MenuItem::new(*label, true, None); - let _ = sub.append(&item); - items.sched.push((item, j.label.clone(), preset.clone())); + None => { + // Not `.ok()`: an icon that never appeared leaves a process running with no + // way to see anything, and the service manager counts it as healthy. Better + // to say so and stop. + match TrayIconBuilder::new() + .with_menu(Box::new(menu)) + .with_icon(icon) + .with_title(title) + .with_tooltip("HyperMnesia") + .build() { + Ok(t) => self.tray = Some(t), + Err(e) => { + eprintln!("hypermnesia: the tray icon could not be created: {e}"); + std::process::exit(1); + } + } } - let _ = sched.append(&sub); } - let _ = menu.append(&sched); } - let _ = menu.append(&PredefinedMenuItem::separator()); - let refresh = MenuItem::new("Refresh now", true, None); - let _ = menu.append(&refresh); - items.refresh = Some(refresh); - let quit = MenuItem::new("Quit", true, None); - let _ = menu.append(&quit); - items.quit = Some(quit); - - self.items = items; - let title = self.title(); // computed before borrowing self.tray - match self.tray.as_mut() { - Some(t) => { - t.set_menu(Some(Box::new(menu))); - let _ = t.set_title(Some(title)); + /// The first line of the menu: how old the numbers are, said in the present tense. + /// + /// Computed here rather than stored, because a stored sentence cannot age. `at` is wall + /// clock, so a machine that slept for two hours reports two hours, not "just now". + fn status_line(&self) -> String { + if !self.status.is_empty() { + return self.status.clone(); } - None => { - // Not `.ok()`: an icon that never appeared leaves a process running with no way - // to see anything, and launchd counts it as healthy. Better to say so and stop. - match TrayIconBuilder::new() - .with_menu(Box::new(menu)) - .with_title(title) - .with_tooltip("HyperMnesia") - .build() { - Ok(t) => self.tray = Some(t), - Err(e) => { - eprintln!("hypermnesia: the menu-bar icon could not be created: {e}"); - std::process::exit(1); + match &self.last { + None => "loading…".into(), + Some(s) => { + let age = view::age_of(s.at); + let took = s.stats.took.as_secs_f32(); + if age < Duration::from_secs(5) { + format!("updated just now, in {took:.1}s") + } else { + format!("updated {} ago, in {took:.1}s", view::ago(age)) } } } } } - /// The first line of the menu: how old the numbers are, said in the present tense. - /// - /// Computed here rather than stored, because a stored sentence cannot age. `at` is wall - /// clock, so a Mac that slept for two hours reports two hours, not "just now". - fn status_line(&self) -> String { - if !self.status.is_empty() { - return self.status.clone(); - } - match &self.last { - None => "loading…".into(), - Some(s) => { - let age = age_of(s.at); - let took = s.stats.took.as_secs_f32(); - if age < Duration::from_secs(5) { - format!("updated just now, in {took:.1}s") - } else { - format!("updated {} ago, in {took:.1}s", ago(age)) - } - } + /// A flat, solid-colour square: everything this tray's icon needs to say is "fine" or "not + /// fine", which two colours already say without a single asset file. 22x22 is a comfortable + /// size for a Linux system tray at typical DPI; macOS scales `set_icon`'s image itself. + fn mark_icon(warned: bool) -> Icon { + const SIZE: u32 = 22; + let (r, g, b) = if warned { (196, 60, 48) } else { (52, 150, 90) }; + let mut rgba = Vec::with_capacity((SIZE * SIZE * 4) as usize); + for _ in 0..(SIZE * SIZE) { + rgba.extend_from_slice(&[r, g, b, 255]); } + Icon::from_rgba(rgba, SIZE, SIZE).expect("a fixed-size solid RGBA buffer is always valid") } +} + +// The tray needs a GUI toolkit: winit + tray-icon on macOS (declared unconditionally, since that +// is where the schedules it edits live), tray-icon + GTK on Linux (declared only behind the +// `tray` feature, so the data layer and the four CLI tools keep building on a machine with no GUI +// libraries at all -- which is what CI's bare runner is). +#[cfg(target_os = "macos")] +mod mac { + use std::sync::mpsc; + use std::time::Duration; - /// What is visible without opening the menu. A mark matters more than a number here: the - /// menu bar is where a problem is NOTICED, not where a report is read. - fn title(&self) -> String { - if self.status_line().starts_with('!') - || self.note.as_ref().is_some_and(|(n, _)| n.starts_with('!')) { - return "HM !".into(); + use winit::application::ApplicationHandler; + use winit::event_loop::{ActiveEventLoop, ControlFlow, EventLoop}; + + use super::app::{self, App, Command, Tick, Update}; + + pub fn main() { + let args: Vec = std::env::args().skip(1).collect(); + if app::handle_early_args(&args) { + return; } - match &self.last { - None => "HM …".into(), - Some(s) => { - // Derived from the lines the menu actually shows, not recomputed beside them. - // Recomputing is what let the two drift: "! not embedded: 4300" sat in the menu - // under a calm "HM", while "! embedding models: 2" -- printed by the same - // function, two lines later -- did mark the title. - // - // A non-zero exit code is deliberately NOT a mark: the freshness job exits 1 to - // mean "discrepancies found", by design, and a permanent "!" is one nobody looks - // at. The code is on the job's own line instead. - let warn = summary(&s.stats).iter().any(|l| l.starts_with('!')) - || s.jobs_error.is_some() - || s.jobs.iter().any(Job::broken) - || s.jobs.iter().any(Job::overdue); - format!("HM{}", if warn { " !" } else { "" }) - } + + let event_loop = EventLoop::builder().build().expect("event loop"); + event_loop.set_control_flow(ControlFlow::WaitUntil( + std::time::Instant::now() + Duration::from_millis(200))); + + let (tx, rx) = mpsc::channel::(); + let (cmd_tx, cmd_rx) = mpsc::channel::(); + app::spawn_worker(tx, cmd_rx); + // Read immediately: an empty menu at startup looks like a broken console. + let _ = cmd_tx.send(Command::Refresh); + + let mut handler = Handler { app: App::new(rx, cmd_tx) }; + if let Err(e) = event_loop.run_app(&mut handler) { + eprintln!("hypermnesia: the event loop ended: {e}"); } } -} -fn summary(s: &Stats) -> Vec { - let (docs, chunks, embedded) = s.corpus.iter() - .filter(|r| !r.is_memory_page()) - .fold((0, 0, 0), |(d, c, e), r| (d + r.docs, c + r.chunks, e + r.embedded)); - let mut out = vec![ - format!("Memory: {} active of {}", s.memories_active, s.memories_total), - format!("Knowledge pages: {}", s.pages), - format!("Corpus: {docs} docs, {chunks} chunks"), - ]; - if embedded < chunks { - out.push(format!("! not embedded: {}", chunks - embedded)); + struct Handler { + app: App, } - if s.embedding_models.len() > 1 { - // Two models in one store means part of the corpus cannot be reached by meaning at all. - out.push(format!("! embedding models: {}", s.embedding_models.len())); - } - // Marked when there is something in it, so the title can be derived from these lines rather - // than from a second list of conditions kept in step by hand. - out.push(if s.review_pending > 0 { - format!("! Review queue: {} (oldest {} days)", s.review_pending, s.review_oldest_days) - } else { - "Review queue: 0".to_string() - }); - out.push(format!("Stale: {}", s.stale)); - out.push(format!("Database: {}", s.db_size)); - out -} -/// How old something is, by the wall clock. `Instant` would stop while the Mac sleeps. -fn age_of(t: SystemTime) -> Duration { - SystemTime::now().duration_since(t).unwrap_or_default() -} + impl ApplicationHandler for Handler { + fn resumed(&mut self, _: &ActiveEventLoop) { + // The first `tick()` builds the tray if it does not exist yet. + let _ = self.app.tick(); + } -fn ago(d: Duration) -> String { - let s = d.as_secs(); - if s < 90 { format!("{s}s") } - else if s < 5400 { format!("{}m", s / 60) } - else if s < 172_800 { format!("{}h", s / 3600) } - else { format!("{}d", s / 86_400) } -} + fn window_event(&mut self, _: &ActiveEventLoop, _: winit::window::WindowId, + _: winit::event::WindowEvent) {} -fn job_line(j: &Job) -> String { - if let Some(why) = &j.fault { - // launchd refused this plist too, so the job is not running. It used to be missing from - // the menu entirely. - let first = why.lines().next().unwrap_or(why); - return format!("{} (! unreadable plist: {first})", j.short()); - } - let state = if j.running() { - "running".to_string() - } else if j.overdue() { - "! overdue".to_string() - } else if j.never_ran() { - "never run yet".to_string() - } else if let Some(_) = j.freshness_unknown() { - // Neither fresh nor stale: the log that would date it is gone or undatable. Drawing - // that as health is how a job that quietly stopped stays invisible. - "? cannot be dated".to_string() - } else if j.last_exit.is_none() { - "not loaded".to_string() - } else if j.runs == Some(0) { - // launchd prints exit 0 both for "finished successfully" and for "never finished at - // all". With no run since this login there is nothing to call successful. - "no run since login".to_string() - } else { - // The exit code belongs here. Without it a job that fails on every single run reads - // exactly like one that works: same schedule, same fresh log timestamp. - let when = match j.since_last_output() { - Some(d) => format!("{} ago", ago(d)), - None => "—".to_string(), - }; - match j.last_exit { - Some(0) => when, - Some(c) => format!("! exit {c}, {when}"), - None => format!("{when}, not loaded"), + fn about_to_wait(&mut self, el: &ActiveEventLoop) { + if let Tick::Quit = self.app.tick() { + el.exit(); + return; + } + el.set_control_flow(ControlFlow::WaitUntil( + std::time::Instant::now() + Duration::from_millis(200))); } - }; - format!("{} ({}, {})", j.short(), j.schedule.human(), state) + } } -#[cfg(test)] -mod tests { - use super::*; - use std::path::PathBuf; - - fn job(schedule: jobs::Schedule) -> Job { - Job { - label: "com.hypermnesia.extract".into(), plist: PathBuf::new(), schedule, - program: vec![], log: Some(PathBuf::from("/nope")), last_exit: Some(0), pid: None, - log_state: jobs::LogState::Written(SystemTime::now() - Duration::from_secs(300)), - runs: Some(9), - installed: Some(SystemTime::now() - Duration::from_secs(90_000)), fault: None, +#[cfg(all(target_os = "linux", feature = "tray"))] +mod linux { + use std::cell::RefCell; + use std::rc::Rc; + use std::sync::mpsc; + + use super::app::{self, App, Command, Tick, Update}; + + pub fn main() { + let args: Vec = std::env::args().skip(1).collect(); + if app::handle_early_args(&args) { + return; } - } - /// A job that fails on every run used to read exactly like one that works: same schedule, - /// same fresh log timestamp, no code anywhere in the menu. - #[test] - fn a_failing_job_says_so_in_its_own_line() { - let mut j = job(jobs::Schedule::Every(14_400)); - assert_eq!(job_line(&j), "extract (every 4 h, 5m ago)"); - - j.last_exit = Some(2); - let line = job_line(&j); - assert!(line.contains("exit 2"), "the exit code has to be on the line: {line}"); - assert!(line.starts_with("extract (every 4 h, !"), "and marked: {line}"); - } + if gtk::init().is_err() { + eprintln!("hypermnesia: could not initialize GTK -- is a display / Wayland or X11 \ + session available?"); + std::process::exit(1); + } - /// An unreadable plist is one launchd also refused, so the job is not running. It used to be - /// absent from the menu altogether. - #[test] - fn an_unreadable_plist_is_a_visible_line() { - let mut j = job(jobs::Schedule::None); - j.fault = Some("plutil: unexpected character\nsecond line".into()); - let line = job_line(&j); - assert!(line.contains("unreadable plist"), "{line}"); - assert!(!line.contains("second line"), "one line only, it is a menu: {line}"); - } + let (tx, rx) = mpsc::channel::(); + let (cmd_tx, cmd_rx) = mpsc::channel::(); + app::spawn_worker(tx, cmd_rx); + let _ = cmd_tx.send(Command::Refresh); + + let app = Rc::new(RefCell::new(App::new(rx, cmd_tx))); + // The first tick builds the tray immediately: an empty menu bar at startup looks like a + // broken console, same reasoning as `resumed()` on macOS. + let _ = app.borrow_mut().tick(); + + glib::source::timeout_add_local(std::time::Duration::from_millis(200), move || { + match app.borrow_mut().tick() { + Tick::Quit => { + gtk::main_quit(); + glib::ControlFlow::Break + } + Tick::Continue => glib::ControlFlow::Continue, + } + }); - #[test] - fn a_job_never_run_is_not_called_fresh() { - let mut j = job(jobs::Schedule::Calendar(jobs::Cal { hour: Some(6), minute: Some(10), day: None, weekday: Some(1), month: None })); - j.runs = Some(0); - j.log_state = jobs::LogState::Missing; - assert!(job_line(&j).contains("never run yet"), "{}", job_line(&j)); + gtk::main(); } } -} #[cfg(target_os = "macos")] fn main() { mac::main() } -#[cfg(not(target_os = "macos"))] +#[cfg(all(target_os = "linux", feature = "tray"))] +fn main() { linux::main() } + +#[cfg(all(target_os = "linux", not(feature = "tray")))] +fn main() { + eprintln!("hypermnesia: this build has no tray -- rebuild console with `--features tray` \ + (needs GTK3 and libayatana-appindicator development headers). \ + hypermnesia-stats, -jobs, -settings and -setup work as built."); + std::process::exit(1); +} + +#[cfg(not(any(target_os = "macos", target_os = "linux")))] fn main() { - eprintln!("hypermnesia: the menu-bar tray is macOS-only -- it reads launchd and draws in the \ - system menu bar. hypermnesia-stats, -jobs, -settings and -setup work here."); + eprintln!("hypermnesia: the tray is built for macOS and Linux only. \ + hypermnesia-stats, -jobs, -settings and -setup work here."); std::process::exit(1); } diff --git a/console/src/jobs.rs b/console/src/jobs/launchd.rs similarity index 58% rename from console/src/jobs.rs rename to console/src/jobs/launchd.rs index dcf0e95..73d00fb 100644 --- a/console/src/jobs.rs +++ b/console/src/jobs/launchd.rs @@ -1,321 +1,15 @@ -//! The pipeline's scheduled jobs: what is configured, when it last worked, and how to run it now. +//! The launchd backend: `*.plist` files under `~/Library/LaunchAgents`, read through +//! `plutil -convert json` and `launchctl list`/`print`/`kickstart`/`bootout`/`bootstrap`. //! -//! All of it lives in launchd, so this is the macOS-specific part of the console. We read the -//! `*.plist` files through `plutil -convert json` and take their state from `launchctl list`. No -//! plist parser of our own: the format is binary about as often as it is text, and a homegrown +//! No plist parser of our own: the format is binary about as often as it is text, and a homegrown //! implementation would break silently on the very first binary file. -//! -//! The important part about "when it last worked": launchd does not keep that. It knows the exit -//! code of the LAST run and nothing else. So the time comes from the mtime of the log file -- and -//! a missing log means "it has not run once since the path was set", not "it is working quietly". -//! This is not a nicety: our reflect job is set to run weekly, and there is no log file at all. use std::collections::BTreeMap; use std::path::PathBuf; use std::process::Command; use std::time::{Duration, SystemTime}; -/// Reverse-DNS prefix of the launchd labels this console manages. -pub const DEFAULT_JOB_PREFIX: &str = "com.hypermnesia"; - -/// Label of the tray's own autostart job, passed to `install_self`. -pub const DEFAULT_TRAY_LABEL: &str = "com.hypermnesia.tray"; - -/// The label prefix to look for: environment, then the config file, then the default -- the same -/// order as every other setting here. -/// -/// The config file matters more than it looks. The tray is started by launchd, which gives it a -/// minimal environment and none of a login shell's variables, so an installation whose jobs are -/// named differently could set HM_JOB_PREFIX in its shell forever and the tray would still show -/// an empty job list -- while every command run by hand showed the right one. -pub fn job_prefix() -> String { - if let Some(v) = std::env::var("HM_JOB_PREFIX").ok().filter(|s| !s.is_empty()) { - return v; - } - // Filtered the same way as the environment: an empty HM_JOB_PREFIX in the config makes - // `starts_with` match every LaunchAgent the person owns, and this module edits and kickstarts - // what it lists. - if let Some(v) = crate::config().get("HM_JOB_PREFIX").filter(|s| !s.is_empty()) { - return v.clone(); - } - DEFAULT_JOB_PREFIX.to_string() -} - -/// One calendar slot, as launchd stores it. -/// -/// Every field is optional AND an omitted field is a WILDCARD, which is the whole reason this is -/// a struct rather than an hour and a minute. `{"Minute": 15}` means every hour at :15, not -/// "daily 00:15"; `{"Day": 1, "Hour": 3}` means the first of each month, not "daily 03:00". -/// Reading the omissions as zeroes printed a plausible wrong time -- and, worse, a wrong period: -/// a healthy monthly job was measured against a day and a half and reported overdue for -/// twenty-nine days out of thirty. -#[derive(Debug, Clone, PartialEq, Default)] -pub struct Cal { - pub minute: Option, - pub hour: Option, - /// Day of the month, 1-31. - pub day: Option, - /// 0 = Sunday. - pub weekday: Option, - pub month: Option, -} - -impl Cal { - pub fn at(hour: u32, minute: u32, weekday: Option) -> Self { - Self { hour: Some(hour), minute: Some(minute), weekday, ..Self::default() } - } - - /// How long between two firings, taken from the COARSEST field that is pinned: everything - /// finer than it repeats inside that cycle, everything coarser is a wildcard. - pub fn period(&self) -> Duration { - let day = 86_400; - Duration::from_secs(if self.month.is_some() { - 365 * day - } else if self.day.is_some() { - 31 * day - } else if self.weekday.is_some() { - 7 * day - } else if self.hour.is_some() { - day - } else if self.minute.is_some() { - 3_600 - } else { - // Nothing pinned at all: launchd fires such a job every minute. - 60 - }) - } - - pub fn human(&self) -> String { - const DAYS: [&str; 7] = ["Sun", "Mon", "Tue", "Wed", "Thu", "Fri", "Sat"]; - const MONTHS: [&str; 12] = ["Jan", "Feb", "Mar", "Apr", "May", "Jun", - "Jul", "Aug", "Sep", "Oct", "Nov", "Dec"]; - // "xx" where launchd would repeat: an hour with no minute fires every minute of it. - let time = match (self.hour, self.minute) { - (Some(h), Some(m)) => format!("{h:02}:{m:02}"), - (Some(h), None) => format!("{h:02}:xx"), - (None, Some(m)) => format!(":{m:02}"), - (None, None) => String::new(), - }; - let month = self.month - .map(|m| format!("{} ", MONTHS[((m.max(1) - 1) as usize) % 12])) - .unwrap_or_default(); - match (self.month, self.day, self.weekday, self.hour, self.minute) { - (_, Some(d), _, _, _) => format!("{} {month}day {d} {time}", - if self.month.is_some() { "yearly" } else { "monthly" }), - (_, None, Some(w), _, _) => format!("weekly {} {time}", DAYS[(w as usize) % 7]), - (_, None, None, Some(_), _) => format!("daily {time}"), - (_, None, None, None, Some(_)) => format!("hourly at {time}"), - _ => "every minute".into(), - } - } -} - -#[derive(Debug, Clone, PartialEq)] -pub enum Schedule { - /// Every N seconds. - Every(u64), - /// One calendar slot. - Calendar(Cal), - /// Several calendar slots: launchd accepts an array of dicts. All of them are kept -- the - /// period has to come from the widest of them, and the first slot is not the schedule. - Several(Vec), - /// No schedule: a service that is simply kept running. - None, -} - -impl Schedule { - /// A one-slot calendar schedule, which is what the console can set. - pub fn at(hour: u32, minute: u32, weekday: Option) -> Self { - Schedule::Calendar(Cal::at(hour, minute, weekday)) - } - - pub fn human(&self) -> String { - match self { - Schedule::Every(s) if *s % 3600 == 0 => format!("every {} h", s / 3600), - Schedule::Every(s) if *s % 60 == 0 => format!("every {} min", s / 60), - Schedule::Every(s) => format!("every {s} s"), - Schedule::Calendar(c) => c.human(), - Schedule::Several(c) if c.len() == 1 => c[0].human(), - Schedule::Several(c) => format!("{} slots ({})", c.len(), c[0].human()), - Schedule::None => "on demand".into(), - } - } -} - -/// What is known about a job's log. -#[derive(Debug, Clone, PartialEq)] -pub enum LogState { - /// The plist names no log: nothing can be dated by it, in either direction. - NotConfigured, - /// A log is configured and it is not there -- the job has not written anything since the - /// path was set. - Missing, - /// The log was last written at this time. - Written(SystemTime), - /// A log is configured, exists, and cannot be dated: unreadable, or stamped in the future - /// (a clock that moved). Not evidence of health and not evidence of failure -- and saying so - /// beats both of the alternatives. - Unknown(String), -} - -impl LogState { - pub fn at(&self) -> Option { - match self { - LogState::Written(t) => Some(*t), - _ => None, - } - } -} - -#[derive(Debug, Clone)] -pub struct Job { - pub label: String, - pub plist: PathBuf, - pub schedule: Schedule, - pub program: Vec, - pub log: Option, - /// Exit code of the last run, as launchd remembers it. `None` -- the job is not loaded. - pub last_exit: Option, - /// When the plist last changed. Needed so as not to raise a false alarm: a job installed - /// later than its own last slot simply has not come due yet. - pub installed: Option, - /// How many times launchd has started it. Zero -- not once. - /// - /// Without this number the status column lies: `launchctl list` prints `0` in the status - /// column both for "finished successfully" and for "never finished at all", so a job that has - /// never started looks like one that ran without errors. Seen on reflect: the list showed - /// "last run ok" with `runs = 0` and `last exit code = (never exited)`. - pub runs: Option, - /// PID, if the job is running right now. - pub pid: Option, - /// What the log says about when the job last worked. The only sign that survives a reboot, - /// and the reason it is three states rather than a timestamp: "no log configured", "a log is - /// configured and is not there" and "unreadable / dated in the future" are three different - /// pieces of knowledge, and collapsing them into None made a rotated-away log read as a - /// healthy job forever. - pub log_state: LogState, - /// Why this job could not be read. A plist in LaunchAgents that `plutil` refuses is a plist - /// launchd also refused at login -- the job is not running, and that is precisely the case - /// the console used to have nothing to say about, because the row was dropped from the list. - pub fault: Option, -} - -impl Job { - /// Short name without the reverse-DNS prefix: on the console's screen `extract`, not - /// `com.hypermnesia.extract`. - pub fn short(&self) -> &str { - self.label.rsplit(['.', '-']).next().unwrap_or(&self.label) - } - - pub fn running(&self) -> bool { - self.pid.is_some() - } - - /// How long until the job repeats. For calendar schedules -- a day or a week; the exact date - /// of the next run is not needed, only the order of magnitude. - pub fn period(&self) -> Option { - match &self.schedule { - Schedule::Every(s) => Some(Duration::from_secs(*s)), - Schedule::Calendar(c) => Some(c.period()), - // The widest of the slots, not the narrowest: two weekly slots are still a week - // apart at worst, and measuring them against a day would cry wolf for six days of it. - Schedule::Several(cals) => cals.iter().map(Cal::period).max(), - Schedule::None => None, - } - } - - /// Something is wrong with the job itself, before any question of when it last ran. - pub fn broken(&self) -> bool { - self.fault.is_some() - } - - /// launchd has never started it. - /// - /// `runs` alone does not answer this. It is per-bootstrap state: launchd loads every - /// LaunchAgent again at each login and the counter starts at zero, while the plist's mtime - /// does not move. Trusting it turned every healthy job into "never ran / overdue" after a - /// reboot. So the log -- the only evidence that survives a restart -- has to agree, and it - /// only counts as evidence if it was written AFTER the job was installed: a log left behind - /// by a previous installation says nothing about this one. - pub fn never_ran(&self) -> bool { - if self.schedule == Schedule::None || self.broken() { - return false; // a service without a schedule -- not a complaint - } - if self.log_after_install() { - return false; // it has written something since: it ran - } - match self.runs { - Some(n) => n == 0, - // The counter is unavailable (the job is not loaded): fall back to the sign "a log is - // configured but not there" -- weaker, but better than silence. - None => self.log_state == LogState::Missing, - } - } - - /// Was the log written after this version of the job was installed? A log from a previous - /// install is not evidence about this one -- in either direction. - fn log_after_install(&self) -> bool { - match (self.log_state.at(), self.installed) { - (Some(out), Some(inst)) => out >= inst, - (Some(_), None) => true, - _ => false, - } - } - - /// Has never run AND would have had time to -- or ran once and then stopped. THAT is the - /// complaint; `never_ran` on its own is NOT. - /// - /// The difference is not cosmetic. A weekly job installed on Monday afternoon honestly shows - /// zero runs until the next Monday, and shouting at it would be a lie. The signal must fire - /// on "a whole period has passed and nothing happened", otherwise people learn to scroll - /// past it and it stops working on the day it is needed. - /// - /// A period and a half is the threshold, not one period: a job whose slot is due about now - /// has not failed yet. - pub fn overdue(&self) -> bool { - let Some(period) = self.period() else { return false }; - let slack = period + period / 2; - if self.never_ran() { - // The only date available is the plist's own. - return self.since_installed().map(|age| age > slack).unwrap_or(true); - } - // It ran at some point. Then the complaint is staleness, dated by the log -- and never - // from before the job was installed, or a fresh install inheriting an old log would be - // called overdue the moment it was made. - match (self.since_last_output(), self.since_installed()) { - (Some(age), Some(since_install)) => age.min(since_install) > slack, - (Some(age), None) => age > slack, - (None, _) => false, - } - } - - /// Can this job's freshness be judged at all? A configured log that is gone or undatable - /// leaves the question open -- and an open question must not be drawn as health. - pub fn freshness_unknown(&self) -> Option { - if self.broken() || self.schedule == Schedule::None { - return None; - } - match &self.log_state { - LogState::Unknown(why) => Some(why.clone()), - // Ran at least once since this boot, yet the log it is supposed to write is not - // there: it was rotated away, or the job is writing nowhere. - LogState::Missing if self.runs.unwrap_or(0) > 0 => - Some("the configured log is not there, so nothing can date its last run".into()), - LogState::NotConfigured if self.runs.unwrap_or(0) > 0 => - Some("no log is configured, so nothing survives a reboot to date its runs".into()), - _ => None, - } - } - - /// How long ago it was installed (by the plist's modification time). - pub fn since_installed(&self) -> Option { - self.installed.and_then(|t| SystemTime::now().duration_since(t).ok()) - } - - pub fn since_last_output(&self) -> Option { - self.log_state.at().and_then(|t| SystemTime::now().duration_since(t).ok()) - } -} +use super::{log_state_of, Cal, Job, LogState, Schedule}; /// Where the user's own LaunchAgents live. /// @@ -373,7 +67,7 @@ fn broken_job(path: &PathBuf, name: &str, fault: String, let (pid, last_exit) = states.get(name).cloned().unwrap_or((None, None)); Job { label: name.to_string(), - plist: path.clone(), + source: path.clone(), schedule: Schedule::None, program: Vec::new(), log: None, @@ -382,6 +76,7 @@ fn broken_job(path: &PathBuf, name: &str, fault: String, runs: runs_of(name), pid, log_state: LogState::NotConfigured, + trigger: super::TriggerState::NotTracked, fault: Some(fault), } } @@ -404,28 +99,12 @@ fn read_plist(path: &PathBuf, states: &BTreeMap, Option) -> LogState { - let Some(p) = log else { return LogState::NotConfigured }; - let md = match std::fs::metadata(p) { - Ok(m) => m, - Err(e) if e.kind() == std::io::ErrorKind::NotFound => return LogState::Missing, - Err(e) => return LogState::Unknown(format!("{}: {e}", p.display())), - }; - match md.modified() { - Ok(t) if t > SystemTime::now() + Duration::from_secs(60) => - LogState::Unknown(format!("{} is stamped in the future -- a clock moved", - p.display())), - Ok(t) => LogState::Written(t), - Err(e) => LogState::Unknown(format!("{}: {e}", p.display())), - } -} - /// What a plist says, before anything is asked of launchd or of the filesystem. Separate from /// `read_plist` so it can be tested against real plutil output without a Mac in the loop. struct Plist { @@ -582,7 +261,7 @@ pub fn set_schedule(job: &Job, sched: &Schedule) -> Result { if let Some(why) = job.fault.as_ref() { return Err(format!("this plist could not be read ({why}) -- I will not edit it")); } - let path = job.plist.to_string_lossy().to_string(); + let path = job.source.to_string_lossy().to_string(); let backup = format!("{path}.bak"); // A .bak already there is the leftover of an edit that did not finish -- a crash between // removing the schedule keys and writing the new ones. Overwriting it with the CURRENT file @@ -821,18 +500,6 @@ pub fn uninstall_self(label: &str) -> Result { } } -/// The last `n` lines of a job's log. -pub fn tail(job: &Job, n: usize) -> Result { - let Some(path) = job.log.as_ref() else { - return Err("the job has no log file configured".into()); - }; - let data = std::fs::read(path) - .map_err(|e| format!("{}: {e}", path.display()))?; - let text = String::from_utf8_lossy(&data); - let lines: Vec<&str> = text.lines().collect(); - Ok(lines[lines.len().saturating_sub(n)..].join("\n")) -} - #[cfg(test)] mod tests { use super::*; @@ -845,27 +512,12 @@ mod tests { fn job(schedule: Schedule) -> Job { Job { - label: "x".into(), plist: PathBuf::new(), schedule, program: vec![], log: None, + label: "x".into(), source: PathBuf::new(), schedule, program: vec![], log: None, last_exit: None, pid: None, log_state: LogState::NotConfigured, runs: None, - installed: None, fault: None, + installed: None, trigger: super::TriggerState::NotTracked, fault: None, } } - /// A job with a log that was written `ago` ago, installed `installed_ago` ago. - fn with_log(schedule: Schedule, runs: Option, written: Option, - installed_ago: Duration) -> Job { - let mut j = job(schedule); - j.log = Some(PathBuf::from("/nope")); - j.runs = runs; - j.last_exit = Some(0); - j.installed = Some(SystemTime::now() - installed_ago); - j.log_state = match written { - Some(d) => LogState::Written(SystemTime::now() - d), - None => LogState::Missing, - }; - j - } - #[test] fn reads_a_weekly_calendar_schedule() { let p = parse_plist_json(REFLECT).expect("parse"); @@ -935,94 +587,6 @@ mod tests { assert!(parse_plist_json("not json at all").is_err()); } - #[test] - fn schedules_render_in_words() { - assert_eq!(Schedule::Every(14400).human(), "every 4 h"); - assert_eq!(Schedule::Every(900).human(), "every 15 min"); - // Something not a whole number of minutes stays in seconds instead of being rounded into - // a lie. - assert_eq!(Schedule::Every(90).human(), "every 90 s"); - assert_eq!(Schedule::at(5, 30, None).human(), "daily 05:30"); - assert_eq!(Schedule::at(6, 10, Some(1)).human(), "weekly Mon 06:10"); - assert_eq!(Schedule::None.human(), "on demand"); - } - - /// A job launchd has never started must be called exactly that, not "successful". - #[test] - fn a_scheduled_job_that_never_ran_is_flagged() { - let mut j = with_log(Schedule::at(6, 10, Some(1)), Some(0), None, - Duration::from_secs(30 * 86_400)); - assert!(j.never_ran()); - - j.runs = Some(44); // launchd started it since the last login - assert!(!j.never_ran()); - - j.runs = Some(0); - j.schedule = Schedule::None; // a service without a schedule -- not a complaint - assert!(!j.never_ran()); - - // The counter is unavailable (the job is not loaded): fall back to "a log is configured - // but not there". - j.schedule = Schedule::Every(3600); - j.runs = None; - assert!(j.never_ran()); - j.log_state = LogState::Written(SystemTime::now()); - assert!(!j.never_ran()); - } - - /// The run counter is per-bootstrap: launchd resets it at every login while the plist's - /// mtime stays where it was. Keying "never ran" on it alone turned every healthy job into a - /// false alarm after a reboot -- which is how a warning stops being read. - #[test] - fn a_reboot_does_not_turn_a_working_job_into_a_false_alarm() { - let j = with_log(Schedule::Every(14_400), Some(0), Some(Duration::from_secs(600)), - Duration::from_secs(30 * 86_400)); - assert!(!j.never_ran(), "the log says it ran ten minutes ago"); - assert!(!j.overdue()); - } - - /// The other direction of the same mistake: a job that ran once and then stopped used to be - /// exempt forever, because a positive counter ended the question. - #[test] - fn a_job_that_ran_once_and_then_stalled_is_still_overdue() { - let mut j = with_log(Schedule::Every(3600), Some(12), Some(Duration::from_secs(600)), - Duration::from_secs(30 * 86_400)); - assert!(!j.overdue(), "ten minutes into an hourly job is fine"); - - j.log_state = LogState::Written(SystemTime::now() - Duration::from_secs(6 * 3600)); - assert!(j.overdue(), "six hours into an hourly job is not"); - } - - /// A log left behind by a PREVIOUS installation is not evidence about this one. Without - /// that rule a job installed a minute ago, pointing at a month-old log, was called overdue - /// on the spot. - #[test] - fn an_inherited_log_does_not_condemn_a_fresh_install() { - let j = with_log(Schedule::at(6, 10, Some(1)), Some(0), - Some(Duration::from_secs(30 * 86_400)), // log: a month old - Duration::from_secs(60)); // installed a minute ago - assert!(!j.overdue(), "it has not had a chance to run yet"); - assert!(j.never_ran(), "and the old log is not evidence that it did"); - } - - /// Freshness that cannot be judged must say so. A rotated-away log used to leave a stalled - /// job reading "last run ok" with an empty timestamp, for ever. - #[test] - fn a_log_that_cannot_date_a_run_is_an_open_question_not_health() { - let j = with_log(Schedule::Every(3600), Some(9), None, Duration::from_secs(86_400)); - assert!(j.freshness_unknown().is_some(), "the log is configured and gone"); - assert!(!j.overdue(), "and that is not the same as proof it stopped"); - - let mut future = with_log(Schedule::Every(3600), Some(9), Some(Duration::from_secs(60)), - Duration::from_secs(86_400)); - future.log_state = LogState::Unknown("stamped in the future".into()); - assert!(future.freshness_unknown().is_some()); - - let healthy = with_log(Schedule::Every(3600), Some(9), Some(Duration::from_secs(60)), - Duration::from_secs(86_400)); - assert!(healthy.freshness_unknown().is_none()); - } - /// A plist that cannot be read is kept as a row with a fault, never dropped. #[test] fn an_unreadable_plist_becomes_a_row_not_a_gap() { @@ -1035,22 +599,6 @@ mod tests { assert!(!j.overdue()); } - #[test] - fn period_matches_the_schedule() { - assert_eq!(job(Schedule::Every(14400)).period(), Some(Duration::from_secs(14400))); - assert_eq!(job(Schedule::at(5, 30, None)).period(), Some(Duration::from_secs(86_400))); - assert_eq!(job(Schedule::at(6, 10, Some(1))).period(), - Some(Duration::from_secs(7 * 86_400))); - assert_eq!(job(Schedule::None).period(), None); - } - - #[test] - fn short_name_drops_the_reverse_dns_prefix() { - let mut j = job(Schedule::None); - j.label = "com.hypermnesia.consolidate".into(); - assert_eq!(j.short(), "consolidate"); - } - /// Quit has to stick. `KeepAlive` set unconditionally relaunches the tray about ten seconds /// after the person chooses Quit, which makes the menu item a lie. Conditional on a failed /// exit keeps the intent -- restart a tray that crashed. diff --git a/console/src/jobs/mod.rs b/console/src/jobs/mod.rs new file mode 100644 index 0000000..f33fdf7 --- /dev/null +++ b/console/src/jobs/mod.rs @@ -0,0 +1,632 @@ +//! The pipeline's scheduled jobs: what is configured, when it last worked, and how to run it now. +//! +//! What is asked here -- a schedule, a last exit code, how long since it last did anything, a way +//! to run it now -- is the same question on every platform. The answer comes from a different +//! place on each: launchd's plists and `launchctl` on macOS, systemd user units and `systemctl +//! --user` on Linux. This module holds the question -- `Cal`, `Schedule`, `LogState`, `Job` and +//! the freshness logic in `impl Job` -- once, and knows nothing about either backend. `launchd` +//! and `systemd` hold the answers, selected by `cfg(target_os)`, and nothing above this module +//! (the tray, `hypermnesia-jobs`) has to know which one is in effect. +//! +//! The important part about "when it last worked" is not the same fact on both platforms. launchd +//! does not keep it: it knows only the exit code of the LAST run, so the time has to come from the +//! mtime of the job's own log file, and a missing log means "it has not run once since the path +//! was set", not "it is working quietly". systemd keeps the answer directly -- a timer's own +//! `LastTriggerUSec` -- which is why `Job` carries both `log_state` and `trigger`: the second is +//! `NotTracked` on macOS and, wherever it is tracked, it is trusted over the log every time. + +use std::path::PathBuf; +use std::time::{Duration, SystemTime}; + +#[cfg(target_os = "macos")] +mod launchd; +#[cfg(target_os = "macos")] +pub use launchd::{autostart_fault, install_self, list, run_now, set_schedule, uninstall_self}; + +#[cfg(target_os = "linux")] +mod systemd; +#[cfg(target_os = "linux")] +pub use systemd::{autostart_fault, install_self, list, run_now, set_schedule, uninstall_self}; + +#[cfg(not(any(target_os = "macos", target_os = "linux")))] +mod unsupported; +#[cfg(not(any(target_os = "macos", target_os = "linux")))] +pub use unsupported::{autostart_fault, install_self, list, run_now, set_schedule, uninstall_self}; + +/// Reverse-DNS prefix of the launchd labels this console manages, on macOS. +#[cfg(target_os = "macos")] +pub const DEFAULT_JOB_PREFIX: &str = "com.hypermnesia"; +/// Prefix of the systemd unit names this console manages, on Linux: `hypermnesia-extract.timer`, +/// not a reverse-DNS label -- that convention is launchd's, not systemd's. +#[cfg(target_os = "linux")] +pub const DEFAULT_JOB_PREFIX: &str = "hypermnesia-"; +#[cfg(not(any(target_os = "macos", target_os = "linux")))] +pub const DEFAULT_JOB_PREFIX: &str = "hypermnesia-"; + +/// Label of the tray's own autostart job, passed to `install_self`. +#[cfg(target_os = "macos")] +pub const DEFAULT_TRAY_LABEL: &str = "com.hypermnesia.tray"; +#[cfg(not(target_os = "macos"))] +pub const DEFAULT_TRAY_LABEL: &str = "hypermnesia-tray"; + +/// What to call the file a job is defined in, in a sentence a person reads. A plist on macOS, a +/// unit file on Linux -- getting this wrong is exactly the kind of small, confident lie the rest +/// of this console refuses to tell. +#[cfg(target_os = "macos")] +pub const UNIT_NOUN: &str = "plist"; +#[cfg(target_os = "linux")] +pub const UNIT_NOUN: &str = "unit"; +#[cfg(not(any(target_os = "macos", target_os = "linux")))] +pub const UNIT_NOUN: &str = "job file"; + +/// The service manager in charge, named the way a person would say it: "launchd would not load +/// it", not "the service manager would not load it". +#[cfg(target_os = "macos")] +pub const BACKEND_NAME: &str = "launchd"; +#[cfg(target_os = "linux")] +pub const BACKEND_NAME: &str = "systemd"; +#[cfg(not(any(target_os = "macos", target_os = "linux")))] +pub const BACKEND_NAME: &str = "the service manager"; + +/// Where job definitions live, for messages that have to name a place. `~`, not `$HOME` +/// expanded: this is prose for a person, not a path handed to a shell. +#[cfg(target_os = "macos")] +pub const UNITS_LOCATION: &str = "~/Library/LaunchAgents"; +#[cfg(target_os = "linux")] +pub const UNITS_LOCATION: &str = "~/.config/systemd/user"; +#[cfg(not(any(target_os = "macos", target_os = "linux")))] +pub const UNITS_LOCATION: &str = "wherever this platform's scheduled jobs live"; + +/// The label prefix to look for: environment, then the config file, then the default -- the same +/// order as every other setting here. +/// +/// The config file matters more than it looks. The tray is started by the platform's own service +/// manager, which gives it a minimal environment and none of a login shell's variables, so an +/// installation whose jobs are named differently could set HM_JOB_PREFIX in its shell forever and +/// the tray would still show an empty job list -- while every command run by hand showed the +/// right one. +pub fn job_prefix() -> String { + if let Some(v) = std::env::var("HM_JOB_PREFIX").ok().filter(|s| !s.is_empty()) { + return v; + } + // Filtered the same way as the environment: an empty HM_JOB_PREFIX in the config makes + // `starts_with` match every job the person owns, and this module edits and restarts what it + // lists. + if let Some(v) = crate::config().get("HM_JOB_PREFIX").filter(|s| !s.is_empty()) { + return v.clone(); + } + DEFAULT_JOB_PREFIX.to_string() +} + +/// One calendar slot, as launchd stores it -- and as a systemd `OnCalendar=` line is parsed into, +/// since it expresses the same wildcard-by-omission idea. +/// +/// Every field is optional AND an omitted field is a WILDCARD, which is the whole reason this is +/// a struct rather than an hour and a minute. `{"Minute": 15}` means every hour at :15, not +/// "daily 00:15"; `{"Day": 1, "Hour": 3}` means the first of each month, not "daily 03:00". +/// Reading the omissions as zeroes printed a plausible wrong time -- and, worse, a wrong period: +/// a healthy monthly job was measured against a day and a half and reported overdue for +/// twenty-nine days out of thirty. +#[derive(Debug, Clone, PartialEq, Default)] +pub struct Cal { + pub minute: Option, + pub hour: Option, + /// Day of the month, 1-31. + pub day: Option, + /// 0 = Sunday. + pub weekday: Option, + pub month: Option, +} + +impl Cal { + pub fn at(hour: u32, minute: u32, weekday: Option) -> Self { + Self { hour: Some(hour), minute: Some(minute), weekday, ..Self::default() } + } + + /// How long between two firings, taken from the COARSEST field that is pinned: everything + /// finer than it repeats inside that cycle, everything coarser is a wildcard. + pub fn period(&self) -> Duration { + let day = 86_400; + Duration::from_secs(if self.month.is_some() { + 365 * day + } else if self.day.is_some() { + 31 * day + } else if self.weekday.is_some() { + 7 * day + } else if self.hour.is_some() { + day + } else if self.minute.is_some() { + 3_600 + } else { + // Nothing pinned at all: fires every minute. + 60 + }) + } + + pub fn human(&self) -> String { + const DAYS: [&str; 7] = ["Sun", "Mon", "Tue", "Wed", "Thu", "Fri", "Sat"]; + const MONTHS: [&str; 12] = ["Jan", "Feb", "Mar", "Apr", "May", "Jun", + "Jul", "Aug", "Sep", "Oct", "Nov", "Dec"]; + // "xx" where the schedule would repeat: an hour with no minute fires every minute of it. + let time = match (self.hour, self.minute) { + (Some(h), Some(m)) => format!("{h:02}:{m:02}"), + (Some(h), None) => format!("{h:02}:xx"), + (None, Some(m)) => format!(":{m:02}"), + (None, None) => String::new(), + }; + let month = self.month + .map(|m| format!("{} ", MONTHS[((m.max(1) - 1) as usize) % 12])) + .unwrap_or_default(); + match (self.month, self.day, self.weekday, self.hour, self.minute) { + (_, Some(d), _, _, _) => format!("{} {month}day {d} {time}", + if self.month.is_some() { "yearly" } else { "monthly" }), + (_, None, Some(w), _, _) => format!("weekly {} {time}", DAYS[(w as usize) % 7]), + (_, None, None, Some(_), _) => format!("daily {time}"), + (_, None, None, None, Some(_)) => format!("hourly at {time}"), + _ => "every minute".into(), + } + } +} + +#[derive(Debug, Clone, PartialEq)] +pub enum Schedule { + /// Every N seconds. + Every(u64), + /// One calendar slot. + Calendar(Cal), + /// Several calendar slots: launchd accepts an array of dicts, and systemd accepts several + /// `OnCalendar=` lines in one unit. All of them are kept -- the period has to come from the + /// widest of them, and the first slot is not the schedule. + Several(Vec), + /// No schedule: a service that is simply kept running. + None, +} + +impl Schedule { + /// A one-slot calendar schedule, which is what the console can set. + pub fn at(hour: u32, minute: u32, weekday: Option) -> Self { + Schedule::Calendar(Cal::at(hour, minute, weekday)) + } + + pub fn human(&self) -> String { + match self { + Schedule::Every(s) if *s % 3600 == 0 => format!("every {} h", s / 3600), + Schedule::Every(s) if *s % 60 == 0 => format!("every {} min", s / 60), + Schedule::Every(s) => format!("every {s} s"), + Schedule::Calendar(c) => c.human(), + Schedule::Several(c) if c.len() == 1 => c[0].human(), + Schedule::Several(c) => format!("{} slots ({})", c.len(), c[0].human()), + Schedule::None => "on demand".into(), + } + } +} + +/// What is known about a job's log. +#[derive(Debug, Clone, PartialEq)] +pub enum LogState { + /// The unit names no log: nothing can be dated by it, in either direction. + NotConfigured, + /// A log is configured and it is not there -- the job has not written anything since the + /// path was set. + Missing, + /// The log was last written at this time. + Written(SystemTime), + /// A log is configured, exists, and cannot be dated: unreadable, or stamped in the future + /// (a clock that moved). Not evidence of health and not evidence of failure -- and saying so + /// beats both of the alternatives. + Unknown(String), +} + +impl LogState { + pub fn at(&self) -> Option { + match self { + LogState::Written(t) => Some(*t), + _ => None, + } + } +} + +/// What the backend's own record says about when the job last fired, where it keeps one at all. +/// +/// A plain `Option` cannot say this: it collapses "this backend keeps no such record" +/// (launchd -- the log's mtime is the only evidence there is) and "this backend keeps one, and it +/// says never" (systemd, a timer whose `LastTriggerUSec` is empty) into the same `None`, and a +/// job read through the second case would then fall through to launchd's log/run-count logic, +/// find nothing there either, and read as calm by accident -- the same quiet lie `LogState` +/// exists to prevent for the log itself. +#[derive(Debug, Clone, PartialEq)] +pub enum TriggerState { + /// This backend keeps no direct record of when a job last fired. + NotTracked, + /// The backend tracks it, and it has never fired. + Never, + /// The backend tracks it, and it last fired at this time. + At(SystemTime), +} + +#[derive(Debug, Clone)] +pub struct Job { + pub label: String, + /// The file the job is defined in: a `.plist` on macOS, a `.service`/`.timer` on Linux. Named + /// `source` rather than `plist` so that a systemd path in this field is not a small lie about + /// which platform produced it. + pub source: PathBuf, + pub schedule: Schedule, + pub program: Vec, + pub log: Option, + /// Exit code of the last run, as the service manager remembers it. `None` -- the job is not + /// loaded. + pub last_exit: Option, + /// When the unit last changed. Needed so as not to raise a false alarm: a job installed + /// later than its own last slot simply has not come due yet. + pub installed: Option, + /// How many times the service manager has started it since this login/boot. Zero -- not once. + /// + /// Without this number the status column lies: launchd prints `0` in the status column both + /// for "finished successfully" and for "never finished at all", so a job that has never + /// started looks like one that ran without errors. `None` when the backend keeps no such + /// counter at all (systemd does not). + pub runs: Option, + /// PID, if the job is running right now. + pub pid: Option, + /// What the log says about when the job last worked. The only sign that survives a reboot on + /// launchd, and the reason it is three states rather than a timestamp: "no log configured", + /// "a log is configured and is not there" and "unreadable / dated in the future" are three + /// different pieces of knowledge, and collapsing them into None made a rotated-away log read + /// as a healthy job forever. + pub log_state: LogState, + /// When the backend's own record says the job last fired, where it keeps one at all. + /// systemd's `LastTriggerUSec` answers this directly and survives a reboot; launchd keeps + /// nothing of the kind, so this is always `NotTracked` there and the log's mtime is what + /// dates a run instead. Preferred over `log_state` wherever both could answer the same + /// question, because it is the more direct evidence. + pub trigger: TriggerState, + /// Why this job could not be read. A unit that the parser or the service manager refuses is + /// one that is not running, and that is precisely the case the console used to have nothing + /// to say about, because the row was dropped from the list. + pub fault: Option, +} + +impl Job { + /// Short name without the reverse-DNS prefix or the `hypermnesia-` prefix: on the console's + /// screen `extract`, not `com.hypermnesia.extract` or `hypermnesia-extract`. + pub fn short(&self) -> &str { + self.label.rsplit(['.', '-']).next().unwrap_or(&self.label) + } + + pub fn running(&self) -> bool { + self.pid.is_some() + } + + /// How long until the job repeats. For calendar schedules -- a day or a week; the exact date + /// of the next run is not needed, only the order of magnitude. + pub fn period(&self) -> Option { + match &self.schedule { + Schedule::Every(s) => Some(Duration::from_secs(*s)), + Schedule::Calendar(c) => Some(c.period()), + // The widest of the slots, not the narrowest: two weekly slots are still a week + // apart at worst, and measuring them against a day would cry wolf for six days of it. + Schedule::Several(cals) => cals.iter().map(Cal::period).max(), + Schedule::None => None, + } + } + + /// Something is wrong with the job itself, before any question of when it last ran. + pub fn broken(&self) -> bool { + self.fault.is_some() + } + + /// The service manager has never started it. + /// + /// `runs` alone does not answer this on launchd: it is per-bootstrap state, reset to zero at + /// every login while the plist's mtime does not move. Trusting it turned every healthy job + /// into "never ran / overdue" after a reboot. `trigger`, where the backend tracks one, is + /// authoritative and settles the question directly in either direction. Otherwise the log -- + /// the only evidence that survives a restart on launchd -- has to agree, and it only counts + /// if it was written AFTER the job was installed: a log left behind by a previous + /// installation says nothing about this one. + pub fn never_ran(&self) -> bool { + if self.schedule == Schedule::None || self.broken() { + return false; // a service without a schedule -- not a complaint + } + match self.trigger { + TriggerState::At(_) => return false, // the backend's own record says it fired + TriggerState::Never => return true, // ...and here, that it never has + TriggerState::NotTracked => {} + } + if self.log_after_install() { + return false; // it has written something since: it ran + } + match self.runs { + Some(n) => n == 0, + // The counter is unavailable: fall back to the sign "a log is configured but not + // there" -- weaker, but better than silence. + None => self.log_state == LogState::Missing, + } + } + + /// Was the log written after this version of the job was installed? A log from a previous + /// install is not evidence about this one -- in either direction. + fn log_after_install(&self) -> bool { + match (self.log_state.at(), self.installed) { + (Some(out), Some(inst)) => out >= inst, + (Some(_), None) => true, + _ => false, + } + } + + /// Has never run AND would have had time to -- or ran once and then stopped. THAT is the + /// complaint; `never_ran` on its own is NOT. + /// + /// The difference is not cosmetic. A weekly job installed on Monday afternoon honestly shows + /// zero runs until the next Monday, and shouting at it would be a lie. The signal must fire + /// on "a whole period has passed and nothing happened", otherwise people learn to scroll + /// past it and it stops working on the day it is needed. + /// + /// A period and a half is the threshold, not one period: a job whose slot is due about now + /// has not failed yet. + pub fn overdue(&self) -> bool { + let Some(period) = self.period() else { return false }; + let slack = period + period / 2; + if self.never_ran() { + // The only date available is the unit's own. + return self.since_installed().map(|age| age > slack).unwrap_or(true); + } + // It ran at some point. Then the complaint is staleness, dated by the best evidence + // available -- and never from before the job was installed, or a fresh install inheriting + // an old log would be called overdue the moment it was made. + match (self.since_last_output(), self.since_installed()) { + (Some(age), Some(since_install)) => age.min(since_install) > slack, + (Some(age), None) => age > slack, + (None, _) => false, + } + } + + /// Can this job's freshness be judged at all? A configured log that is gone or undatable + /// leaves the question open -- and an open question must not be drawn as health. + /// + /// `trigger`, where the backend tracks one, answers the question directly either way -- there + /// is nothing left unknown, whether it says "fired" or "never has". + pub fn freshness_unknown(&self) -> Option { + if self.broken() || self.schedule == Schedule::None + || matches!(self.trigger, TriggerState::At(_) | TriggerState::Never) { + return None; + } + match &self.log_state { + LogState::Unknown(why) => Some(why.clone()), + // Ran at least once since this boot, yet the log it is supposed to write is not + // there: it was rotated away, or the job is writing nowhere. + LogState::Missing if self.runs.unwrap_or(0) > 0 => + Some("the configured log is not there, so nothing can date its last run".into()), + LogState::NotConfigured if self.runs.unwrap_or(0) > 0 => + Some("no log is configured, so nothing survives a reboot to date its runs".into()), + _ => None, + } + } + + /// How long ago it was installed (by the unit's modification time). + pub fn since_installed(&self) -> Option { + self.installed.and_then(|t| SystemTime::now().duration_since(t).ok()) + } + + /// How long since the job last did something, from the best evidence available: the backend's + /// own trigger record if it keeps one, the log's mtime otherwise. + pub fn since_last_output(&self) -> Option { + let at = match self.trigger { + TriggerState::At(t) => Some(t), + TriggerState::Never | TriggerState::NotTracked => self.log_state.at(), + }?; + SystemTime::now().duration_since(at).ok() + } +} + +/// What the log can tell us, kept as the three different things it can be. Shared by both +/// backends: a log path and a filesystem are the same kind of fact on either platform. +pub(crate) fn log_state_of(log: Option<&PathBuf>) -> LogState { + let Some(p) = log else { return LogState::NotConfigured }; + let md = match std::fs::metadata(p) { + Ok(m) => m, + Err(e) if e.kind() == std::io::ErrorKind::NotFound => return LogState::Missing, + Err(e) => return LogState::Unknown(format!("{}: {e}", p.display())), + }; + match md.modified() { + Ok(t) if t > SystemTime::now() + Duration::from_secs(60) => + LogState::Unknown(format!("{} is stamped in the future -- a clock moved", + p.display())), + Ok(t) => LogState::Written(t), + Err(e) => LogState::Unknown(format!("{}: {e}", p.display())), + } +} + +/// The last `n` lines of a job's log. +pub fn tail(job: &Job, n: usize) -> Result { + let Some(path) = job.log.as_ref() else { + return Err("the job has no log file configured".into()); + }; + let data = std::fs::read(path) + .map_err(|e| format!("{}: {e}", path.display()))?; + let text = String::from_utf8_lossy(&data); + let lines: Vec<&str> = text.lines().collect(); + Ok(lines[lines.len().saturating_sub(n)..].join("\n")) +} + +#[cfg(test)] +mod tests { + use super::*; + + fn job(schedule: Schedule) -> Job { + Job { + label: "x".into(), source: PathBuf::new(), schedule, program: vec![], log: None, + last_exit: None, pid: None, log_state: LogState::NotConfigured, runs: None, + installed: None, trigger: TriggerState::NotTracked, fault: None, + } + } + + /// A job with a log that was written `ago` ago, installed `installed_ago` ago. + fn with_log(schedule: Schedule, runs: Option, written: Option, + installed_ago: Duration) -> Job { + let mut j = job(schedule); + j.log = Some(PathBuf::from("/nope")); + j.runs = runs; + j.last_exit = Some(0); + j.installed = Some(SystemTime::now() - installed_ago); + j.log_state = match written { + Some(d) => LogState::Written(SystemTime::now() - d), + None => LogState::Missing, + }; + j + } + + /// A job whose backend dates runs directly (systemd's `LastTriggerUSec`), with no log at all. + fn with_trigger(schedule: Schedule, ago: Option) -> Job { + let mut j = job(schedule); + j.last_exit = Some(0); + j.installed = Some(SystemTime::now() - Duration::from_secs(30 * 86_400)); + j.trigger = match ago { + Some(d) => TriggerState::At(SystemTime::now() - d), + None => TriggerState::Never, + }; + j + } + + #[test] + fn schedules_render_in_words() { + assert_eq!(Schedule::Every(14400).human(), "every 4 h"); + assert_eq!(Schedule::Every(900).human(), "every 15 min"); + // Something not a whole number of minutes stays in seconds instead of being rounded into + // a lie. + assert_eq!(Schedule::Every(90).human(), "every 90 s"); + assert_eq!(Schedule::at(5, 30, None).human(), "daily 05:30"); + assert_eq!(Schedule::at(6, 10, Some(1)).human(), "weekly Mon 06:10"); + assert_eq!(Schedule::None.human(), "on demand"); + } + + /// A job the service manager has never started must be called exactly that, not + /// "successful". + #[test] + fn a_scheduled_job_that_never_ran_is_flagged() { + let mut j = with_log(Schedule::at(6, 10, Some(1)), Some(0), None, + Duration::from_secs(30 * 86_400)); + assert!(j.never_ran()); + + j.runs = Some(44); // started since the last login + assert!(!j.never_ran()); + + j.runs = Some(0); + j.schedule = Schedule::None; // a service without a schedule -- not a complaint + assert!(!j.never_ran()); + + // The counter is unavailable (the job is not loaded): fall back to "a log is configured + // but not there". + j.schedule = Schedule::Every(3600); + j.runs = None; + assert!(j.never_ran()); + j.log_state = LogState::Written(SystemTime::now()); + assert!(!j.never_ran()); + } + + /// The run counter is per-bootstrap on launchd: it resets at every login while the unit's + /// mtime stays where it was. Keying "never ran" on it alone turned every healthy job into a + /// false alarm after a reboot -- which is how a warning stops being read. + #[test] + fn a_reboot_does_not_turn_a_working_job_into_a_false_alarm() { + let j = with_log(Schedule::Every(14_400), Some(0), Some(Duration::from_secs(600)), + Duration::from_secs(30 * 86_400)); + assert!(!j.never_ran(), "the log says it ran ten minutes ago"); + assert!(!j.overdue()); + } + + /// The other direction of the same mistake: a job that ran once and then stopped used to be + /// exempt forever, because a positive counter ended the question. + #[test] + fn a_job_that_ran_once_and_then_stalled_is_still_overdue() { + let mut j = with_log(Schedule::Every(3600), Some(12), Some(Duration::from_secs(600)), + Duration::from_secs(30 * 86_400)); + assert!(!j.overdue(), "ten minutes into an hourly job is fine"); + + j.log_state = LogState::Written(SystemTime::now() - Duration::from_secs(6 * 3600)); + assert!(j.overdue(), "six hours into an hourly job is not"); + } + + /// A log left behind by a PREVIOUS installation is not evidence about this one. Without + /// that rule a job installed a minute ago, pointing at a month-old log, was called overdue + /// on the spot. + #[test] + fn an_inherited_log_does_not_condemn_a_fresh_install() { + let j = with_log(Schedule::at(6, 10, Some(1)), Some(0), + Some(Duration::from_secs(30 * 86_400)), // log: a month old + Duration::from_secs(60)); // installed a minute ago + assert!(!j.overdue(), "it has not had a chance to run yet"); + assert!(j.never_ran(), "and the old log is not evidence that it did"); + } + + /// Freshness that cannot be judged must say so. A rotated-away log used to leave a stalled + /// job reading "last run ok" with an empty timestamp, for ever. + #[test] + fn a_log_that_cannot_date_a_run_is_an_open_question_not_health() { + let j = with_log(Schedule::Every(3600), Some(9), None, Duration::from_secs(86_400)); + assert!(j.freshness_unknown().is_some(), "the log is configured and gone"); + assert!(!j.overdue(), "and that is not the same as proof it stopped"); + + let mut future = with_log(Schedule::Every(3600), Some(9), Some(Duration::from_secs(60)), + Duration::from_secs(86_400)); + future.log_state = LogState::Unknown("stamped in the future".into()); + assert!(future.freshness_unknown().is_some()); + + let healthy = with_log(Schedule::Every(3600), Some(9), Some(Duration::from_secs(60)), + Duration::from_secs(86_400)); + assert!(healthy.freshness_unknown().is_none()); + } + + #[test] + fn period_matches_the_schedule() { + assert_eq!(job(Schedule::Every(14400)).period(), Some(Duration::from_secs(14400))); + assert_eq!(job(Schedule::at(5, 30, None)).period(), Some(Duration::from_secs(86_400))); + assert_eq!(job(Schedule::at(6, 10, Some(1))).period(), + Some(Duration::from_secs(7 * 86_400))); + assert_eq!(job(Schedule::None).period(), None); + } + + #[test] + fn short_name_drops_the_reverse_dns_prefix() { + let mut j = job(Schedule::None); + j.label = "com.hypermnesia.consolidate".into(); + assert_eq!(j.short(), "consolidate"); + } + + #[test] + fn short_name_drops_the_hyphen_prefix() { + let mut j = job(Schedule::None); + j.label = "hypermnesia-consolidate".into(); + assert_eq!(j.short(), "consolidate"); + } + + /// A job whose backend answers "when did it last fire" directly is neither never-ran, nor + /// overdue, nor unknown -- the question is settled, and drawing it as open or as healthy by + /// accident would both be wrong. + #[test] + fn a_backend_supplied_trigger_answers_freshness_directly() { + let fresh = with_trigger(Schedule::Every(3600), Some(Duration::from_secs(300))); + assert!(!fresh.never_ran()); + assert!(!fresh.overdue()); + assert!(fresh.freshness_unknown().is_none()); + // Not an exact `assert_eq!`: two `SystemTime::now()` calls a few instructions apart are a + // handful of microseconds apart too, and the age has to allow for that. + let age = fresh.since_last_output().expect("a trigger time was set"); + assert!(age >= Duration::from_secs(300) && age < Duration::from_secs(301), + "expected ~300s, got {age:?}"); + + let stale = with_trigger(Schedule::Every(3600), Some(Duration::from_secs(6 * 3600))); + assert!(!stale.never_ran(), "it did fire once, just long ago"); + assert!(stale.overdue()); + assert!(stale.freshness_unknown().is_none(), "the timestamp answers the question"); + + // The backend tracks triggers AND says it has never fired -- `TriggerState::Never`, not + // an absence of evidence. Must be a settled "never ran", not a fall-through to the + // log/run-count logic that has nothing to say here (no log, no run counter) and would + // otherwise read this job as calm by accident. + let never = with_trigger(Schedule::Every(3600), None); + assert!(never.never_ran()); + assert!(never.freshness_unknown().is_none(), "never-fired is a settled answer too"); + } +} diff --git a/console/src/jobs/systemd.rs b/console/src/jobs/systemd.rs new file mode 100644 index 0000000..876ca29 --- /dev/null +++ b/console/src/jobs/systemd.rs @@ -0,0 +1,618 @@ +//! The systemd backend: user units under `~/.config/systemd/user/*.{service,timer}`, read through +//! `systemctl --user show`. +//! +//! **Read-only, deliberately.** `list()` is real; `run_now`, `set_schedule`, `install_self` and +//! `uninstall_self` all return "not implemented on Linux yet" rather than a silent no-op -- +//! writing a systemd timer safely (the launchd side of this file is the backup-edit-lint-reload +//! dance in `launchd.rs`, about half its length) is real work this prototype does not yet do. +//! `autostart_fault` returns `None`: an honest "nothing to report" for a check the tray does not +//! yet perform, not a permanent warning about a missing feature -- a warning nobody can act on is +//! a warning nobody reads. +//! +//! systemd answers two of the three questions this module needs directly, and better than +//! launchd does: `ExecMainStatus`, `MainPID` and `LastTriggerUSec` are exactly "exit code", "is it +//! running" and "when did it last fire", and the last of those survives a reboot -- launchd keeps +//! nothing of the kind, which is why the launchd backend has to date runs by a log file's mtime +//! instead. What systemd does NOT keep is a run counter (`runs` stays `None` here always): the +//! nearest thing, `NRestarts`, counts crash-restarts of a still-loaded service, not lifetime runs, +//! and showing it under that label would be a different number wearing this one's name. +//! +//! All of `show`'s timestamp properties are parsed here rather than shelled out to `date`, and +//! that is only safe because every `systemctl --user show` call in this file pins `TZ=UTC` first: +//! with the zone fixed, the civil-calendar arithmetic in `parse_systemd_timestamp` needs no +//! timezone table at all, which is what keeps this data layer at zero dependencies. + +use std::path::PathBuf; +use std::process::Command; +use std::time::{Duration, SystemTime}; + +use super::{log_state_of, Cal, Job, LogState, Schedule, TriggerState}; + +fn units_dir() -> Result { + let home = std::env::var("HOME") + .map_err(|_| "HOME is not set, so I cannot tell where your systemd user units are" + .to_string())?; + if home.is_empty() { + return Err("HOME is empty, so I cannot tell where your systemd user units are".into()); + } + Ok(PathBuf::from(home).join(".config/systemd/user")) +} + +/// Every job with the given prefix, in name order. +/// +/// A `.timer` is the unit of record: it is the schedule, and it is what the console lets a person +/// change (once that is implemented). A `.timer` with no matching `.service` is still listed -- +/// `systemctl show` on the service side simply answers `LoadState=not-found`, which becomes a +/// fault on the row, not a silent drop. +/// +/// Two failures are `Err`, not an empty `Ok(vec![])`, because an empty *reading* and an empty +/// *store* must never look alike: no session bus reachable (the tray started before one existed, +/// or is running somewhere the bus is not shared), and a units directory that exists but cannot +/// be read (permissions). +pub fn list(prefix: &str) -> Result, String> { + if let Err(why) = bus_reachable() { + return Err(why); + } + let dir = units_dir()?; + let mut out = Vec::new(); + let entries = match std::fs::read_dir(&dir) { + Ok(e) => e, + // A units directory that has never been created is an honest empty store, same as + // launchd's LaunchAgents directory would be. Anything else -- permissions, a file where a + // directory belongs -- is a fault: the console cannot tell "no jobs" from "could not look". + Err(e) if e.kind() == std::io::ErrorKind::NotFound => return Ok(out), + Err(e) => return Err(format!("cannot read {}: {e}", dir.display())), + }; + for e in entries.flatten() { + let p = e.path(); + if p.extension().and_then(|x| x.to_str()) != Some("timer") { + continue; + } + let stem = p.file_stem().and_then(|x| x.to_str()).unwrap_or("").to_string(); + if !stem.starts_with(prefix) { + continue; + } + out.push(read_unit(&stem, &p)); + } + out.sort_by(|a, b| a.label.cmp(&b.label)); + Ok(out) +} + +/// Is `systemctl --user` reachable at all? Distinguishes "no session bus" (a real fault) from +/// "no units defined yet" (an honest empty list), which an unconditional +/// `Ok(read_dir(...).unwrap_or_default())` would have collapsed into the same "every job +/// healthy" screen shown for a total blackout. +fn bus_reachable() -> Result<(), String> { + let o = Command::new("systemctl").args(["--user", "show", "-p", "Version", "--", "-.mount"]) + .output().map_err(|e| format!("could not run systemctl: {e}"))?; + if o.status.success() { + return Ok(()); + } + let err = String::from_utf8_lossy(&o.stderr).trim().to_string(); + Err(if err.is_empty() { + "systemctl --user did not answer".into() + } else { + format!("systemctl --user is not reachable: {err}") + }) +} + +/// Read one job from its `.timer` file's stem, filling in what `systemctl --user show` knows +/// about both the timer and the service it triggers. +fn read_unit(stem: &str, timer_path: &PathBuf) -> Job { + let timer_label = format!("{stem}.timer"); + let service_label = format!("{stem}.service"); + let timer_rows = match show(&timer_label) { + Ok(r) => r, + Err(why) => return broken_job(stem, timer_path, why), + }; + if let Some(err) = load_fault(&timer_rows) { + return broken_job(stem, timer_path, err); + } + let service_rows = show(&service_label).unwrap_or_default(); + + let schedule = parse_schedule(&timer_rows); + // Empty means "never fired" -- systemd's own answer, not an absence of one -- and must stay + // distinguishable from `NotTracked`, or a job read on this backend could fall through to + // launchd's log/run-count logic, find nothing there either (systemd jobs rarely set either), + // and read as calm by accident. + let trigger = match get(&timer_rows, "LastTriggerUSec").unwrap_or("") { + "" => TriggerState::Never, + raw => match parse_systemd_timestamp(raw) { + Some(t) => TriggerState::At(t), + // A non-empty LastTriggerUSec that fails to parse is worth a fault of its own: the + // job DID fire, by the backend's own account, and silently falling back to "never + // ran" would be exactly the quiet substitution this console exists to refuse. + None => return broken_job(stem, timer_path, + format!("the timer fired, but its own timestamp could not be read: {raw:?}")), + }, + }; + + let log = get(&service_rows, "StandardOutput") + .and_then(|v| v.strip_prefix("append:")) + .map(PathBuf::from) + .or_else(|| service_log_from_file(stem)); + let last_exit = get(&service_rows, "ExecMainStatus").and_then(|v| v.parse().ok()); + let pid = get(&service_rows, "MainPID").and_then(|v| v.parse().ok()).filter(|p| *p != 0); + let installed = get(&timer_rows, "FragmentPath") + .map(PathBuf::from) + .and_then(|p| std::fs::metadata(p).ok()) + .and_then(|m| m.modified().ok()); + let program = get(&service_rows, "ExecStart") + .and_then(|v| extract_between(v, "path")) + .map(|p| vec![p]) + .unwrap_or_default(); + + Job { + // The bare stem, not `timer_label`: `Job::short()` strips a trailing `.something` the + // same way it strips the reverse-DNS prefix, and a label ending in `.timer` fed that + // logic a name to cut at, dropping the job's actual name ("extract") and leaving "timer" + // behind instead -- on every single job, since every systemd label ends the same way. + label: stem.to_string(), + source: timer_path.clone(), + schedule, + program, + log_state: log_state_of(log.as_ref()), + log, + last_exit, + installed, + runs: None, // systemd keeps no lifetime run counter + pid, + trigger, + fault: enablement_fault(&timer_rows), + } +} + +/// Not scheduled by anything, despite having a valid unit: disabled, masked, or the timer is +/// loaded but not active. No launchd analogue -- launchd loads whatever is in LaunchAgents, but a +/// systemd unit can exist, parse cleanly, and still not be armed. Rendering that as a healthy +/// schedule would be the same quiet lie as an unreadable file read as an empty list. +fn enablement_fault(timer_rows: &[(String, String)]) -> Option { + let file_state = get(timer_rows, "UnitFileState").unwrap_or(""); + if matches!(file_state, "disabled" | "masked" | "masked-runtime") { + return Some(format!("the timer is {file_state}, so it is not scheduled by anything")); + } + let active = get(timer_rows, "ActiveState").unwrap_or(""); + if active != "active" { + let sub = get(timer_rows, "SubState").unwrap_or(""); + return Some(format!("the timer is loaded but not armed (ActiveState={active}, \ + SubState={sub})")); + } + None +} + +fn load_fault(rows: &[(String, String)]) -> Option { + match get(rows, "LoadState") { + Some("loaded") | None => None, + Some(other) => { + let detail = get(rows, "LoadError").filter(|s| !s.is_empty()); + Some(match detail { + Some(d) => format!("{other}: {d}"), + None => other.to_string(), + }) + } + } +} + +fn broken_job(stem: &str, path: &PathBuf, fault: String) -> Job { + Job { + label: stem.to_string(), // see the comment on the same field in read_unit + source: path.clone(), + schedule: Schedule::None, + program: Vec::new(), + log: None, + log_state: LogState::NotConfigured, + last_exit: None, + installed: std::fs::metadata(path).ok().and_then(|m| m.modified().ok()), + runs: None, + pid: None, + trigger: TriggerState::NotTracked, + fault: Some(fault), + } +} + +/// `systemctl --user show `, with the timezone pinned so every timestamp property comes +/// back in the one format `parse_systemd_timestamp` understands. +fn show(unit: &str) -> Result, String> { + let out = Command::new("systemctl") + .env("TZ", "UTC") + .args(["--user", "show", unit]) + .output() + .map_err(|e| format!("could not run systemctl: {e}"))?; + if !out.status.success() { + let err = String::from_utf8_lossy(&out.stderr).trim().to_string(); + return Err(if err.is_empty() { "systemctl show failed".into() } else { err }); + } + Ok(parse_show(&String::from_utf8_lossy(&out.stdout))) +} + +/// `KEY=VALUE`, one per line. A key can repeat -- `TimersCalendar` does, once per `OnCalendar=` +/// slot in the unit -- so every row is kept rather than folded into a map. +fn parse_show(text: &str) -> Vec<(String, String)> { + text.lines() + .filter_map(|line| line.split_once('=').map(|(k, v)| (k.to_string(), v.to_string()))) + .collect() +} + +fn get<'a>(rows: &'a [(String, String)], key: &str) -> Option<&'a str> { + rows.iter().find(|(k, _)| k == key).map(|(_, v)| v.as_str()) +} + +fn get_all<'a, 'k>(rows: &'a [(String, String)], key: &'k str) + -> impl Iterator + 'k where 'a: 'k { + rows.iter().filter(move |(k, _)| k == key).map(|(_, v)| v.as_str()) +} + +/// The schedule, from `TimersCalendar` (one row per `OnCalendar=` slot) or `TimersMonotonic` +/// (one `OnUnitActiveSec=`-style interval). Calendar wins if both are somehow present, since that +/// is the more specific thing to have said. +fn parse_schedule(timer_rows: &[(String, String)]) -> Schedule { + let cals: Vec = get_all(timer_rows, "TimersCalendar") + .filter_map(|raw| extract_between(raw, "OnCalendar")) + .filter_map(|spec| parse_oncalendar(&spec)) + .collect(); + match cals.len() { + 0 => {} + 1 => return Schedule::Calendar(cals.into_iter().next().unwrap()), + _ => return Schedule::Several(cals), + } + if let Some(raw) = get(timer_rows, "TimersMonotonic") { + if let Some(spec) = extract_between(raw, "OnUnitActiveUSec") { + if let Some(secs) = parse_systemd_duration(&spec) { + return Schedule::Every(secs); + } + } + } + Schedule::None +} + +/// Pulls `key=...` out of systemd's `{ key=value ; next_elapse=... }` compound property value. +fn extract_between(raw: &str, key: &str) -> Option { + let after = raw.split_once(&format!("{key}="))?.1; + let end = after.find(" ;").unwrap_or(after.len()); + let v = after[..end].trim(); + (!v.is_empty()).then(|| v.to_string()) +} + +/// An `OnCalendar=` spec, always in the normalized form systemd echoes back: +/// `[" "]"-- "::`. Every field systemd does not pin is a +/// literal `*` here too -- `*:15:00` is "every hour at :15", not "00:15" -- the same +/// omission-is-a-wildcard rule launchd's plists use, and the same bug (a plausible wrong time, +/// and a monthly job measured against a day and a half) if it is read as a zero instead. +fn parse_oncalendar(spec: &str) -> Option { + const DAYS: [&str; 7] = ["Sun", "Mon", "Tue", "Wed", "Thu", "Fri", "Sat"]; + let mut parts = spec.split_whitespace(); + let first = parts.next()?; + let (weekday, date_tok) = match DAYS.iter().position(|d| *d == first) { + Some(i) => (Some(i as u32), parts.next()?), + None => (None, first), + }; + let (month, day) = parse_calendar_field(date_tok); + let (hour, minute) = parts.next().map(parse_time_field).unwrap_or((None, None)); + Some(Cal { minute, hour, day, weekday, month }) +} + +/// `Y-M-D`, each component `*` or a number. The year is read and discarded: `Cal` has no field +/// for it, and every schedule this console can set or has ever been asked to render is yearly at +/// most. +fn parse_calendar_field(s: &str) -> (Option, Option) { + let mut it = s.split('-'); + let _year = it.next(); + let month = it.next().and_then(wildcard_num); + let day = it.next().and_then(wildcard_num); + (month, day) +} + +/// `H:M:S`, each component `*` or a number. Seconds are read and discarded: nothing in `Cal` +/// tracks them, and no schedule this console offers is finer than a minute. +fn parse_time_field(s: &str) -> (Option, Option) { + let mut it = s.split(':'); + let hour = it.next().and_then(wildcard_num); + let minute = it.next().and_then(wildcard_num); + (hour, minute) +} + +fn wildcard_num(tok: &str) -> Option { + if tok == "*" { None } else { tok.parse().ok() } +} + +/// A duration the way `systemctl show` writes one back: compound, largest unit first -- +/// `"1h 30min"`, `"1min 30s"`, `"4h"`. Not the `4h`/`15m` shorthand a person types (that is +/// `hypermnesia-jobs`' own `parse_interval`, on the way in); this is systemd's own way of writing +/// the same number back out, on the way out. +fn parse_systemd_duration(s: &str) -> Option { + const UNITS: &[(&str, u64)] = &[ + ("usec", 0), ("us", 0), ("msec", 0), ("ms", 0), + ("seconds", 1), ("second", 1), ("sec", 1), ("s", 1), + ("minutes", 60), ("minute", 60), ("min", 60), + ("hours", 3600), ("hour", 3600), ("h", 3600), + ("days", 86_400), ("day", 86_400), ("d", 86_400), + ("weeks", 604_800), ("week", 604_800), ("w", 604_800), + ("months", 2_592_000), ("month", 2_592_000), + ("years", 31_536_000), ("year", 31_536_000), ("y", 31_536_000), + ]; + let mut total = 0u64; + let mut saw_any = false; + for tok in s.split_whitespace() { + let digits_end = tok.find(|c: char| !c.is_ascii_digit() && c != '.')?; + let (num, unit) = tok.split_at(digits_end); + // Fractional seconds below our resolution: truncated, not rounded up into a lie about + // precision this console never had. + let n: f64 = num.parse().ok()?; + let (_, mult) = UNITS.iter().find(|(u, _)| *u == unit)?; + total += (n * (*mult as f64)) as u64; + saw_any = true; + } + saw_any.then_some(total) +} + +/// `"Tue 2026-09-22 16:45:31 UTC"` -> a `SystemTime`. Safe to hand-roll only because every caller +/// in this file forces `TZ=UTC` on the `systemctl` invocation first: with the zone pinned, this +/// needs plain Gregorian civil-calendar arithmetic and no timezone database, which is what keeps +/// this data layer dependency-free. `None` on anything unexpected -- a name-based zone abbreviation +/// slipping through despite the pin, for instance -- and the caller turns that into a fault rather +/// than a silent "never ran". +fn parse_systemd_timestamp(s: &str) -> Option { + let mut parts = s.split_whitespace(); + let _weekday = parts.next()?; + let date = parts.next()?; + let time = parts.next()?; + let zone = parts.next()?; + if zone != "UTC" { + return None; + } + let mut d = date.split('-'); + let y: i64 = d.next()?.parse().ok()?; + let mo: u32 = d.next()?.parse().ok()?; + let da: u32 = d.next()?.parse().ok()?; + let mut t = time.split(':'); + let h: u64 = t.next()?.parse().ok()?; + let mi: u64 = t.next()?.parse().ok()?; + // Seconds may carry a fractional part; truncate it the same way parse_systemd_duration does. + let se: u64 = t.next()?.split('.').next()?.parse().ok()?; + if mo == 0 || mo > 12 || da == 0 || da > 31 || h > 23 || mi > 59 || se > 60 { + return None; + } + let days = days_from_civil(y, mo, da); + let secs = days.checked_mul(86_400)?.checked_add((h * 3600 + mi * 60 + se) as i64)?; + if secs < 0 { + return None; // this console has no use for a date before 1970 + } + Some(SystemTime::UNIX_EPOCH + Duration::from_secs(secs as u64)) +} + +/// Days since 1970-01-01, for a proleptic Gregorian civil date. Howard Hinnant's +/// `days_from_civil` algorithm (public domain) -- chosen over pulling in a time crate for the one +/// calculation this file needs, in keeping with the data layer's zero dependencies. +fn days_from_civil(y: i64, m: u32, d: u32) -> i64 { + let y = if m <= 2 { y - 1 } else { y }; + let era = if y >= 0 { y } else { y - 399 } / 400; + let yoe = y - era * 400; // [0, 399] + let mp = (m as i64 + 9) % 12; // [0, 11], Mar-based + let doy = (153 * mp + 2) / 5 + d as i64 - 1; // [0, 365] + let doe = yoe * 365 + yoe / 4 - yoe / 100 + doy; // [0, 146096] + era * 146_097 + doe - 719_468 +} + +fn service_log_from_file(stem: &str) -> Option { + let path = units_dir().ok()?.join(format!("{stem}.service")); + let text = std::fs::read_to_string(path).ok()?; + parse_standard_output(&text) +} + +/// `StandardOutput=append:/path` as written in the unit file, for the (probably rare) case where +/// `systemctl show` cannot be reached but the file can -- mirrors `show`'s own stripped-prefix +/// reading of the property, so both paths agree. +fn parse_standard_output(unit_file: &str) -> Option { + unit_file.lines() + .find_map(|l| l.trim().strip_prefix("StandardOutput=")?.strip_prefix("append:")) + .map(PathBuf::from) +} + +pub fn run_now(_label: &str) -> Result { + Err("running a job on demand is not implemented on Linux yet -- use `systemctl --user start \ + .service` directly".into()) +} + +pub fn set_schedule(_job: &Job, _sched: &Schedule) -> Result { + Err("changing a schedule is not implemented on Linux yet -- edit the `.timer` unit and run \ + `systemctl --user daemon-reload`".into()) +} + +pub fn install_self(_label: &str) -> Result { + Err("installing the tray into autostart is not implemented on Linux yet -- add a \ + `hypermnesia-tray.service` user unit with `WantedBy=default.target` by hand".into()) +} + +pub fn uninstall_self(_label: &str) -> Result { + Err("removing the tray from autostart is not implemented on Linux yet".into()) +} + +/// Not implemented, so there is nothing to be wrong about -- `None`, not a permanent warning. +/// `bin/tray.rs` draws whatever this returns on every single menu rebuild, and the project's own +/// rule is that a warning shown forever is a warning nobody reads. +pub fn autostart_fault(_label: &str) -> Option { + None +} + +#[cfg(test)] +mod tests { + use super::*; + + // Real output of `TZ=UTC systemctl --user show .timer`, shortened to the properties + // this module reads. Captured on systemd 259 (Fedora 44). + const WEEKLY: &[(&str, &str)] = &[ + ("Id", "hypermnesia-reflect.timer"), + ("LoadState", "loaded"), + ("ActiveState", "active"), + ("SubState", "waiting"), + ("UnitFileState", "enabled"), + ("FragmentPath", "/home/you/.config/systemd/user/hypermnesia-reflect.timer"), + ("TimersCalendar", "{ OnCalendar=Mon *-*-* 06:10:00 ; next_elapse=Mon 2026-09-28 06:10:00 UTC }"), + ("LastTriggerUSec", "Tue 2026-09-22 16:45:31 UTC"), + ]; + + fn rows(pairs: &[(&str, &str)]) -> Vec<(String, String)> { + pairs.iter().map(|(k, v)| (k.to_string(), v.to_string())).collect() + } + + #[test] + fn reads_a_weekly_calendar_schedule() { + let r = rows(WEEKLY); + assert_eq!(parse_schedule(&r), Schedule::at(6, 10, Some(1))); + } + + /// An omitted field in an `OnCalendar=` spec is a WILDCARD, exactly as in a launchd plist: + /// `*:15:00` means every hour at :15, not 00:15. Reading it as a zero gives a plausible wrong + /// time and, far worse, a wrong period -- a monthly job measured against a day and a half + /// reads as overdue for twenty-nine days out of thirty. + #[test] + fn an_omitted_calendar_field_is_a_wildcard_not_a_zero() { + let hourly = rows(&[("TimersCalendar", + "{ OnCalendar=*-*-* *:15:00 ; next_elapse=(null) }")]); + let s = parse_schedule(&hourly); + assert_eq!(s, Schedule::Calendar(Cal { minute: Some(15), ..Cal::default() })); + assert_eq!(s.human(), "hourly at :15"); + + let monthly = rows(&[("TimersCalendar", + "{ OnCalendar=*-*-01 03:00:00 ; next_elapse=(null) }")]); + let s = parse_schedule(&monthly); + assert_eq!(s.human(), "monthly day 1 03:00"); + + let every_minute = rows(&[("TimersCalendar", + "{ OnCalendar=*-*-* *:*:00 ; next_elapse=(null) }")]); + assert_eq!(parse_schedule(&every_minute).human(), "every minute"); + } + + /// Several `OnCalendar=` lines are several rows of the same key -- the period must come from + /// the WIDEST slot, exactly as for launchd's array form, or two weekly slots would be + /// measured against a day. + #[test] + fn several_calendar_lines_keep_their_widest_period() { + let r = rows(&[ + ("TimersCalendar", "{ OnCalendar=*-*-* 06:10:00 ; next_elapse=(null) }"), + ("TimersCalendar", "{ OnCalendar=*-*-* 18:10:00 ; next_elapse=(null) }"), + ]); + let s = parse_schedule(&r); + assert!(matches!(&s, Schedule::Several(c) if c.len() == 2)); + } + + #[test] + fn reads_a_monotonic_schedule() { + let r = rows(&[("TimersMonotonic", "{ OnUnitActiveUSec=4h ; next_elapse=0 }")]); + assert_eq!(parse_schedule(&r), Schedule::Every(14_400)); + + let compound = rows(&[("TimersMonotonic", "{ OnUnitActiveUSec=1h 30min ; next_elapse=0 }")]); + assert_eq!(parse_schedule(&compound), Schedule::Every(5_400)); + + let seconds = rows(&[("TimersMonotonic", "{ OnUnitActiveUSec=1min 30s ; next_elapse=0 }")]); + assert_eq!(parse_schedule(&seconds), Schedule::Every(90)); + } + + #[test] + fn a_timestamp_round_trips_through_civil_arithmetic() { + // Checked against `date -u -d "2026-09-22 16:45:31 UTC" +%s` => 1790095531. + let t = parse_systemd_timestamp("Tue 2026-09-22 16:45:31 UTC").expect("parses"); + let secs = t.duration_since(SystemTime::UNIX_EPOCH).unwrap().as_secs(); + assert_eq!(secs, 1_790_095_531); + } + + /// A second date, in a different month and a leap year, to catch a days-in-month or + /// era-boundary mistake the first sample would not exercise. + #[test] + fn a_second_timestamp_in_a_leap_year_also_round_trips() { + // `date -u -d "2028-02-29 00:00:00 UTC" +%s`. + let t = parse_systemd_timestamp("Tue 2028-02-29 00:00:00 UTC").expect("parses"); + let secs = t.duration_since(SystemTime::UNIX_EPOCH).unwrap().as_secs(); + assert_eq!(secs, 1_835_395_200); + } + + #[test] + fn a_non_utc_timestamp_is_refused_rather_than_misread() { + // Every caller pins TZ=UTC before asking systemctl for one of these; if that guard were + // ever dropped, a name-based zone abbreviation must not be silently read as UTC. + assert!(parse_systemd_timestamp("Tue 2026-09-22 18:15:31 IST").is_none()); + assert!(parse_systemd_timestamp("Tue 2026-09-22 12:45:31 +04").is_none()); + } + + #[test] + fn an_empty_last_trigger_means_never_fired_not_unparseable() { + // list()'s own handling of "" lives in read_unit, not here -- but the string itself must + // not be handed to the parser, which would (correctly) refuse it. + assert!(parse_systemd_timestamp("").is_none()); + } + + #[test] + fn a_disabled_timer_is_a_fault_not_a_healthy_schedule() { + let r = rows(&[("UnitFileState", "disabled"), ("ActiveState", "inactive")]); + assert!(enablement_fault(&r).is_some()); + } + + #[test] + fn a_loaded_but_inactive_timer_is_a_fault() { + let r = rows(&[("UnitFileState", "enabled"), ("ActiveState", "inactive"), + ("SubState", "dead")]); + assert!(enablement_fault(&r).is_some()); + } + + #[test] + fn an_active_enabled_timer_has_no_enablement_fault() { + let r = rows(&[("UnitFileState", "enabled"), ("ActiveState", "active"), + ("SubState", "waiting")]); + assert!(enablement_fault(&r).is_none()); + } + + #[test] + fn a_bad_unit_file_is_a_load_fault_with_the_reason_attached() { + let r = rows(&[("LoadState", "bad-setting"), + ("LoadError", "org.freedesktop.systemd1.BadUnitSetting \"bad setting\"")]); + let f = load_fault(&r).expect("fault"); + assert!(f.contains("bad setting"), "{f}"); + } + + #[test] + fn a_not_found_unit_is_a_load_fault() { + let r = rows(&[("LoadState", "not-found")]); + assert!(load_fault(&r).is_some()); + } + + #[test] + fn a_loaded_unit_has_no_load_fault() { + assert!(load_fault(&rows(&[("LoadState", "loaded")])).is_none()); + } + + #[test] + fn the_exec_start_path_is_pulled_out_of_the_compound_property() { + let raw = "{ path=/usr/bin/python3 ; argv[]=/usr/bin/python3 /x/mem_extract.py ; \ + ignore_errors=no }"; + assert_eq!(extract_between(raw, "path").as_deref(), Some("/usr/bin/python3")); + } + + #[test] + fn standard_output_append_path_is_read_from_the_unit_file() { + let unit = "[Service]\nType=oneshot\nStandardOutput=append:/home/you/.hypermnesia/logs/x.log\n"; + assert_eq!(parse_standard_output(unit), + Some(PathBuf::from("/home/you/.hypermnesia/logs/x.log"))); + assert_eq!(parse_standard_output("[Service]\nType=oneshot\n"), None); + } + + #[test] + fn write_actions_refuse_rather_than_pretend() { + assert!(run_now("hypermnesia-extract").is_err()); + let job = Job { + label: "x".into(), source: PathBuf::new(), schedule: Schedule::None, + program: vec![], log: None, log_state: LogState::NotConfigured, last_exit: None, + installed: None, runs: None, pid: None, trigger: TriggerState::NotTracked, fault: None, + }; + assert!(set_schedule(&job, &Schedule::Every(3600)).is_err()); + assert!(install_self("hypermnesia-tray").is_err()); + assert!(uninstall_self("hypermnesia-tray").is_err()); + } + + #[test] + fn autostart_fault_is_none_not_a_permanent_warning() { + // Not implemented is not the same claim as "checked, and it's fine" -- but a permanent + // warning drawn on every menu rebuild is a warning nobody reads, so this returns the + // same `None` a passing check would. + assert_eq!(autostart_fault("hypermnesia-tray"), None); + } +} diff --git a/console/src/jobs/unsupported.rs b/console/src/jobs/unsupported.rs new file mode 100644 index 0000000..10beb25 --- /dev/null +++ b/console/src/jobs/unsupported.rs @@ -0,0 +1,31 @@ +//! Neither launchd nor systemd: a platform this console has no scheduled-jobs backend for. +//! +//! `list()` errors rather than returning an empty list -- "not supported here" and "no jobs +//! configured" are different facts, and the console must not print the second when it means the +//! first. + +use super::{Job, Schedule}; + +pub fn list(_prefix: &str) -> Result, String> { + Err("scheduled jobs are not supported on this platform".into()) +} + +pub fn run_now(_label: &str) -> Result { + Err("scheduled jobs are not supported on this platform".into()) +} + +pub fn set_schedule(_job: &Job, _sched: &Schedule) -> Result { + Err("scheduled jobs are not supported on this platform".into()) +} + +pub fn install_self(_label: &str) -> Result { + Err("autostart is not supported on this platform".into()) +} + +pub fn uninstall_self(_label: &str) -> Result { + Err("autostart is not supported on this platform".into()) +} + +pub fn autostart_fault(_label: &str) -> Option { + None +} diff --git a/console/src/lib.rs b/console/src/lib.rs index d55e7a9..cab8052 100644 --- a/console/src/lib.rs +++ b/console/src/lib.rs @@ -15,6 +15,7 @@ use std::time::{Duration, Instant}; pub mod jobs; pub mod settings; +pub mod view; pub const STATS_SQL: &str = include_str!("stats.sql"); diff --git a/console/src/view.rs b/console/src/view.rs new file mode 100644 index 0000000..d513d3d --- /dev/null +++ b/console/src/view.rs @@ -0,0 +1,159 @@ +//! How the tray renders what it knows, shared by every platform's menu shell. +//! +//! Nothing here calls a GUI library or a platform's service manager. It turns a `Stats` and a +//! `[Job]` into the same lines and the same single yes/no "does this need a mark" verdict on +//! macOS and on Linux, so the two menu shells cannot drift into disagreeing about what a person +//! should be shown. `warned()` is deliberately *derived* from the lines this module renders, +//! never recomputed beside them: a title mark that disagrees with the menu under it is the same +//! quiet lie as a stale reading drawn as a fresh one. + +use std::time::{Duration, SystemTime}; + +use crate::jobs::{self, Job}; +use crate::Stats; + +/// What is visible without opening the menu. What is visible without opening the menu. A mark +/// matters more than a number here: the title/icon is where a problem is NOTICED, not where a +/// report is read. +/// +/// Derived from the same lines `summary` and `job_line` produce, not from a second list of +/// conditions kept in step by hand -- recomputing separately is what once let a warning line sit +/// in the menu under a calm mark. +pub fn warned(stats: &Stats, jobs: &[Job], jobs_error: &Option) -> bool { + summary(stats).iter().any(|l| l.starts_with('!')) + || jobs_error.is_some() + || jobs.iter().any(Job::broken) + || jobs.iter().any(Job::overdue) +} + +/// The lines describing the store's numbers, for the menu body. +pub fn summary(s: &Stats) -> Vec { + let (docs, chunks, embedded) = s.corpus.iter() + .filter(|r| !r.is_memory_page()) + .fold((0, 0, 0), |(d, c, e), r| (d + r.docs, c + r.chunks, e + r.embedded)); + let mut out = vec![ + format!("Memory: {} active of {}", s.memories_active, s.memories_total), + format!("Knowledge pages: {}", s.pages), + format!("Corpus: {docs} docs, {chunks} chunks"), + ]; + if embedded < chunks { + out.push(format!("! not embedded: {}", chunks - embedded)); + } + if s.embedding_models.len() > 1 { + // Two models in one store means part of the corpus cannot be reached by meaning at all. + out.push(format!("! embedding models: {}", s.embedding_models.len())); + } + // Marked when there is something in it, so the title can be derived from these lines rather + // than from a second list of conditions kept in step by hand. + out.push(if s.review_pending > 0 { + format!("! Review queue: {} (oldest {} days)", s.review_pending, s.review_oldest_days) + } else { + "Review queue: 0".to_string() + }); + out.push(format!("Stale: {}", s.stale)); + out.push(format!("Database: {}", s.db_size)); + out +} + +/// How old something is, by the wall clock. `Instant` would stop while the machine sleeps. +pub fn age_of(t: SystemTime) -> Duration { + SystemTime::now().duration_since(t).unwrap_or_default() +} + +pub fn ago(d: Duration) -> String { + let s = d.as_secs(); + if s < 90 { format!("{s}s") } + else if s < 5400 { format!("{}m", s / 60) } + else if s < 172_800 { format!("{}h", s / 3600) } + else { format!("{}d", s / 86_400) } +} + +/// One job's line in the "Run now" / status listing. +pub fn job_line(j: &Job) -> String { + if let Some(why) = &j.fault { + // The backend refused this unit too, so the job is not running. It used to be missing + // from the menu entirely. + let first = why.lines().next().unwrap_or(why); + return format!("{} (! unreadable {}: {first})", j.short(), jobs::UNIT_NOUN); + } + let state = if j.running() { + "running".to_string() + } else if j.overdue() { + "! overdue".to_string() + } else if j.never_ran() { + "never run yet".to_string() + } else if let Some(_) = j.freshness_unknown() { + // Neither fresh nor stale: the log that would date it is gone or undatable. Drawing + // that as health is how a job that quietly stopped stays invisible. + "? cannot be dated".to_string() + } else if j.last_exit.is_none() { + "not loaded".to_string() + } else if j.runs == Some(0) { + // The service manager prints exit 0 both for "finished successfully" and for "never + // finished at all". With no run since this login there is nothing to call successful. + "no run since login".to_string() + } else { + // The exit code belongs here. Without it a job that fails on every single run reads + // exactly like one that works: same schedule, same fresh log timestamp. + let when = match j.since_last_output() { + Some(d) => format!("{} ago", ago(d)), + None => "—".to_string(), + }; + match j.last_exit { + Some(0) => when, + Some(c) => format!("! exit {c}, {when}"), + None => format!("{when}, not loaded"), + } + }; + format!("{} ({}, {})", j.short(), j.schedule.human(), state) +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::jobs::{LogState, Schedule}; + use std::path::PathBuf; + + fn job(schedule: Schedule) -> Job { + Job { + label: "com.hypermnesia.extract".into(), source: PathBuf::new(), schedule, + program: vec![], log: Some(PathBuf::from("/nope")), last_exit: Some(0), pid: None, + log_state: LogState::Written(SystemTime::now() - Duration::from_secs(300)), + runs: Some(9), trigger: jobs::TriggerState::NotTracked, + installed: Some(SystemTime::now() - Duration::from_secs(90_000)), fault: None, + } + } + + /// A job that fails on every run used to read exactly like one that works: same schedule, + /// same fresh log timestamp, no code anywhere in the menu. + #[test] + fn a_failing_job_says_so_in_its_own_line() { + let mut j = job(Schedule::Every(14_400)); + assert_eq!(job_line(&j), "extract (every 4 h, 5m ago)"); + + j.last_exit = Some(2); + let line = job_line(&j); + assert!(line.contains("exit 2"), "the exit code has to be on the line: {line}"); + assert!(line.starts_with("extract (every 4 h, !"), "and marked: {line}"); + } + + /// A unit the backend also refused is one that is not running. It used to be absent from the + /// menu altogether. + #[test] + fn an_unreadable_unit_is_a_visible_line() { + let mut j = job(Schedule::None); + j.fault = Some("plutil: unexpected character\nsecond line".into()); + let line = job_line(&j); + assert!(line.contains(&format!("unreadable {}", jobs::UNIT_NOUN)), "{line}"); + assert!(!line.contains("second line"), "one line only, it is a menu: {line}"); + } + + #[test] + fn a_job_never_run_is_not_called_fresh() { + let mut j = job(Schedule::Calendar(jobs::Cal { + hour: Some(6), minute: Some(10), day: None, weekday: Some(1), month: None })); + j.runs = Some(0); + j.log_state = LogState::Missing; + assert!(job_line(&j).contains("never run yet"), "{}", job_line(&j)); + } +} diff --git a/docs/DIAGNOSTICS.md b/docs/DIAGNOSTICS.md index 55eb4ce..77f12bf 100644 --- a/docs/DIAGNOSTICS.md +++ b/docs/DIAGNOSTICS.md @@ -114,24 +114,32 @@ have. What it says, and what each line means: | `! not updated: , showing state from 4m ago` | the reading failed; these numbers are old and this is how old | | `! no answer about a reading for 2m -- the reader is stuck` | the worker thread is alive and has stopped answering | | `! the reader thread has died` | nothing will update again; restart the tray | -| `never ran (installed 3 d ago)` | launchd has started this job zero times and it has written no log | -| `! the period has already passed and launchd still never ran it` | a whole period and a half with no run — the complaint, as opposed to the observation above | +| `never ran (installed 3 d ago)` | the service manager has started this job zero times and it has written no log | +| `! the period has already passed and it still never ran` | a whole period and a half with no run — the complaint, as opposed to the observation above | | `! it ran at some point, but the log has not moved for more than a period` | it worked once and stopped | -| `! unreadable plist: ` | launchd refused this job too; it is not running | +| `! unreadable plist: ` (macOS) | launchd refused this job too; it is not running | +| `! unreadable unit: ` (Linux) | the same fault, from `systemctl --user show` or the parser instead of `plutil` | +| `! the timer is disabled, so it is not scheduled by anything` (Linux) | the `.timer` unit exists and parses, but is not enabled — no launchd analogue: a plist in LaunchAgents is loaded by definition, but a systemd timer can exist and simply not be armed | +| `! the timer is loaded but not armed (ActiveState=..., SubState=...)` (Linux) | loaded and enabled, but not actually active right now | | `! exit 2, 5m ago` | the last run failed. `freshness` exits 1 by design, meaning discrepancies were found | -| `loaded; no run since this login` | launchd has the job and has not started it since the last login. Its exit column says 0, which it prints for "never finished" as well as for "finished well" | +| `loaded; no run since this login` (macOS) | launchd has the job and has not started it since the last login. Its exit column says 0, which it prints for "never finished" as well as for "finished well" | | `? the configured log is not there` | the job ran, but the file that would date its last run is gone. Neither health nor failure: the evidence is missing, and that is its own answer | | `! the hooks IGNORE this file entirely: ` | the settings file is out of force; the defaults are running | | `file (IGNORED)` / `5 (file: 9)` | the file says 9, the default 5 is what is actually in effect | -| `this shell` as a source | set here, but launchd's jobs do not inherit it — the line below says what they use | +| `this shell` as a source | set here, but the jobs the service manager starts do not inherit it — the line below says what they use | | `the answer parsed but has no "memories"` | something answered, but it was not this query's result. Not a store full of zeros | +| `systemctl --user is not reachable: ` (Linux) | no session bus, not "no jobs configured" — the two must never look alike | +| `... is not implemented on Linux yet` | `run`, `every`, `at`, `--install`, `--uninstall`: read-only for now on the systemd backend | -Schedules are read the way launchd means them: an omitted `StartCalendarInterval` field is a -wildcard, so `{Minute: 15}` is "hourly at :15" and `{Day: 1, Hour: 3}` is "monthly day 1 03:00". -The period follows the coarsest pinned field, which is what keeps a monthly job from being called -overdue every day of the month. +Schedules are read the way each backend means them. On macOS an omitted `StartCalendarInterval` +field is a wildcard, so `{Minute: 15}` is "hourly at :15" and `{Day: 1, Hour: 3}` is "monthly day 1 +03:00". On Linux the same rule applies to `OnCalendar=`: `*:15:00` is the same "hourly at :15", not +"00:15". Either way the period follows the coarsest pinned field, which is what keeps a monthly job +from being called overdue every day of the month. Two readings that are NOT complaints, and are deliberately not marked: a job with no schedule -(`on demand`) never being "overdue", and a positive run counter with no log to date it by. The -run counter is per-bootstrap — launchd resets it at every login — so it can only ever prove that -something ran, never that nothing did. +(`on demand`) never being "overdue", and — on macOS — a positive run counter with no log to date it +by. The run counter is per-bootstrap — launchd resets it at every login — so it can only ever prove +that something ran, never that nothing did. systemd keeps no such counter at all; where it answers +"when did this last fire" directly (`LastTriggerUSec`), that answer is trusted over the log, and it +settles the question in either direction — "never" is as final an answer there as a timestamp is. diff --git a/docs/INSTALL.md b/docs/INSTALL.md index d2e4109..bea8ba3 100644 --- a/docs/INSTALL.md +++ b/docs/INSTALL.md @@ -221,7 +221,7 @@ Three ways to run it without remembering to: | When | How | |------|-----| | on every commit / pull | a `post-commit` and `post-merge` hook in the repo, calling the line above | -| on a timer | cron, a systemd timer, or a launchd agent; on macOS `hypermnesia-jobs` lists and reschedules those | +| on a timer | cron, a systemd timer, or a launchd agent; `hypermnesia-jobs` lists them on both macOS and Linux, and reschedules them on macOS | | in CI | a step on push, if the runner can reach the database | A timer is the simplest and a git hook is the better one: it runs when something actually changed, @@ -232,7 +232,7 @@ documents whose ingested commit is behind HEAD, and component globs that match n doctor` covers the other half -- a missing index, unembedded chunks, two embedding models in one store. Neither needs a schedule to be useful, but both are worth one. -## The console (optional, macOS menu bar + CLI) +## The console (optional, menu bar / system tray + CLI) `console/` is the operator's side: what the store holds, whether the scheduled passes are running, and what the tunables are set to. @@ -244,6 +244,30 @@ cd console && cargo build --release ./target/release/hypermnesia --install # the menu-bar tray, at login (macOS) ``` +The four command-line tools build the same way everywhere; the data layer has no dependencies at +all. The tray itself needs a GUI toolkit, declared per platform so a machine with none of it +installed still builds everything else: + +```bash +# macOS: the tray is part of the default build. +cargo build --release + +# Linux: behind a feature, since it needs GTK3 and libayatana-appindicator development headers +# that a plain `cargo build` should not have to assume are present. +sudo apt install libgtk-3-dev libayatana-appindicator3-dev libxdo-dev # Debian/Ubuntu +sudo dnf install gtk3-devel libayatana-appindicator-gtk3-devel libxdo-devel # Fedora +cargo build --release --features tray +``` + +**Linux, what works today:** the tray icon, the store's numbers, and reading systemd user units +(`~/.config/systemd/user/*.timer`) -- schedule, last exit code, whether it is running, and when it +last fired, read straight from `systemctl --user show` with no date-parsing crate. Editing a +schedule, running a job on demand, and installing the tray into autostart are not implemented on +Linux yet; `hypermnesia-jobs` and the tray say so rather than doing nothing silently. See +[the console message table](DIAGNOSTICS.md#what-the-console-tells-you-about-itself) for the exact +Linux-side lines, and `HM_JOB_PREFIX` below for how job names differ (`hypermnesia-extract.timer`, +not `com.hypermnesia.extract.plist`). + It reaches the database through ONE setting: a command that receives SQL on stdin and prints unaligned rows. The wizard offers the four usual shapes — direct psql, `docker exec`, `kubectl exec`, ssh to a machine that has kubectl — and tries the command before saving it. @@ -256,7 +280,8 @@ unaligned rows. The wizard offers the four usual shapes — direct psql, `docker Both are **refused entirely** — loudly, falling back to defaults — unless the file is yours, is unwritable by any other account, and sits in a directory with the same property. The first holds a command run through `sh -c` at every refresh and at login with nobody present; the second sets the -environment of the jobs launchd starts. `HM_PSQL_CMD` in the environment beats the config file. +environment of the jobs the service manager starts. `HM_PSQL_CMD` in the environment beats the +config file. A password inside the psql command ends up in psql's argv, where any process running as you can read it. `~/.pgpass` or `PGPASSWORD` in the command's own environment avoids that. @@ -265,7 +290,8 @@ read it. `~/.pgpass` or `PGPASSWORD` in the command's own environment avoids tha variables; names that decide what gets EXECUTED (`HM_PYTHON`, `HM_MEM_OPS`, `HM_SEARCH`, `HM_RERANK`, `HM_LLM_CMD`, `PATH`, `PYTHONPATH`) are refused by name. A variable already set in the environment beats the file — but "the environment" means the shell that started the process, and -launchd gives its jobs none of yours, so for the scheduled passes the file is what is in force. +neither launchd nor a systemd user unit gives its jobs any of yours, so for the scheduled passes +the file is what is in force. ## Configuration reference @@ -315,7 +341,7 @@ shared settings file below, or from the default — in that order. `~/.claude/hypermnesia.env` — one `KEY=value` file for the tunables above, written by `hypermnesia-settings set `, loaded on import by `hooks/_mem_common.py` and `ingest/_common.py`. `HYPERMNESIA_ENV_FILE` moves it. The rules it is read under are in -[The console](#the-console-optional-macos-menu-bar--cli) above: refused whole unless it is +[The console](#the-console-optional-menu-bar--system-tray--cli) above: refused whole unless it is your own private file, and only the exact names listed here are accepted. ### The console @@ -323,6 +349,6 @@ your own private file, and only the exact names listed here are accepted. | Env | Default | Meaning | |-----|---------|---------| | `HM_PSQL_CMD` | `psql "$DATABASE_URL" -tAX -v ON_ERROR_STOP=1` | the command the console sends SQL to on stdin. The one setting that differs between deployments | -| `HM_JOB_PREFIX` | `com.hypermnesia` | the launchd label prefix the job tools manage | +| `HM_JOB_PREFIX` | `com.hypermnesia` (macOS) / `hypermnesia-` (Linux) | the label/unit-name prefix the job tools manage | | `HM_CONSOLE_CONFIG` | `~/.config/hypermnesia/console.conf` | where the two settings above are stored. Refused whole if another account can write it: it holds a command the tray runs at login | | `HM_TIMEOUT_SECS` | `30` | ceiling on one reading of the store | From 6ea173851cd0ae5864f8cba0328b04bff881c5e9 Mon Sep 17 00:00:00 2001 From: Vladimir nett00n Budylnikov Date: Wed, 23 Sep 2026 09:59:18 +0400 Subject: [PATCH 2/2] full interactivity for the linux tray --- README.md | 30 +- console/src/bin/jobs.rs | 101 +++--- console/src/bin/tray.rs | 301 ++++++++++++----- console/src/jobs/launchd.rs | 13 +- console/src/jobs/mod.rs | 29 +- console/src/jobs/systemd.rs | 558 +++++++++++++++++++++++++++++--- console/src/jobs/unsupported.rs | 6 + console/src/view.rs | 1 + docs/DIAGNOSTICS.md | 7 +- docs/INSTALL.md | 21 +- 10 files changed, 889 insertions(+), 178 deletions(-) diff --git a/README.md b/README.md index 0220b06..ae11d50 100644 --- a/README.md +++ b/README.md @@ -286,7 +286,7 @@ cannot be told from an empty store. | Command | What | |---------|------| -| `hypermnesia` | the menu-bar tray (macOS) / system tray (Linux, `--features tray`; jobs are read-only there so far) | +| `hypermnesia` | the menu-bar tray (macOS) / system tray (Linux, `--features tray`) | | `hypermnesia-stats` | the same numbers on stdout | | `hypermnesia-jobs` | scheduled passes: what is configured, when each last worked, run one now, change a schedule | | `hypermnesia-settings` | the tunables, each with the value in force and where that value came from | @@ -308,17 +308,27 @@ What is in the menu: - **the first line is the age of what you are reading.** Not a status light: "updated 4m ago, in 1.2s", and after a failure "! not updated: , showing state from 12m ago". It is computed when the menu is drawn, from the wall clock, so a Mac that slept for two hours says two hours. + The store's own numbers refresh every minute; the jobs — a local, cheap read on either platform + — refresh every few seconds, faster still for a little while after a button is pressed, so + pressing one does not mean waiting out the idle cadence to see whether it worked. - **the volumes** — memories active of total, knowledge pages, documents and chunks, how many chunks have no embedding, how many embedding models are in the store, the review queue, the stale count, the database size. - **Run now** — every scheduled pass, with its schedule, when it last wrote to its log, and its - last exit code. On macOS, pressing one runs it through launchd and reports what launchd then - did, not that the request was accepted. On Linux this is read-only so far: the row shows the - same information, read from `systemctl --user show`, but pressing it reports "not implemented". -- **Schedule** — the common intervals and times, per job. On macOS it edits the plist, validates - it, and reloads the job, because launchd keeps its own copy from the moment it loaded it: writing - the file without the reload would show a new schedule while the old one is in force. Not - implemented on Linux yet. + last exit code. Pressing one runs it — through launchd on macOS, through `systemctl --user start + --no-block` on Linux — and reports what the service manager then did, not that the request was + accepted: a program that is not there fails immediately, and that is read back rather than taken + on faith. +- **Schedule** — the common intervals and times, per job. It writes the unit, validates it, and + reloads the job, because both launchd and systemd keep their own copy from the moment they + loaded it: writing the file without the reload would show a new schedule while the old one is + still in force. A backup goes down next to the file first, and on any slip along the way it is + restored — what actually landed is read back and compared to what was asked before the tray + calls it done. +- **Timers** (Linux only) — arm or disarm a job's schedule without removing it. systemd tracks + "loaded but not armed" as a state of its own, which the read side of this tray could already + show as a fault; this is what fixes it. launchd has no such state — a job it has loaded runs on + its schedule, full stop — so there is nothing here to switch on macOS. - **Refresh now** and **Quit**. The tray carries one mark for "something here needs a look", derived from the lines below rather @@ -327,8 +337,8 @@ design in one detail — the mark is where a problem is noticed, so it must not with the menu. On macOS the mark is in the menu-bar title, `HM !`; a system tray icon has no text of its own, so on Linux it is the icon itself, calm green or warned red. -`hypermnesia --install` puts it in launchd to start at login, restarted if it crashes and not if -you quit it, on macOS. `--uninstall` takes it back out. Not implemented on Linux yet. +`hypermnesia --install` puts it in autostart to start at login, restarted if it crashes and not if +you quit it — launchd on macOS, a systemd user unit on Linux. `--uninstall` takes it back out. **Try it without a database.** The console's only connection setting is a command that prints the query's answer, so a file works: diff --git a/console/src/bin/jobs.rs b/console/src/bin/jobs.rs index 108a65b..4c25b34 100644 --- a/console/src/bin/jobs.rs +++ b/console/src/bin/jobs.rs @@ -1,8 +1,12 @@ -//! The pipeline's scheduled jobs: show them and run them. +//! The pipeline's scheduled jobs: show them, run them, and change when they run. //! //! hypermnesia-jobs what is configured and when it last worked //! hypermnesia-jobs run run now, without touching the schedule //! hypermnesia-jobs log the last lines of the log +//! hypermnesia-jobs every 4h an interval schedule +//! hypermnesia-jobs at 05:30 a calendar schedule +//! hypermnesia-jobs enable arm the schedule (systemd only) +//! hypermnesia-jobs disable disarm it without removing it (systemd only) //! //! The short name is the tail of the label: extract, consolidate, reflect, freshness, rerank. @@ -69,7 +73,21 @@ fn main() { let Some((hour, minute)) = parse_hhmm(time) else { fail("a time is written as HH:MM") }; apply(find(&all, name), &jobs::Schedule::at(hour, minute, day)); } - Some("-h") | Some("--help") => print!("{HELP}"), + Some("enable") => { + let Some(name) = args.get(1) else { fail("give the job's short name") }; + match jobs::set_enabled(find(&all, name), true) { + Ok(msg) => println!("{msg}"), + Err(e) => fail(&e), + } + } + Some("disable") => { + let Some(name) = args.get(1) else { fail("give the job's short name") }; + match jobs::set_enabled(find(&all, name), false) { + Ok(msg) => println!("{msg}"), + Err(e) => fail(&e), + } + } + Some("-h") | Some("--help") => print!("{}", help_text()), Some(other) => fail(&format!("unknown command: {other}")), } } @@ -116,9 +134,12 @@ fn apply(j: &Job, sched: &jobs::Schedule) { Ok(msg) => { println!("{msg}"); println!("was: {was}"); - // launchd keeps its own copy of the schedule from the moment it loaded the job, so the - // job is reloaded. That resets the run counter -- which has to be said out loud, or - // the next look at the list will be alarming: a zero where everything is fine. + // The service manager keeps its own copy of the schedule from the moment it loaded + // the job, so the job is reloaded. That resets whatever run counter it keeps -- which + // has to be said out loud on the platform that has one, or the next look at the list + // will be alarming: a zero where everything is fine. systemd keeps no such counter + // (`Job::runs` is always `None` there), so there is nothing to warn about on Linux. + #[cfg(target_os = "macos")] println!("the job was reloaded into launchd; the run counter started over"); } Err(e) => fail(&e), @@ -159,7 +180,7 @@ fn parse_weekday(s: &str) -> Option { fn find<'a>(all: &'a [Job], name: &str) -> &'a Job { match all.iter().find(|j| j.short() == name || j.label == name) { - Some(j) if j.broken() => fail(&format!( + Some(j) if j.unwritable() => fail(&format!( "{name}: this {} cannot be read ({}). {} may still be running the copy it loaded \ before the file broke -- what it will do at the next login is the question. Fix \ the file first.", jobs::UNIT_NOUN, j.fault.clone().unwrap_or_default(), @@ -239,6 +260,9 @@ fn list(all: &[Job]) { println!(); println!("run — run now; log [N] — the tail of the log; \ every/at … — change the schedule (--help)"); + if jobs::SUPPORTS_ENABLE { + println!("enable/disable — arm or disarm the schedule without removing it"); + } } fn runs_note(j: &Job) -> String { @@ -256,45 +280,46 @@ fn ago(d: std::time::Duration) -> String { else { format!("{} d ago", s / 86400) } } -// The schedule-editing mechanism described below is real on macOS -- a plist is written, linted -// and reloaded into launchd, with a `.bak` for safety -- and is simply not implemented yet on -// Linux (`jobs::set_schedule` returns "not implemented" there). Two texts, not one interpolated -// with `jobs::BACKEND_NAME`, because the difference is not a noun, it is a paragraph that is -// false on one of the two platforms. +// One text, not two: both backends now write, validate, reload and read back a schedule, and +// differ only in the nouns `jobs::UNIT_NOUN`/`BACKEND_NAME`/`UNITS_LOCATION` already carry. +// `enable`/`disable` is the one command that is not portable -- launchd has no separate "loaded +// but not armed" state, so it is documented as systemd-only rather than interpolated away. #[cfg(target_os = "macos")] -const HELP: &str = "\ -hypermnesia-jobs — the memory pipeline's scheduled jobs (launchd). +const RUN_NOW: &str = "run immediately (launchctl kickstart -k)"; +#[cfg(target_os = "linux")] +const RUN_NOW: &str = "run immediately (systemctl --user start --no-block)"; +#[cfg(not(any(target_os = "macos", target_os = "linux")))] +const RUN_NOW: &str = "run immediately"; + +fn help_text() -> String { + // launchd has no separate "loaded but not armed" state -- a job it has loaded is armed, full + // stop -- so `enable`/`disable` is documented only where `jobs::SUPPORTS_ENABLE` says it does + // something, rather than interpolating a noun into a paragraph that would be false on macOS. + let enable = if jobs::SUPPORTS_ENABLE { + format!(" hypermnesia-jobs enable arm the schedule\n\ + \x20 hypermnesia-jobs disable disarm it without removing it\n") + } else { + String::new() + }; + format!("\ +hypermnesia-jobs — the memory pipeline's scheduled jobs ({backend}). hypermnesia-jobs what is configured, when it worked, the last exit code - hypermnesia-jobs run run immediately (launchctl kickstart -k) - hypermnesia-jobs log [N] the last N lines of the log (40 by default) + hypermnesia-jobs run {run_now} + hypermnesia-jobs log [N] the last N lines of the log (40 by default), where one \ +is configured hypermnesia-jobs every 4h an interval schedule (30s, 15m, 4h, 2d) hypermnesia-jobs at 05:30 daily at that time hypermnesia-jobs at mon 06:10 weekly on that day - -Editing a schedule writes the plist, validates it and reloads the job into launchd: without the -reload launchd keeps working from its own copy, and the new schedule would be displayed where the -old one is in force. Before the edit a .bak is put down next to the file, and on any slip the +{enable} +Editing a schedule writes the {noun}, validates it and reloads the job into {backend}: without +the reload {backend} keeps working from its own copy, and the new schedule would be shown where +the old one is in force. Before the edit a .bak is put down next to the file, and on any slip the file is restored. The name is the tail of the label: extract, consolidate, reflect, freshness, rerank. The label prefix comes from HM_JOB_PREFIX, then the console config file, then the -default com.hypermnesia. -"; - -#[cfg(not(target_os = "macos"))] -const HELP: &str = "\ -hypermnesia-jobs — the memory pipeline's scheduled jobs (systemd user units). - - hypermnesia-jobs what is configured, when it worked, the last exit code - hypermnesia-jobs log [N] the last N lines of the log (40 by default), where one is \ -configured - -`run`, `every` and `at` are read on this platform, but not yet implemented: this build can list -systemd user timers and services, not edit or trigger them. Use `systemctl --user start -.service` and `systemctl --user edit .timer` directly for now. - -The name is the tail of the unit's file stem: extract, consolidate, reflect, freshness, rerank. -The unit prefix comes from HM_JOB_PREFIX, then the console config file, then the -default hypermnesia-. -"; +default {prefix}. +", backend = jobs::BACKEND_NAME, run_now = RUN_NOW, noun = jobs::UNIT_NOUN, + prefix = jobs::DEFAULT_JOB_PREFIX) +} diff --git a/console/src/bin/tray.rs b/console/src/bin/tray.rs index 6164c2e..cd94844 100644 --- a/console/src/bin/tray.rs +++ b/console/src/bin/tray.rs @@ -33,19 +33,28 @@ mod app { use tray_icon::menu::{Menu, MenuEvent, MenuItem, PredefinedMenuItem, Submenu}; use tray_icon::{Icon, TrayIcon, TrayIconBuilder}; - /// How often to refresh on its own. A minute: the numbers move slowly and every reading is a - /// round trip to the store. + /// How often to refresh the store's own numbers on its own. A minute: they move slowly and + /// every reading is a round trip over whatever transport the store is configured with. pub const REFRESH: Duration = Duration::from_secs(60); + /// How often to re-read the jobs on their own, apart from the store. A local `systemctl`/ + /// `launchctl` call is cheap -- cheap enough that a person watching a job they just started + /// should not be stuck behind the store's minute-long cadence to see it finish. + const JOBS_REFRESH: Duration = Duration::from_secs(5); + + /// How long after a write command (run, schedule, enable/disable) to keep polling the jobs at + /// the faster `BURST_REFRESH` cadence, so the result of pressing a button shows up in a + /// couple of seconds rather than waiting out `JOBS_REFRESH`. + const BURST_WINDOW: Duration = Duration::from_secs(10); + const BURST_REFRESH: Duration = Duration::from_secs(1); + /// How long the worker may be silent after being asked something before the menu says so. The /// store's own timeout is 30 s by default, so this is well past any normal answer: the point /// is to notice a worker that will never answer at all, not to hurry a slow one. const WORKER_PATIENCE: Duration = Duration::from_secs(90); - /// Ready-made schedules offered in the menu. Exactly what is usually wanted and nothing more: - /// the rare case belongs on the command line -- where, on Linux today, it is the only place: - /// `jobs::set_schedule` refuses with "not implemented on Linux yet", and picking one of these - /// items just shows that refusal in the note line rather than silently doing nothing. + /// Ready-made schedules offered in the menu. Exactly what is usually wanted and nothing more + /// -- the rare case belongs on the command line, on every platform this tray runs on now. const PRESETS: &[(&str, jobs::Schedule)] = &[ ("hourly", jobs::Schedule::Every(3600)), ("every 4 hours", jobs::Schedule::Every(14_400)), @@ -58,13 +67,13 @@ mod app { hypermnesia — the memory console in the menu bar / system tray. hypermnesia run the tray - hypermnesia --install add to autostart, where supported - hypermnesia --uninstall remove from autostart, where supported + hypermnesia --install add to autostart + hypermnesia --uninstall remove from autostart The numbers and the buttons are the same ones hypermnesia-stats and hypermnesia-jobs give: the -tray calls the same functions. Connection settings come from hypermnesia-setup. Autostart and -schedule editing are launchd-only for now; on Linux those actions say so rather than doing -nothing silently. +tray calls the same functions. Connection settings come from hypermnesia-setup. The store's own +numbers refresh every minute; the jobs -- a local, cheap read -- refresh every few seconds, faster +still for a little while after a button is pressed. "; /// Handle `--install` / `--uninstall` / `--help` before any window or GTK loop exists: all @@ -99,16 +108,12 @@ nothing silently. Failed(String), /// The outcome of a button press: a line to show at the top of the menu. Note(String), + /// A jobs-only reading: local and cheap, refreshed on its own cadence from the store's. + Jobs(Vec, Option), } pub struct Snapshot { stats: Stats, - jobs: Vec, - /// Why the job list is missing, if it is. `unwrap_or_default()` used to turn "could not - /// read the job directory" into an empty list, which rendered as empty Run-now and - /// Schedule submenus and a calm title -- "no jobs readable" and "every job healthy" - /// looked identical. - jobs_error: Option, /// Wall clock, not `Instant`: the age of the data has to survive the machine sleeping, /// and a monotonic clock stops while it does. That made a two-hour nap look like a fresh /// reading. @@ -116,39 +121,54 @@ nothing silently. } pub enum Command { + /// Both the store and the jobs. Refresh, + /// The jobs alone -- local, fast, and asked for far more often than `Refresh`. Does not + /// touch `awaiting`: this is a background heartbeat, not a question the watchdog should + /// hold the tray to. + RefreshJobs, RunJob(String), SetSchedule(String, jobs::Schedule), + SetEnabled(String, bool), + } + + fn read_jobs() -> (Vec, Option) { + // Through jobs::job_prefix rather than a cached value: the tray runs for weeks, and a + // prefix fixed in the config file in the meantime has to reach it without a restart. + match jobs::list(&jobs::job_prefix()) { + Ok(j) => (j, None), + Err(e) => (Vec::new(), Some(e)), + } } /// The worker: every slow thing lives here. The main thread only posts commands to it. pub fn spawn_worker(tx: mpsc::Sender, rx: mpsc::Receiver) { std::thread::spawn(move || { - let read_all = |tx: &mpsc::Sender| { + let read_store = |tx: &mpsc::Sender| { // Re-read the settings on every pass rather than once at startup. The tray runs // for weeks; a connection fixed with the setup wizard in the meantime has to // reach the running tray, or the fix looks like it did not work. - let target = Target::default(); - let prefix = jobs::job_prefix(); - // Jobs are read locally and fast; the store is read over whatever transport was - // configured and may be slow. If the store does not answer, the jobs are still - // worth showing. - let (job_list, jobs_error) = match jobs::list(&prefix) { - Ok(j) => (j, None), - Err(e) => (Vec::new(), Some(e)), - }; - match fetch(&target) { + match fetch(&Target::default()) { Ok(stats) => { - let _ = tx.send(Update::Data(Box::new(Snapshot { - stats, jobs: job_list, jobs_error, at: SystemTime::now(), - }))); + let _ = tx.send(Update::Data(Box::new(Snapshot { stats, at: SystemTime::now() }))); } Err(e) => { let _ = tx.send(Update::Failed(e)); } } }; + let read_jobs_into = |tx: &mpsc::Sender| { + let (jobs, err) = read_jobs(); + let _ = tx.send(Update::Jobs(jobs, err)); + }; while let Ok(cmd) = rx.recv() { match cmd { - Command::Refresh => read_all(&tx), + Command::Refresh => { + // Jobs first: they are local and fast, and worth showing even if the + // store -- reached over whatever transport is configured, possibly slow + // -- does not answer. + read_jobs_into(&tx); + read_store(&tx); + } + Command::RefreshJobs => read_jobs_into(&tx), Command::RunJob(label) => { let msg = match jobs::run_now(&label) { Ok(m) => m, @@ -156,14 +176,13 @@ nothing silently. Err(e) => format!("! did not start: {e}"), }; let _ = tx.send(Update::Note(msg)); - read_all(&tx); + read_jobs_into(&tx); } Command::SetSchedule(label, sched) => { - // Look the job up again: the list the menu was drawn from may be a - // minute old, and editing a schedule from a stale record means editing + // Look the job up again: the list the menu was drawn from may be several + // seconds old, and editing a schedule from a stale record means editing // something other than what the person saw. - let msg = match jobs::list(&jobs::job_prefix()).ok() - .and_then(|all| all.into_iter().find(|j| j.label == label)) { + let msg = match read_jobs().0.into_iter().find(|j| j.label == label) { Some(j) => match jobs::set_schedule(&j, &sched) { Ok(m) => m, Err(e) => format!("! {e}"), @@ -171,7 +190,18 @@ nothing silently. None => format!("! the job {label} is gone"), }; let _ = tx.send(Update::Note(msg)); - read_all(&tx); + read_jobs_into(&tx); + } + Command::SetEnabled(label, enabled) => { + let msg = match read_jobs().0.into_iter().find(|j| j.label == label) { + Some(j) => match jobs::set_enabled(&j, enabled) { + Ok(m) => m, + Err(e) => format!("! {e}"), + }, + None => format!("! the job {label} is gone"), + }; + let _ = tx.send(Update::Note(msg)); + read_jobs_into(&tx); } } } @@ -188,6 +218,9 @@ nothing silently. quit: Option, run: Vec<(MenuItem, String)>, sched: Vec<(MenuItem, String, jobs::Schedule)>, + /// One item per job that can be armed/disarmed, carrying the label and the state a click + /// on it should set (the opposite of what it currently is). + enable: Vec<(MenuItem, String, bool)>, } pub struct App { @@ -196,6 +229,12 @@ nothing silently. rx: mpsc::Receiver, cmd: mpsc::Sender, last: Option, + jobs: Vec, + /// Why the job list is missing, if it is. `unwrap_or_default()` used to turn "could not + /// read the job directory" into an empty list, which rendered as empty Run-now and + /// Schedule submenus and a calm title -- "no jobs readable" and "every job healthy" + /// looked identical. + jobs_error: Option, status: String, /// The outcome of the last button press, on its own line and with its own age. /// @@ -204,11 +243,26 @@ nothing silently. /// "updated just now". The outcome of something a person did is the last thing that /// should be overwritten. note: Option<(String, SystemTime)>, - /// What the worker was last asked, and when. Cleared by any answer. + /// What the worker was last asked, and when. Cleared by any answer to it. Only `Refresh` + /// and the write commands set this -- a background `RefreshJobs` tick is not a question + /// the watchdog should hold the tray to. awaiting: Option<(String, Instant)>, /// The worker thread is gone: nothing will ever be read again. worker_dead: bool, last_refresh: Instant, + last_jobs_refresh: Instant, + /// A `RefreshJobs` was sent and no `Update::Jobs` has answered it yet. Without this a + /// worker stuck on a slow `systemctl`/`launchctl` call would be sent a new request every + /// tick, piling up behind the one already running. + jobs_in_flight: bool, + /// Poll the jobs at `BURST_REFRESH` rather than `JOBS_REFRESH` until this instant -- + /// set after every write command, so its result shows up in a couple of seconds rather + /// than waiting out the idle cadence. + burst_until: Instant, + /// The content last drawn, so an unchanged jobs-only reading at `JOBS_REFRESH` cadence + /// does not tear down and rebuild the menu -- which can close one open under the cursor + /// -- for no visible difference. + last_render: Option, } /// What a platform's event loop should do after a tick. @@ -221,25 +275,33 @@ nothing silently. impl App { pub fn new(rx: mpsc::Receiver, cmd: mpsc::Sender) -> Self { + let now = Instant::now(); App { tray: None, items: Items::default(), rx, cmd, last: None, + jobs: Vec::new(), + jobs_error: None, status: "loading…".to_string(), note: None, - awaiting: Some(("the first reading".into(), Instant::now())), + awaiting: Some(("the first reading".into(), now)), worker_dead: false, - last_refresh: Instant::now(), + last_refresh: now, + last_jobs_refresh: now, + jobs_in_flight: false, + burst_until: now, + last_render: None, } } - /// Post a command and remember that an answer is owed. Every path to the worker goes - /// through here so that nothing can be asked without the watchdog knowing about it. + /// Post a command and remember that an answer is owed. Every path to the worker that the + /// watchdog should track goes through here. fn ask(&mut self, cmd: Command, status: &str, what: &str) { self.status = status.to_string(); self.awaiting = Some((what.to_string(), Instant::now())); + self.burst_until = Instant::now() + BURST_WINDOW; let _ = self.cmd.send(cmd); } @@ -248,36 +310,47 @@ nothing silently. pub fn tick(&mut self) -> Tick { if self.tray.is_none() { self.rebuild(); + self.last_render = Some(self.content_fingerprint()); } let mut dirty = false; loop { match self.rx.try_recv() { - Ok(u) => { + Ok(Update::Data(s)) => { self.awaiting = None; - match u { - Update::Data(s) => { - // Not a sentence: the sentence is rendered at rebuild time from - // `last.at`. Frozen, it said "updated just now" for the whole - // refresh cycle -- and after the machine slept, for as long as - // the nap lasted, because the refresh timer is monotonic and - // stops with it. - self.status = String::new(); - self.last = Some(*s); - } - Update::Failed(e) => { - // The old numbers stay -- they beat an empty screen -- but they - // are labelled with the reason and with HOW OLD they are. - // Without the age you cannot tell "could not reach it a second - // ago" from "stuck since yesterday". - let age = self.last.as_ref() - .map(|s| format!(", showing state from {} ago", - view::ago(view::age_of(s.at)))) - .unwrap_or_default(); - self.status = format!("! not updated: {e}{age}"); - } - Update::Note(n) => self.note = Some((n, SystemTime::now())), - } + // Not a sentence: the sentence is rendered at rebuild time from `last.at`. + // Frozen, it said "updated just now" for the whole refresh cycle -- and + // after the machine slept, for as long as the nap lasted, because the + // refresh timer is monotonic and stops with it. + self.status = String::new(); + self.last = Some(*s); + dirty = true; + } + Ok(Update::Failed(e)) => { + self.awaiting = None; + // The old numbers stay -- they beat an empty screen -- but they are + // labelled with the reason and with HOW OLD they are. Without the age you + // cannot tell "could not reach it a second ago" from "stuck since + // yesterday". + let age = self.last.as_ref() + .map(|s| format!(", showing state from {} ago", + view::ago(view::age_of(s.at)))) + .unwrap_or_default(); + self.status = format!("! not updated: {e}{age}"); + dirty = true; + } + Ok(Update::Note(n)) => { + self.awaiting = None; + self.note = Some((n, SystemTime::now())); + dirty = true; + } + Ok(Update::Jobs(jobs, err)) => { + // Deliberately does NOT clear `awaiting`: a background jobs poll is not + // an answer to whatever `Refresh` or a write command asked, and treating + // it as one would let a hung store fetch hide behind it forever. + self.jobs_in_flight = false; + self.jobs = jobs; + self.jobs_error = err; dirty = true; } Err(mpsc::TryRecvError::Empty) => break, @@ -330,6 +403,14 @@ nothing silently. dirty = true; continue; } + if let Some((_, label, enabled)) = + self.items.enable.iter().find(|(i, _, _)| i.id() == &ev.id) { + let (label, enabled) = (label.clone(), *enabled); + let what = format!("{} {label}", if enabled { "enabling" } else { "disabling" }); + self.ask(Command::SetEnabled(label, enabled), &format!("{what}…"), &what); + dirty = true; + continue; + } // A press on an item from a menu that has since been rebuilt: the ids are new, // so nothing above matched. Silently dropping it left a person pressing a button // that did nothing, with no way to tell that from a button that did nothing @@ -343,12 +424,60 @@ nothing silently. self.last_refresh = Instant::now(); self.ask(Command::Refresh, &self.status.clone(), "a reading"); } + // The jobs, cheap and local, on their own faster cadence -- faster still for a while + // after a button was pressed, so pressing it does not mean waiting out the idle + // cadence to see whether it worked. + let jobs_cadence = if Instant::now() < self.burst_until { BURST_REFRESH } + else { JOBS_REFRESH }; + if !self.jobs_in_flight && self.last_jobs_refresh.elapsed() >= jobs_cadence { + self.last_jobs_refresh = Instant::now(); + self.jobs_in_flight = true; + let _ = self.cmd.send(Command::RefreshJobs); + } if dirty { - self.rebuild(); + let fp = self.content_fingerprint(); + if self.last_render.as_ref() != Some(&fp) { + self.rebuild(); + self.last_render = Some(fp); + } } Tick::Continue } + /// Everything that would change what is drawn, boiled down to one string -- compared + /// against the last one drawn so a `RefreshJobs` tick that changed nothing does not tear + /// down and rebuild the menu, which can close one open under the cursor, for no visible + /// difference. Deliberately leaves out anything that only ages (the "(N ago)" suffixes): + /// those are cosmetic, and `status_line`'s own coarse buckets already force a redraw when + /// the words themselves would change. + fn content_fingerprint(&self) -> String { + let mut fp = self.status_line(); + fp.push('\n'); + if let Some((note, _)) = &self.note { + fp.push_str(note); + fp.push('\n'); + } + if let Some(why) = jobs::autostart_fault(jobs::DEFAULT_TRAY_LABEL) { + fp.push_str(&why); + fp.push('\n'); + } + if let Some(s) = &self.last { + for line in view::summary(&s.stats) { + fp.push_str(&line); + fp.push('\n'); + } + } + if let Some(why) = &self.jobs_error { + fp.push_str(why); + fp.push('\n'); + } + for j in &self.jobs { + fp.push_str(&view::job_line(j)); + fp.push_str(&format!(" armed={:?}\n", j.armed)); + } + fp + } + fn rebuild(&mut self) { let menu = Menu::new(); let mut items = Items::default(); @@ -372,12 +501,12 @@ nothing silently. } let _ = menu.append(&PredefinedMenuItem::separator()); - if let Some(why) = &s.jobs_error { + if let Some(why) = &self.jobs_error { let _ = menu.append(&MenuItem::new(format!("! jobs unreadable: {why}"), false, None)); } let run = Submenu::new("Run now", true); - for j in &s.jobs { + for j in &self.jobs { let item = MenuItem::new(view::job_line(j), true, None); let _ = run.append(&item); items.run.push((item, j.label.clone())); @@ -385,8 +514,8 @@ nothing silently. let _ = menu.append(&run); let sched = Submenu::new("Schedule", true); - for j in &s.jobs { - if matches!(j.schedule, jobs::Schedule::None) || j.broken() { + for j in &self.jobs { + if matches!(j.schedule, jobs::Schedule::None) || j.unwritable() { continue; // nothing to change, or nothing readable to change } let sub = Submenu::new(format!("{} ({})", j.short(), j.schedule.human()), @@ -400,7 +529,33 @@ nothing silently. } let _ = menu.append(&sched); - warned = view::warned(&s.stats, &s.jobs, &s.jobs_error); + // A backend that tracks "armed" at all (systemd; not launchd, where a loaded job + // is armed by definition and there is nothing here to switch) gets a submenu to + // fix the one fault the read side could already see but not act on. + if jobs::SUPPORTS_ENABLE { + let timers = Submenu::new("Timers", true); + let mut any = false; + for j in &self.jobs { + // A job the backend could not load at all (`armed` is only ever `None` + // there) has nothing to arm; a merely disabled one -- `armed == + // Some(false)`, which also carries a `fault` from `enablement_fault` -- + // is exactly the case this submenu exists to fix, so it is NOT excluded + // here the way `j.broken()` alone would exclude it. + let Some(armed) = j.armed else { continue }; + any = true; + let label = format!("{} ({})", j.short(), + if armed { "enabled" } else { "disabled" }); + let action = if armed { "Disable" } else { "Enable" }; + let item = MenuItem::new(format!("{label} -- {action}"), true, None); + let _ = timers.append(&item); + items.enable.push((item, j.label.clone(), !armed)); + } + if any { + let _ = menu.append(&timers); + } + } + + warned = view::warned(&s.stats, &self.jobs, &self.jobs_error); } // The status line and the last note are as much a part of "does this need a mark" as // the store's own numbers: a failed refresh has no `Snapshot` to derive a warning diff --git a/console/src/jobs/launchd.rs b/console/src/jobs/launchd.rs index 73d00fb..e90728e 100644 --- a/console/src/jobs/launchd.rs +++ b/console/src/jobs/launchd.rs @@ -78,6 +78,7 @@ fn broken_job(path: &PathBuf, name: &str, fault: String, log_state: LogState::NotConfigured, trigger: super::TriggerState::NotTracked, fault: Some(fault), + armed: None, } } @@ -102,6 +103,7 @@ fn read_plist(path: &PathBuf, states: &BTreeMap, Option Option { (installed != me).then(|| format!("autostart starts {installed}, not this binary ({me})")) } +/// launchd has no separate "loaded but not armed" state: a job it has loaded runs on its +/// schedule, full stop. There is nothing here for a menu toggle to switch. +pub const SUPPORTS_ENABLE: bool = false; + +pub fn set_enabled(_job: &Job, _enabled: bool) -> Result { + Err("launchd has no separate enable/disable step -- a loaded job is armed; unload it with \ + launchctl bootout instead".into()) +} + /// Take the console out of autostart. pub fn uninstall_self(label: &str) -> Result { let home = std::env::var("HOME").map_err(|_| "no HOME".to_string())?; @@ -514,7 +525,7 @@ mod tests { Job { label: "x".into(), source: PathBuf::new(), schedule, program: vec![], log: None, last_exit: None, pid: None, log_state: LogState::NotConfigured, runs: None, - installed: None, trigger: super::TriggerState::NotTracked, fault: None, + installed: None, trigger: super::TriggerState::NotTracked, fault: None, armed: None, } } diff --git a/console/src/jobs/mod.rs b/console/src/jobs/mod.rs index f33fdf7..4bfab83 100644 --- a/console/src/jobs/mod.rs +++ b/console/src/jobs/mod.rs @@ -21,17 +21,20 @@ use std::time::{Duration, SystemTime}; #[cfg(target_os = "macos")] mod launchd; #[cfg(target_os = "macos")] -pub use launchd::{autostart_fault, install_self, list, run_now, set_schedule, uninstall_self}; +pub use launchd::{autostart_fault, install_self, list, run_now, set_enabled, set_schedule, + uninstall_self, SUPPORTS_ENABLE}; #[cfg(target_os = "linux")] mod systemd; #[cfg(target_os = "linux")] -pub use systemd::{autostart_fault, install_self, list, run_now, set_schedule, uninstall_self}; +pub use systemd::{autostart_fault, install_self, list, run_now, set_enabled, set_schedule, + uninstall_self, SUPPORTS_ENABLE}; #[cfg(not(any(target_os = "macos", target_os = "linux")))] mod unsupported; #[cfg(not(any(target_os = "macos", target_os = "linux")))] -pub use unsupported::{autostart_fault, install_self, list, run_now, set_schedule, uninstall_self}; +pub use unsupported::{autostart_fault, install_self, list, run_now, set_enabled, set_schedule, + uninstall_self, SUPPORTS_ENABLE}; /// Reverse-DNS prefix of the launchd labels this console manages, on macOS. #[cfg(target_os = "macos")] @@ -285,6 +288,12 @@ pub struct Job { /// one that is not running, and that is precisely the case the console used to have nothing /// to say about, because the row was dropped from the list. pub fault: Option, + /// Whether the schedule is armed, where the backend can say. `Some(true)` on systemd is a + /// timer that is `UnitFileState=enabled` and `ActiveState=active`; `Some(false)` is loaded but + /// not armed -- exactly what `enablement_fault` already flags as a fault. `None` on launchd, + /// which has no separate "loaded but not armed" state to report: a plist launchd has loaded is + /// armed, full stop, so there is nothing here to switch. + pub armed: Option, } impl Job { @@ -316,6 +325,18 @@ impl Job { self.fault.is_some() } + /// Wrong enough that a write action should refuse to touch the unit -- unlike + /// `armed == Some(false)`, which means the unit loaded and parsed just fine and is simply + /// disabled, a state `set_enabled`/`set_schedule` exist to fix rather than a reason to + /// refuse them. `armed` is `Some(_)` on systemd whenever the unit could be loaded at all, and + /// `None` only when it could not -- the same distinction `read_unit`/`broken_job` already + /// draw there. On launchd, where enablement is not a separate state, `armed` is always + /// `None`, so this collapses back to plain `broken()` -- exactly what that backend's own + /// `fault` has always meant. + pub fn unwritable(&self) -> bool { + self.fault.is_some() && self.armed.is_none() + } + /// The service manager has never started it. /// /// `runs` alone does not answer this on launchd: it is per-bootstrap state, reset to zero at @@ -458,7 +479,7 @@ mod tests { Job { label: "x".into(), source: PathBuf::new(), schedule, program: vec![], log: None, last_exit: None, pid: None, log_state: LogState::NotConfigured, runs: None, - installed: None, trigger: TriggerState::NotTracked, fault: None, + installed: None, trigger: TriggerState::NotTracked, fault: None, armed: None, } } diff --git a/console/src/jobs/systemd.rs b/console/src/jobs/systemd.rs index 876ca29..0d16756 100644 --- a/console/src/jobs/systemd.rs +++ b/console/src/jobs/systemd.rs @@ -1,13 +1,10 @@ //! The systemd backend: user units under `~/.config/systemd/user/*.{service,timer}`, read through -//! `systemctl --user show`. -//! -//! **Read-only, deliberately.** `list()` is real; `run_now`, `set_schedule`, `install_self` and -//! `uninstall_self` all return "not implemented on Linux yet" rather than a silent no-op -- -//! writing a systemd timer safely (the launchd side of this file is the backup-edit-lint-reload -//! dance in `launchd.rs`, about half its length) is real work this prototype does not yet do. -//! `autostart_fault` returns `None`: an honest "nothing to report" for a check the tray does not -//! yet perform, not a permanent warning about a missing feature -- a warning nobody can act on is -//! a warning nobody reads. +//! `systemctl --user show` and written through the same backup -> edit -> verify -> reload -> +//! read-back -> rollback discipline `launchd.rs` uses for its plists. Nothing here claims success +//! from an exit code alone: every write is checked against what `systemctl --user show` reports +//! afterwards, because a `0` from `systemctl` means "the request was accepted", not "the change +//! took" -- a timer already loaded holds its own copy of the schedule until it is reloaded, and a +//! disabled unit still accepts `daemon-reload` without complaint. //! //! systemd answers two of the three questions this module needs directly, and better than //! launchd does: `ExecMainStatus`, `MainPID` and `LastTriggerUSec` are exactly "exit code", "is it @@ -158,6 +155,8 @@ fn read_unit(stem: &str, timer_path: &PathBuf) -> Job { runs: None, // systemd keeps no lifetime run counter pid, trigger, + armed: Some(get(&timer_rows, "UnitFileState") == Some("enabled") + && get(&timer_rows, "ActiveState") == Some("active")), fault: enablement_fault(&timer_rows), } } @@ -206,6 +205,7 @@ fn broken_job(stem: &str, path: &PathBuf, fault: String) -> Job { runs: None, pid: None, trigger: TriggerState::NotTracked, + armed: None, // could not be read, so nothing is known about it either fault: Some(fault), } } @@ -255,11 +255,16 @@ fn parse_schedule(timer_rows: &[(String, String)]) -> Schedule { 1 => return Schedule::Calendar(cals.into_iter().next().unwrap()), _ => return Schedule::Several(cals), } - if let Some(raw) = get(timer_rows, "TimersMonotonic") { - if let Some(spec) = extract_between(raw, "OnUnitActiveUSec") { - if let Some(secs) = parse_systemd_duration(&spec) { - return Schedule::Every(secs); - } + // `TimersMonotonic` is one row PER monotonic directive -- a unit with `OnBootSec=` next to + // `OnUnitActiveSec=` (which `set_schedule` writes together, so an interval timer also fires + // after a boot) reports two rows, and `get`'s first-match would as likely hand back the + // `OnBootUSec` one, which has no `OnUnitActiveUSec=` key and reads as "no schedule at all". + // `OnUnitActiveUSec` is what this console's own interval schedules are -- and about the only + // monotonic key it would ever need to read back. + if let Some(spec) = get_all(timer_rows, "TimersMonotonic") + .find_map(|raw| extract_between(raw, "OnUnitActiveUSec")) { + if let Some(secs) = parse_systemd_duration(&spec) { + return Schedule::Every(secs); } } Schedule::None @@ -408,30 +413,393 @@ fn parse_standard_output(unit_file: &str) -> Option { .map(PathBuf::from) } -pub fn run_now(_label: &str) -> Result { - Err("running a job on demand is not implemented on Linux yet -- use `systemctl --user start \ - .service` directly".into()) +/// A unit name is a shell argument and a path component both. Rejecting anything but the +/// characters systemd itself allows in an instance/template name keeps every `Command::new` +/// below from ever needing to worry about an escape. +fn valid_stem(stem: &str) -> Result<(), String> { + if stem.is_empty() { + return Err("the job name is empty".into()); + } + let ok = stem.chars().all(|c| c.is_ascii_alphanumeric() || matches!(c, '-' | '_' | '.')) + && !stem.starts_with('-') && !stem.contains(".."); + if ok { + Ok(()) + } else { + Err(format!("{stem:?} is not a valid systemd unit name")) + } } -pub fn set_schedule(_job: &Job, _sched: &Schedule) -> Result { - Err("changing a schedule is not implemented on Linux yet -- edit the `.timer` unit and run \ - `systemctl --user daemon-reload`".into()) +/// Run the job immediately, without touching its schedule. +/// +/// `--no-block` matters: these are `Type=oneshot` services, and a plain `start` waits for the +/// unit to finish before returning -- which would sit on the tray's worker thread for as long as +/// the job takes, well past the 90 s the tray gives up waiting for an answer at all. +/// +/// systemd's `InvocationID` is the counterpart of launchd's run counter: a fresh UUID every time +/// the unit starts, whether or not it succeeds. The state is read back rather than trusted from +/// the exit code of `start`, which only says the request was accepted -- a program that is not +/// there fails immediately and systemd records that asynchronously, the same race `launchd.rs` +/// documents for `kickstart`. +pub fn run_now(label: &str) -> Result { + valid_stem(label)?; + let service = format!("{label}.service"); + let before_id = show(&service).ok() + .and_then(|r| get(&r, "InvocationID").map(str::to_string)); + + let o = Command::new("systemctl").args(["--user", "start", "--no-block", "--", &service]) + .output().map_err(|e| format!("could not call systemctl: {e}"))?; + if !o.status.success() { + let err = String::from_utf8_lossy(&o.stderr).trim().to_string(); + return Err(if err.is_empty() { format!("systemctl start {service} failed") } else { err }); + } + + // A short look-back: long enough for a program that is not there to have failed and for + // systemd to have noticed, short enough that a person does not notice the wait. + let deadline = std::time::Instant::now() + Duration::from_secs(2); + let after = loop { + let rows = show(&service).ok(); + let moved = rows.as_deref().and_then(|r| get(r, "InvocationID")) + .is_some_and(|id| Some(id) != before_id.as_deref()); + let running = rows.as_deref().and_then(|r| get(r, "MainPID")) + .and_then(|v| v.parse::().ok()).is_some_and(|p| p != 0); + if moved || running || std::time::Instant::now() >= deadline { + break rows; + } + std::thread::sleep(Duration::from_millis(150)); + }; + let Some(after) = after else { + return Ok(format!("systemctl accepted the request for {service}; its state could not be \ + read back")); + }; + if get(&after, "MainPID").and_then(|v| v.parse::().ok()).is_some_and(|p| p != 0) { + return Ok(format!("running: {label}")); + } + // The exit code systemd remembers belongs to a run, and only an InvocationID that MOVED + // proves which run that is -- without that check a job that failed yesterday and has not + // started yet would be reported as having just exited. + let moved = get(&after, "InvocationID").is_some_and(|id| Some(id) != before_id.as_deref()); + let exit = get(&after, "ExecMainStatus").and_then(|v| v.parse::().ok()); + match (moved, exit) { + (true, Some(0)) => Ok(format!("ran: {label} (exit 0)")), + (true, Some(c)) => Err(format!("{label} ran and exited {c}")), + (true, None) => Ok(format!("ran: {label} (systemd reports no exit code yet)")), + (false, _) => Err(format!( + "systemctl accepted the request but {label} has not started -- its invocation id \ + has not moved. Check the log.")), + } } -pub fn install_self(_label: &str) -> Result { - Err("installing the tray into autostart is not implemented on Linux yet -- add a \ - `hypermnesia-tray.service` user unit with `WantedBy=default.target` by hand".into()) +/// Replace the schedule keys inside a unit file's `[Timer]` section. Pure text in, text out -- no +/// filesystem, no `systemctl` -- so this is where the tests live rather than against a real +/// system. +fn rewrite_timer(text: &str, sched: &Schedule) -> Result { + let new_line = match sched { + Schedule::Every(secs) => format!("OnUnitActiveSec={secs}\nOnBootSec={secs}"), + Schedule::Calendar(c) => format!("OnCalendar={}", render_oncalendar(c)), + Schedule::Several(_) | Schedule::None => return Err( + "this console sets one schedule at a time; several calendar slots or no schedule \ + at all has to be edited in the unit by hand".into()), + }; + + let lines: Vec<&str> = text.lines().collect(); + let section = lines.iter().position(|l| l.trim() == "[Timer]") + .ok_or_else(|| "the unit has no [Timer] section to edit".to_string())?; + let end = lines[section + 1..].iter().position(|l| l.trim_start().starts_with('[')) + .map(|i| section + 1 + i).unwrap_or(lines.len()); + + const SCHEDULE_KEYS: [&str; 4] = + ["OnCalendar=", "OnUnitActiveSec=", "OnUnitInactiveSec=", "OnBootSec="]; + let mut out: Vec = lines[..=section].iter().map(|s| s.to_string()).collect(); + for l in &lines[section + 1..end] { + if !SCHEDULE_KEYS.iter().any(|k| l.trim_start().starts_with(k)) { + out.push(l.to_string()); + } + } + out.push(new_line); + out.extend(lines[end..].iter().map(|s| s.to_string())); + let mut result = out.join("\n"); + if text.ends_with('\n') { + result.push('\n'); + } + Ok(result) } -pub fn uninstall_self(_label: &str) -> Result { - Err("removing the tray from autostart is not implemented on Linux yet".into()) +/// The inverse of `parse_oncalendar`: an omitted `Cal` field is a wildcard `*`, never a zero -- +/// the same rule `edit_schedule` in `launchd.rs` writes plist calendar keys by. The year is always +/// `*`: `Cal` carries none, and every schedule this console sets is yearly at most. +fn render_oncalendar(c: &Cal) -> String { + const DAYS: [&str; 7] = ["Sun", "Mon", "Tue", "Wed", "Thu", "Fri", "Sat"]; + let num = |v: Option| v.map(|n| format!("{n:02}")).unwrap_or_else(|| "*".to_string()); + let weekday = c.weekday.map(|w| format!("{} ", DAYS[(w as usize) % 7])).unwrap_or_default(); + format!("{weekday}*-{}-{} {}:{}:00", num(c.month), num(c.day), num(c.hour), num(c.minute)) } -/// Not implemented, so there is nothing to be wrong about -- `None`, not a permanent warning. -/// `bin/tray.rs` draws whatever this returns on every single menu rebuild, and the project's own -/// rule is that a warning shown forever is a warning nobody reads. -pub fn autostart_fault(_label: &str) -> Option { - None +/// Check the edited unit before anything is reloaded into systemd -- the counterpart of +/// `plutil -lint`. A unit that fails this check simply is not accepted by systemd, so reloading +/// it anyway would leave the timer quietly unarmed rather than on the new schedule. +/// +/// `systemd-analyze` is not present on every system (some minimal distros ship `systemctl` +/// without it); the fallback asks the same question a different way, through `daemon-reload` and +/// the `LoadState` this module already knows how to read. +fn verify_unit(path: &str) -> Result<(), String> { + match Command::new("systemd-analyze").args(["--user", "verify", path]).output() { + Ok(o) if o.status.success() => Ok(()), + Ok(o) => { + let err = String::from_utf8_lossy(&o.stderr).trim().to_string(); + Err(if err.is_empty() { "systemd-analyze verify failed".into() } else { err }) + } + Err(_) => { + let _ = Command::new("systemctl").args(["--user", "daemon-reload"]).output(); + let unit = std::path::Path::new(path).file_name() + .and_then(|f| f.to_str()).unwrap_or("").to_string(); + match show(&unit) { + Ok(rows) => load_fault(&rows).map_or(Ok(()), Err), + Err(e) => Err(e), + } + } + } +} + +/// Reload systemd's view of the unit files and re-arm the timer. A reload alone is not enough: +/// a running timer holds its own copy of the schedule from the moment it was armed, the same trap +/// `launchd.rs` documents for `bootstrap` -- so the timer is restarted, not merely reloaded. +fn reload_and_restart(label: &str) -> Result<(), String> { + let o = Command::new("systemctl").args(["--user", "daemon-reload"]).output() + .map_err(|e| format!("could not call systemctl: {e}"))?; + if !o.status.success() { + let err = String::from_utf8_lossy(&o.stderr).trim().to_string(); + return Err(if err.is_empty() { "daemon-reload failed".into() } else { err }); + } + let timer = format!("{label}.timer"); + let o = Command::new("systemctl").args(["--user", "restart", "--", &timer]).output() + .map_err(|e| format!("could not call systemctl: {e}"))?; + if !o.status.success() { + let err = String::from_utf8_lossy(&o.stderr).trim().to_string(); + return Err(if err.is_empty() { format!("restart {timer} failed") } else { err }); + } + Ok(()) +} + +/// Change a job's schedule: edit the `.timer` unit and reload it into systemd. +/// +/// Mirrors `launchd.rs::set_schedule`'s discipline, not its mechanics: a `.bak` copy is made +/// first (refusing if one already exists -- the leftover of an edit that did not finish, and the +/// only copy of the working schedule worth trusting), the file is rewritten, checked, reloaded, +/// and finally **read back** through `systemctl --user show` -- the point of every step before it. +/// Any failure restores the backup and says, in words, whether the restore itself worked. +pub fn set_schedule(job: &Job, sched: &Schedule) -> Result { + if matches!(sched, Schedule::None | Schedule::Several(_)) { + return Err("this console changes one calendar slot or one interval at a time; several \ + slots or no schedule at all has to be edited in the unit by hand".into()); + } + if job.unwritable() { + return Err(format!("this unit could not be read ({}) -- I will not edit it", + job.fault.as_deref().unwrap_or(""))); + } + valid_stem(&job.label)?; + let path = job.source.to_string_lossy().to_string(); + let backup = format!("{path}.bak"); + if std::fs::metadata(&backup).is_ok() { + return Err(format!("{backup} already exists: an earlier edit did not finish. Compare it \ + with {path} and remove it before editing again -- it may be the only \ + copy of the working schedule.")); + } + let original = std::fs::read_to_string(&path).map_err(|e| format!("cannot read {path}: {e}"))?; + std::fs::copy(&path, &backup).map_err(|e| format!("cannot make a backup copy {backup}: {e}"))?; + + let rewritten = match rewrite_timer(&original, sched) { + Ok(text) => text, + Err(e) => { + let _ = std::fs::remove_file(&backup); + return Err(e); + } + }; + if let Err(e) = std::fs::write(&path, &rewritten) { + return Err(match std::fs::copy(&backup, &path) { + Ok(_) => { let _ = std::fs::remove_file(&backup); + format!("could not write the unit ({e}); the file was restored from the \ + backup copy") } + Err(c) => format!("could not write the unit ({e}) AND it could not be restored \ + ({c}). The good copy is at {backup}."), + }); + } + + if let Err(e) = verify_unit(&path) { + return Err(match std::fs::copy(&backup, &path) { + Ok(_) => { let _ = std::fs::remove_file(&backup); + format!("the unit does not pass the check after the edit ({e}); the file \ + was restored from the backup copy, the schedule was not changed") } + Err(c) => format!("the unit does not pass the check after the edit ({e}) AND it \ + could NOT be restored ({c}). The good copy is at {backup}."), + }); + } + + if let Err(e) = reload_and_restart(&job.label) { + if let Err(c) = std::fs::copy(&backup, &path) { + return Err(format!("systemd did not accept the new schedule ({e}) AND the file \ + could not be restored ({c}). The good copy is at {backup}; the file on disk \ + carries the new schedule that systemd refused.")); + } + let reloaded = reload_and_restart(&job.label); + let _ = std::fs::remove_file(&backup); + return Err(match reloaded { + Ok(()) => format!("systemd did not accept the new schedule ({e}); the file was \ + restored and the timer is armed again"), + Err(e2) => format!("systemd did not accept the new schedule ({e}); the file was \ + restored BUT the timer could not be reloaded either ({e2}). Fix it with: \ + systemctl --user daemon-reload && systemctl --user restart {}.timer", job.label), + }); + } + + // Read back what actually landed, not what was asked for -- the point of every step above. + let timer_label = format!("{}.timer", job.label); + let landed = show(&timer_label).ok().map(|rows| parse_schedule(&rows)); + if landed.as_ref() != Some(sched) { + let _ = std::fs::copy(&backup, &path); + let _ = reload_and_restart(&job.label); + let _ = std::fs::remove_file(&backup); + return Err(format!( + "systemd reloaded the unit but reports a different schedule than was set ({}) -- \ + the file was restored", landed.map(|s| s.human()).unwrap_or_else(|| "none".into()))); + } + + let _ = std::fs::remove_file(&backup); + Ok(format!("{}: {}", job.short(), sched.human())) +} + +/// Arm or disarm a timer without touching its schedule -- the action `enablement_fault` can +/// already see is needed but the read-only tray could not take. +pub const SUPPORTS_ENABLE: bool = true; + +pub fn set_enabled(job: &Job, enabled: bool) -> Result { + valid_stem(&job.label)?; + let timer = format!("{}.timer", job.label); + let before = show(&timer)?; + if matches!(get(&before, "UnitFileState"), Some("masked" | "masked-runtime")) { + return Err(format!("{timer} is masked -- unmask it first: \ + systemctl --user unmask {timer}")); + } + let verb = if enabled { "enable" } else { "disable" }; + let o = Command::new("systemctl").args(["--user", verb, "--now", "--", &timer]).output() + .map_err(|e| format!("could not call systemctl: {e}"))?; + if !o.status.success() { + let err = String::from_utf8_lossy(&o.stderr).trim().to_string(); + return Err(if err.is_empty() { format!("{verb} {timer} failed") } else { err }); + } + let after = show(&timer)?; + let now_active = get(&after, "ActiveState") == Some("active"); + if now_active != enabled { + return Err(format!("{timer} was asked to {verb}, but its ActiveState is now {:?}", + get(&after, "ActiveState").unwrap_or(""))); + } + Ok(format!("{}: {}", job.short(), if enabled { "enabled" } else { "disabled" })) +} + +/// The tray's own autostart unit, as text. A function so a test can read what is actually +/// written, the same role `tray_plist` plays in `launchd.rs`. +/// +/// `Restart=on-failure`, not `always`: an unconditional restart makes Quit a menu item that lies +/// -- the tray reappearing seconds after someone chose to close it -- which is exactly the trap +/// `launchd.rs`'s `KeepAlive` comment documents for the plist side of this same unit. +/// `WantedBy=default.target` rather than a graphical-session target: not every desktop environment +/// reaches the latter, and `default.target` is the one every session manager pulls in. +fn tray_unit(exe: &str, log: &str) -> String { + format!( + "[Unit]\n\ + Description=HyperMnesia tray\n\ + \n\ + [Service]\n\ + Type=simple\n\ + ExecStart={exe}\n\ + Restart=on-failure\n\ + RestartSec=5\n\ + StandardOutput=append:{log}\n\ + StandardError=append:{log}\n\ + \n\ + [Install]\n\ + WantedBy=default.target\n") +} + +/// Put the console itself into autostart: a unit pointing at the current binary, enabled but not +/// started -- starting it now would run a second tray beside the one already open. +pub fn install_self(label: &str) -> Result { + valid_stem(label)?; + let exe = std::env::current_exe() + .map_err(|e| format!("cannot determine the path to myself: {e}"))? + .to_string_lossy().to_string(); + let home = std::env::var("HOME").map_err(|_| "no HOME".to_string())?; + let dir = units_dir()?; + std::fs::create_dir_all(&dir).map_err(|e| format!("{}: {e}", dir.display()))?; + let path = dir.join(format!("{label}.service")); + let log = format!("{home}/.hypermnesia/logs/tray.log"); + if let Some(d) = PathBuf::from(&log).parent() { + let _ = std::fs::create_dir_all(d); + } + std::fs::write(&path, tray_unit(&exe, &log)).map_err(|e| format!("{}: {e}", path.display()))?; + + if let Err(e) = Command::new("systemctl").args(["--user", "daemon-reload"]).output() { + let _ = std::fs::remove_file(&path); + return Err(format!("could not call systemctl: {e}")); + } + let unit = format!("{label}.service"); + let o = Command::new("systemctl").args(["--user", "enable", "--", &unit]).output() + .map_err(|e| format!("could not call systemctl: {e}"))?; + if !o.status.success() { + let err = String::from_utf8_lossy(&o.stderr).trim().to_string(); + let _ = std::fs::remove_file(&path); + let _ = Command::new("systemctl").args(["--user", "daemon-reload"]).output(); + return Err(format!("the unit was written but systemd would not enable it: {}", + if err.is_empty() { "enable failed".into() } else { err })); + } + Ok(format!("{} installed into autostart\nbinary: {exe}\nlog: {log}\nit starts at your next \ + login; to start it now: systemctl --user start {unit}", path.display())) +} + +/// Take the console out of autostart. +pub fn uninstall_self(label: &str) -> Result { + valid_stem(label)?; + let path = units_dir()?.join(format!("{label}.service")); + let unit = format!("{label}.service"); + let _ = Command::new("systemctl").args(["--user", "disable", "--now", "--", &unit]).output(); + match std::fs::remove_file(&path) { + Ok(()) => { + let _ = Command::new("systemctl").args(["--user", "daemon-reload"]).output(); + Ok(format!("{} removed from autostart", path.display())) + } + Err(e) if e.kind() == std::io::ErrorKind::NotFound => + Err(format!("{} is not there anyway", path.display())), + Err(e) => Err(format!("{}: {e}", path.display())), + } +} + +/// Is the installed autostart unit still pointing at this binary, and still armed? +/// +/// A unit that has never been installed is `None`, not a fault -- `install_self` not having been +/// run is not the same claim as "checked, and something is wrong". A unit that exists but points +/// somewhere else, or is masked/disabled, is the same quiet failure `launchd.rs`'s version of this +/// function was written to catch: a tray moved, rebuilt elsewhere, or installed from a copy that +/// has since been deleted, failing at every login with nothing on screen to say so. +pub fn autostart_fault(label: &str) -> Option { + let path = units_dir().ok()?.join(format!("{label}.service")); + let text = match std::fs::read_to_string(&path) { + Ok(t) => t, + Err(e) if e.kind() == std::io::ErrorKind::NotFound => return None, + Err(e) => return Some(format!("{}: {e}", path.display())), + }; + let installed = text.lines().find_map(|l| l.trim().strip_prefix("ExecStart="))?.to_string(); + if !std::path::Path::new(&installed).exists() { + return Some(format!("autostart points at {installed}, which is not there any more")); + } + let me = std::env::current_exe().ok()?.to_string_lossy().to_string(); + if installed != me { + return Some(format!("autostart starts {installed}, not this binary ({me})")); + } + let rows = show(&format!("{label}.service")).ok()?; + match get(&rows, "UnitFileState") { + Some("disabled" | "masked" | "masked-runtime") => Some(format!( + "the tray's autostart unit is {}", get(&rows, "UnitFileState").unwrap_or(""))), + _ => None, + } } #[cfg(test)] @@ -508,6 +876,21 @@ mod tests { assert_eq!(parse_schedule(&seconds), Schedule::Every(90)); } + /// `set_schedule` writes `OnBootSec=` next to `OnUnitActiveSec=`, so systemd reports TWO + /// `TimersMonotonic` rows -- one per directive. Reproduced against a real unit: `systemctl + /// show` lists the `OnBootUSec` row first, and `get`'s first-match used to hand that one + /// back, find no `OnUnitActiveUSec=` key in it, and read the whole timer as unscheduled -- + /// which would have made `set_schedule`'s own read-back check fail every single interval + /// schedule it had just written. + #[test] + fn a_monotonic_schedule_is_found_even_behind_an_onboot_row() { + let r = rows(&[ + ("TimersMonotonic", "{ OnBootUSec=1h ; next_elapse=1h }"), + ("TimersMonotonic", "{ OnUnitActiveUSec=1h ; next_elapse=12h }"), + ]); + assert_eq!(parse_schedule(&r), Schedule::Every(3_600)); + } + #[test] fn a_timestamp_round_trips_through_civil_arithmetic() { // Checked against `date -u -d "2026-09-22 16:45:31 UTC" +%s` => 1790095531. @@ -595,24 +978,109 @@ mod tests { assert_eq!(parse_standard_output("[Service]\nType=oneshot\n"), None); } - #[test] - fn write_actions_refuse_rather_than_pretend() { - assert!(run_now("hypermnesia-extract").is_err()); - let job = Job { - label: "x".into(), source: PathBuf::new(), schedule: Schedule::None, + fn broken(label: &str, schedule: Schedule) -> Job { + Job { + label: label.into(), source: PathBuf::new(), schedule, program: vec![], log: None, log_state: LogState::NotConfigured, last_exit: None, - installed: None, runs: None, pid: None, trigger: TriggerState::NotTracked, fault: None, - }; - assert!(set_schedule(&job, &Schedule::Every(3600)).is_err()); - assert!(install_self("hypermnesia-tray").is_err()); - assert!(uninstall_self("hypermnesia-tray").is_err()); + installed: None, runs: None, pid: None, trigger: TriggerState::NotTracked, + armed: None, fault: Some("could not be read".into()), + } + } + + #[test] + fn set_schedule_refuses_a_job_with_a_fault_without_touching_the_filesystem() { + assert!(set_schedule(&broken("x", Schedule::Every(3600)), &Schedule::Every(3600)).is_err()); + } + + #[test] + fn set_schedule_refuses_none_and_several_before_touching_anything() { + let mut j = broken("x", Schedule::None); + j.fault = None; + assert!(set_schedule(&j, &Schedule::None).is_err()); + assert!(set_schedule(&j, &Schedule::Several(vec![Cal::default()])).is_err()); + } + + /// A unit name is a path component and a shell argument both; anything that could escape + /// either is refused before a single `Command` is built. + #[test] + fn valid_stem_rejects_traversal_and_shell_metacharacters() { + assert!(valid_stem("hypermnesia-extract").is_ok()); + assert!(valid_stem("").is_err()); + assert!(valid_stem("../etc/passwd").is_err()); + assert!(valid_stem("a/b").is_err()); + assert!(valid_stem("-rf").is_err()); + assert!(valid_stem("a b").is_err()); + assert!(valid_stem("a;rm -rf ~").is_err()); + } + + /// The inverse of `parse_oncalendar`: every field it can produce round-trips, and an omitted + /// field renders as `*`, never a zero -- the same bug `parse_oncalendar`'s own tests guard + /// against, from the writing side this time. + #[test] + fn render_oncalendar_round_trips_through_parse_oncalendar() { + let weekly = Cal { hour: Some(6), minute: Some(10), day: None, weekday: Some(1), + month: None }; + assert_eq!(render_oncalendar(&weekly), "Mon *-*-* 06:10:00"); + assert_eq!(parse_oncalendar(&render_oncalendar(&weekly)), Some(weekly)); + + let monthly = Cal { day: Some(1), hour: Some(3), minute: Some(0), weekday: None, + month: None }; + assert_eq!(render_oncalendar(&monthly), "*-*-01 03:00:00"); + assert_eq!(parse_oncalendar(&render_oncalendar(&monthly)), Some(monthly)); + + let hourly = Cal { minute: Some(15), ..Cal::default() }; + assert_eq!(render_oncalendar(&hourly), "*-*-* *:15:00"); + assert_eq!(parse_oncalendar(&render_oncalendar(&hourly)), Some(hourly)); + } + + #[test] + fn rewrite_timer_replaces_an_existing_oncalendar_rather_than_doubling_it() { + let unit = "[Unit]\nDescription=x\n\n[Timer]\nOnCalendar=*-*-* 05:00:00\nPersistent=true\n\ + \n[Install]\nWantedBy=timers.target\n"; + let sched = Schedule::at(6, 10, Some(1)); + let out = rewrite_timer(unit, &sched).expect("rewrites"); + assert_eq!(out.matches("OnCalendar=").count(), 1, "{out}"); + assert!(out.contains("OnCalendar=Mon *-*-* 06:10:00"), "{out}"); + assert!(out.contains("Persistent=true"), "unrelated keys survive: {out}"); + assert!(out.contains("WantedBy=timers.target"), "sections after [Timer] survive: {out}"); + } + + #[test] + fn rewrite_timer_switches_an_interval_schedule_and_sets_onbootsec_too() { + let unit = "[Timer]\nOnCalendar=*-*-* 05:00:00\n\n[Install]\nWantedBy=timers.target\n"; + let out = rewrite_timer(unit, &Schedule::Every(14_400)).expect("rewrites"); + assert!(!out.contains("OnCalendar="), "{out}"); + assert!(out.contains("OnUnitActiveSec=14400"), "{out}"); + assert!(out.contains("OnBootSec=14400"), "so an interval timer also fires after a boot: {out}"); + } + + #[test] + fn rewrite_timer_refuses_without_a_timer_section() { + assert!(rewrite_timer("[Unit]\nDescription=x\n", &Schedule::Every(3600)).is_err()); + } + + #[test] + fn rewrite_timer_refuses_none_and_several() { + assert!(rewrite_timer("[Timer]\n", &Schedule::None).is_err()); + assert!(rewrite_timer("[Timer]\n", &Schedule::Several(vec![Cal::default()])).is_err()); + } + + /// `Restart=on-failure`, not `always`: an unconditional restart makes Quit a button that + /// lies, the same trap `launchd.rs`'s `KeepAlive` comment documents for the plist side. + #[test] + fn tray_unit_restarts_on_failure_only_and_points_at_the_binary() { + let unit = tray_unit("/usr/local/bin/hypermnesia", "/home/you/.hypermnesia/logs/tray.log"); + assert!(unit.contains("Restart=on-failure"), "{unit}"); + assert!(!unit.contains("Restart=always"), "{unit}"); + assert!(unit.contains("ExecStart=/usr/local/bin/hypermnesia"), "{unit}"); + assert!(unit.contains("WantedBy=default.target"), "{unit}"); } + /// A unit that was never installed is an honest "nothing to report", not a fault -- the same + /// rule the read-only version of this function was already written to. #[test] - fn autostart_fault_is_none_not_a_permanent_warning() { - // Not implemented is not the same claim as "checked, and it's fine" -- but a permanent - // warning drawn on every menu rebuild is a warning nobody reads, so this returns the - // same `None` a passing check would. - assert_eq!(autostart_fault("hypermnesia-tray"), None); + fn autostart_fault_is_none_when_nothing_was_ever_installed() { + // A label no real installation would use, so the read simply finds no file. + assert_eq!(autostart_fault("hypermnesia-tray-selftest-does-not-exist"), None); } } diff --git a/console/src/jobs/unsupported.rs b/console/src/jobs/unsupported.rs index 10beb25..d9b6653 100644 --- a/console/src/jobs/unsupported.rs +++ b/console/src/jobs/unsupported.rs @@ -18,6 +18,12 @@ pub fn set_schedule(_job: &Job, _sched: &Schedule) -> Result { Err("scheduled jobs are not supported on this platform".into()) } +pub const SUPPORTS_ENABLE: bool = false; + +pub fn set_enabled(_job: &Job, _enabled: bool) -> Result { + Err("scheduled jobs are not supported on this platform".into()) +} + pub fn install_self(_label: &str) -> Result { Err("autostart is not supported on this platform".into()) } diff --git a/console/src/view.rs b/console/src/view.rs index d513d3d..1a9eeed 100644 --- a/console/src/view.rs +++ b/console/src/view.rs @@ -121,6 +121,7 @@ mod tests { log_state: LogState::Written(SystemTime::now() - Duration::from_secs(300)), runs: Some(9), trigger: jobs::TriggerState::NotTracked, installed: Some(SystemTime::now() - Duration::from_secs(90_000)), fault: None, + armed: None, } } diff --git a/docs/DIAGNOSTICS.md b/docs/DIAGNOSTICS.md index 77f12bf..163c06c 100644 --- a/docs/DIAGNOSTICS.md +++ b/docs/DIAGNOSTICS.md @@ -129,7 +129,12 @@ have. What it says, and what each line means: | `this shell` as a source | set here, but the jobs the service manager starts do not inherit it — the line below says what they use | | `the answer parsed but has no "memories"` | something answered, but it was not this query's result. Not a store full of zeros | | `systemctl --user is not reachable: ` (Linux) | no session bus, not "no jobs configured" — the two must never look alike | -| `... is not implemented on Linux yet` | `run`, `every`, `at`, `--install`, `--uninstall`: read-only for now on the systemd backend | +| `.bak already exists: an earlier edit did not finish` | a previous `set_schedule` was interrupted before it could clean up; compare the `.bak` with the live file and remove it before editing again | +| `the unit does not pass the check after the edit (...); the file was restored` (Linux) | `systemd-analyze --user verify` (or, failing that, a `daemon-reload` + `LoadState` check) refused the edit before anything was reloaded; nothing changed | +| `systemd reloaded the unit but reports a different schedule than was set (...) -- the file was restored` (Linux) | the read-back after a reload did not match what was asked; the console does not call an edit done on the strength of an exit code alone | +| ` ran and exited ` | `run` succeeded in the sense of starting; the program itself failed. The exit code is read back from `ExecMainStatus`/`InvocationID` after the request, not assumed from `systemctl start`'s own success | +| ` is masked -- unmask it first: systemctl --user unmask ` (Linux) | `enable`/`disable` refuses a masked unit rather than fighting the mask | +| `autostart starts , not this binary ()` | the installed autostart unit points at a binary that has moved, been rebuilt elsewhere, or been deleted — it will fail at every login with nothing on screen to say so until this line does | Schedules are read the way each backend means them. On macOS an omitted `StartCalendarInterval` field is a wildcard, so `{Minute: 15}` is "hourly at :15" and `{Day: 1, Hour: 3}` is "monthly day 1 diff --git a/docs/INSTALL.md b/docs/INSTALL.md index bea8ba3..9e06d1b 100644 --- a/docs/INSTALL.md +++ b/docs/INSTALL.md @@ -241,7 +241,7 @@ and what the tunables are set to. cd console && cargo build --release ./target/release/hypermnesia-setup # the walkthrough, including the connection ./target/release/hypermnesia-stats # the numbers, once -./target/release/hypermnesia --install # the menu-bar tray, at login (macOS) +./target/release/hypermnesia --install # the tray, at login (launchd on macOS, a systemd user unit on Linux) ``` The four command-line tools build the same way everywhere; the data layer has no dependencies at @@ -259,11 +259,20 @@ sudo dnf install gtk3-devel libayatana-appindicator-gtk3-devel libxdo-devel # F cargo build --release --features tray ``` -**Linux, what works today:** the tray icon, the store's numbers, and reading systemd user units -(`~/.config/systemd/user/*.timer`) -- schedule, last exit code, whether it is running, and when it -last fired, read straight from `systemctl --user show` with no date-parsing crate. Editing a -schedule, running a job on demand, and installing the tray into autostart are not implemented on -Linux yet; `hypermnesia-jobs` and the tray say so rather than doing nothing silently. See +**Linux backend:** systemd user units (`~/.config/systemd/user/*.{timer,service}`), read through +`systemctl --user show` with no date-parsing crate. Both platforms now read, run, reschedule and +arm/disarm jobs, and install the tray itself into autostart: + +| | macOS (launchd) | Linux (systemd) | +|---|---|---| +| run a job now | `launchctl kickstart -k` | `systemctl --user start --no-block` | +| change a schedule | edit the plist, `plutil -lint`, bootout/bootstrap | edit the unit, `systemd-analyze --user verify`, daemon-reload + restart | +| arm / disarm | n/a — a loaded job is armed, full stop | `systemctl --user enable\|disable --now` (`hypermnesia-jobs enable\|disable`, and the tray's **Timers** submenu) | +| install the tray | `~/Library/LaunchAgents/com.hypermnesia.tray.plist`, `RunAtLoad` | `~/.config/systemd/user/hypermnesia-tray.service`, `WantedBy=default.target`, enabled but not started (avoids running a second tray beside the one already open) | + +Every write is backed up, validated before it is reloaded, and read back afterward to confirm what +actually landed rather than trusting the exit code of the command that asked for it; any failure +along the way restores the backup and says, in words, whether the restore itself worked. See [the console message table](DIAGNOSTICS.md#what-the-console-tells-you-about-itself) for the exact Linux-side lines, and `HM_JOB_PREFIX` below for how job names differ (`hypermnesia-extract.timer`, not `com.hypermnesia.extract.plist`).