Skip to content

Remote Brain writer authority is never published in production; failover stops at Unavailable #2710

Description

@ScriptedAlchemy

Problem

The shipped daemon has no production path that publishes a Remote Brain writer authority. Because of that, tracedecay remote failover (and any replay that needs the current writer) cannot get past the writer lookup on a real daemon.

RemoteSqliteStorageV1::publish_authority (crates/tracedecay-rusqlite-runtime/src/remote/mod.rs, around line 231) is the only code that INSERTs into the RemoteNode remote_authorities table. On master 9b79b48 it has no non-test caller:

$ rg -n "publish_authority" crates --type rust | rg -v "fn publish_authority"
crates/tracedecay-rusqlite-runtime/src/remote/recovery_authority.rs:230:        self.publish_authority(&authority, frontier_sequence)   # a different fn: RemoteRecoverySqliteAuthorityV1::publish_authority, which writes remote_recovery_authorities
crates/tracedecay-rusqlite-runtime/src/remote/tests.rs:660/709/735/1298                  # tests only

Every production reader expects that row to exist:

  • RemoteSqliteStorageV1::recovery_writer / recovery_writer_for_lineage (remote/recovery_authority.rs, lines 268 and 290), which failover uses through DaemonRemoteRecoveryPhysicalEffectsV1::current_authority / promote
  • remote/replay_authority.rs, lines 30 and 92
  • remote/recovery_authority/journal.rs::promote_primary_writer_in, which only UPDATEs an existing row

Consequence

On a real daemon, /failover → RemoteRecoverySqliteAuthorityV1::promote → ensure_authority_seeded → effects.current_authority → storage.recovery_writer finds no row and returns RemoteRecoveryOperationErrorV1::Unavailable. The rest of the promotion (writer-fence read, fence install, sink publication) cannot run from the shipped CLI. The only coverage is in-process tests that call publish_authority directly, for example daemon::store_runtime_tests::remote_failover_reports_corrupt_fence_text_naming_cancel_as_corruption_not_cancellation, added in the root-cause-remote-recovery lane.

Expected

Either a real production journey (enrollment, capture, or an operator authority-seeding operation) publishes the current writer authority with typed validation, or failover and replay report a typed "no writer authority published" state instead of generic Unavailable. Any fix needs a production-daemon journey test.

Found while typing the remote-recovery fence errors (lane root-cause-remote-recovery).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions