Skip to content

docs(floxhub-onprem): scaffold the on-prem section - #103

Open
imkarrer wants to merge 8 commits into
docs/floxhub-onpremfrom
isaac/ent-151-onprem-skeleton
Open

imkarrer wants to merge 8 commits into
docs/floxhub-onpremfrom
isaac/ent-151-onprem-skeleton

Conversation

@imkarrer

@imkarrer imkarrer commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Scaffolds a FloxHub on-prem section in the sidebar — 33 pages across
installation, reference architectures, administration, and upgrades — and
writes the first seven.

What to review: the section structure (docs.json), and the seven pages
that carry real content:

Page Issue
floxhub-onprem/intro — what the offering is, how it differs from hosted FloxHub, and what the operator is responsible for ENT-338
.../administration/monitoring/health-check — every service's health endpoint, how to probe it, and what a 200 does and does not prove ENT-350
.../administration/users/identity — attaching an identity provider, with working connector examples ENT-361
.../administration/users/overview — how a person becomes a user, and which operator controls exist ENT-360
.../install/requirements — the hard prerequisites: host, database, URL and TLS, network, identity provider ENT-340
.../reference-architectures/overview — the supported topology, what runs where, and what single-host does not give you ENT-344
.../administration/object-storage — attaching an S3-compatible store, signing, and consumer trust ENT-352

The remaining 26 are placeholders carrying an under-construction banner and a
short statement of planned scope. Each commit cites the sub-issue it closes.

Tracked by ENT-151.

Targets docs/floxhub-onprem, the integration branch for this section, not
main. Merging here publishes nothing: the section reaches flox.dev/docs
only when that branch merges to main, once on-prem is ready to be public.

Why a skeleton first

The section is thirty-odd cross-linked pages accumulating on an integration
branch. Landing them one at a time without an agreed tree means every page
either links to nothing or invents a path the next page has to match. Putting the structure in first makes the
page naming reviewable while it is still cheap to change, lets each
subsequent page land as its own small PR, and makes the two landing pages
(introduction and reference architectures) writable against something real
instead of guesswork.

The placeholders are deliberate rather than empty: each states what the page
will cover, so the scope of the section is reviewable now.

Structure and where it comes from

The tree mirrors the GitLab documentation's split between
installation and
administration, which is the
structure ENT-151 asks for. Page-for-page, each placeholder corresponds to a
sub-issue of ENT-151.

The group sits directly after Imageless Kubernetes, the site's other
self-hosted-product section, rather than under Customer — that section is
three standalone pages, not a product tree.

Health check sits under Administer → Monitoring, matching GitLab's
/administration/monitoring/health_check/. The Configure page links to it.

How the health check page was written

Every endpoint, port, and path was read out of the service source rather than
inferred, including the mount prefixes that determine the full paths. Notable
findings that shaped the page:

  • Only two endpoints do real work. catalog-server's /status/healthcheck
    runs a show, a search, and a resolve against the database and reports
    per-operation timings; build-coordinator's /health returns 503 naming
    any background loop that has exited. Everything else returns a fixed
    literal.
  • The /ready paths on accounts, factory, and build-coordinator check
    nothing. They are documented as placeholders rather than as readiness
    signals, because treating them as readiness is the mistake this page exists
    to prevent.
  • web-bff's /api/health/details reports no dependencies on-prem — its
    only dependency check targets a service on-prem deployments do not use — so
    it is documented as a second liveness probe.
  • Through the front door only /web-bff/api/health/status and the two
    discovery documents answer without a credential. Everything else returns
    401, which is a useful signal about the front door and no signal at all
    about the service behind it.

The Limits section is as much of the page as the endpoint list, because
the failure mode in practice is reading a 200 as "the deployment works".

How the Users pages were written

Sourced from the deployment's own connector guide and the accounts service,
with the endpoint and provisioning behavior read from the code rather than
inferred.

Three findings drove how the pages are organized:

  • The handle is derived, not chosen. It comes from the
    preferred_username claim (falling back to name), sanitized to the handle
    grammar. So a provider sending neither cannot have its users provisioned, and
    two usernames that sanitize to the same handle block the second person's
    sign-in with no self-service way out. Both are documented as pre-rollout
    checks, because they are cheap to check and expensive to hit during a
    rollout.
  • The external URL is load-bearing. It is baked into every token the
    deployment issues and advertised as the broker's issuer, so it has to be
    correct before the first sign-in. That is stated before the connector
    instructions rather than after.
  • Several expected controls do not exist — no invite flow, no user
    deactivation, no organization deletion, no directory group links, no
    administrative user listing. These are recorded explicitly so an operator
    plans around them instead of hunting for a setting. Revocation belongs at the
    identity provider, and does not reach tokens already issued.

Organization and membership management is named but deliberately left to the
follow-up scoped to it, matching how ENT-360 is written.

How the requirements and architecture pages were written

Both are grounded in the base environment and a working site configuration
rather than in intent.

  • TLS is not the deployment's. The front door listens on plain HTTP and has
    no certificate configuration at all, so termination belongs to the operator's
    own edge. Stated as a warning, because assuming otherwise means discovering
    it at cutover.
  • The Nix feature requirement is retained, not install-time. Catalog
    population and every periodic refresh evaluate flake outputs, so the features
    cannot be enabled for installation and turned off afterwards.
  • Host platform and served systems are different things. One x86_64 Linux
    host serves a catalog for macOS and Linux clients on both architectures, and
    the two are configured separately.
  • Sizing is absent on purpose. The test harness runs in a small VM against a
    fixture catalog, which would badly mislead as a recommendation. Both pages say
    the figures are being measured instead of guessing.

The architecture page states that one topology is documented and that it offers
no redundancy, so availability rests on the host and database the operator
supplies plus their ability to restore. Separated, redundant, and
managed-component arrangements are named as undocumented rather than left
ambiguous.

The diagram is Mermaid, rendered with mermaid-cli to confirm it parses and
that every node label survives — not eyeballed.

How the object storage page was written

Grounded in the exercised reference-architecture deployment in
flox/playground, not only in the base environment — the base leaves object
storage to the site, so reading it alone suggests the capability is absent when
it is the site's to supply.

Two behaviors shape the page because they differ from the hosted service:

  • FloxHub mints no object-store credentials. On a nix-copy store the
    publish-info path returns the ingress URI and no credential, where the hosted
    publisher type returns both. So whoever publishes authenticates to the
    bucket themselves with standard AWS environment variables, and there is no
    per-user scoping inside FloxHub. Recorded as a warning, since it is a
    credential-distribution problem an operator inherits.
  • Signing is load-bearing, not optional. Nothing substitutes an unsigned
    artifact, so the key and the trust clients place in its name are part of
    making a store work at all. The key name is recorded in every signature, so
    the page says to pick a stable one.

The verification section insists on two machines with independent stores. A
single-machine publish and install succeeds with no object store configured,
because the local store already holds the build — so the obvious check proves
nothing.

The per-catalog store API is documented in place of the tool named by the 422
error an unconfigured catalog returns, which on-prem operators do not have
(HUB-310). The page carries a
short note to that effect, which can go once that is fixed.

Issues raised while writing these pages

Reading the services closely enough to document them surfaced two defects. Both
are filed and neither blocks this PR.

  • HUB-310 — publishing to a
    catalog with no store configured returns a 422 instructing the operator to
    run catalog-util, which does not ship in the on-prem environment and appears
    nowhere outside that error string. An operator meets this the first time they
    publish. The object-storage page documents the store-config API instead.
  • DEV-372 — the published
    Signing keys page named the flox publish flag as --signing-key rather than
    --signing-private-key. Already fixed by
    #104. It matters here because the
    object-storage page deliberately defers the cross-platform trusted-key
    mechanics to that page rather than restating them, so it wants that page
    correct.

Also noted but not filed: RUNBOOK.md in floxhub contradicts the
installation-requirements page in places and is the stale side — it still opens
"Status: PoC - Not for production use" and describes the superseded
per-customer environment-copy model, while other parts of it are current.

Verification
  • vale — clean across all 33 pages. loopback, substituter, and unconfigured added alongside the others.
  • mermaid-cli — the architecture diagram renders, with all node labels present.
    liveness, Entra, Okta, lowercased, misconfigured, rollout, and
    loopback added to the project vocabulary; two phrasings reworded rather than
    adding a word for them.
  • mint broken-links — no broken links introduced. The three reported are
    pre-existing changelog/rss.xml references.
  • mint dev — every new page returns 200 locally, confirming the MDX and
    the components parse.
  • llms.txt regenerated with scripts/generate-llms-txt.sh, as
    check-llms-txt requires.

Add a `FloxHub on-prem` navigation group covering installation, reference
architectures, administration, and upgrades, with a page per planned topic.

Two pages carry real content:

- `floxhub-onprem/intro` — what the offering is, how it differs from hosted
  FloxHub, and which responsibilities belong to the operator.
- `floxhub-onprem/administration/monitoring/health-check` — the health
  endpoint each service exposes, how to probe it from the host and through
  the front door, and what a `200` does and does not prove.

The rest are placeholders marked under construction. They exist so the
navigation tree and cross-links are complete while the remaining pages are
written, rather than landing thirty pages of links to nowhere.

`liveness` joins the Vale vocabulary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mintlify

mintlify Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
flox 🟢 Ready View Preview Sep 28, 2026, 2:34 PM

💡 Tip: Enable Automations to automatically generate PRs for you.

@imkarrer
imkarrer changed the base branch from main to docs/floxhub-onprem September 28, 2026 17:21
imkarrer and others added 2 commits September 28, 2026 16:36
Replace the placeholder with the connector configuration an operator needs:
the three-step flow for attaching a provider, minimal working examples for
Okta, Entra ID, AD FS, and LDAP, and why SAML is not offered.

Two facts get prominence because getting them wrong is expensive. The external
URL is baked into every token the deployment issues, so it has to be right
before the first sign-in rather than after. And the deployment derives a user's
handle from the `preferred_username` claim, so a provider that sends neither
that nor `name` produces an authenticated identity that cannot be provisioned.

Also states what authenticating does not buy: directory group membership does
not map to organizations or roles, so a group grant yields a working sign-in
and nothing more.

`Entra`, `Okta`, `lowercased`, `misconfigured`, and `rollout` join the Vale
vocabulary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replace the placeholder with how a person becomes a user of a deployment:
provisioning as a side effect of a first successful sign-in, how the handle is
derived and what grammar it must satisfy, and the single namespace users and
organizations share.

The failure modes get a section of their own. A handle collision has no
self-service resolution, because the handle is derived rather than chosen, so
two directory usernames that sanitize to the same handle block the second
person's sign-in until the provider sends something different. That is cheap to
check before a rollout and expensive to discover during one.

Records the operator controls that do not exist — no invite flow, no
deactivation, no organization deletion, no directory group links, no
administrative user listing — so they are planned around rather than searched
for. Revocation belongs at the identity provider, and does not reach tokens
already issued.

Organization and membership management is named but left to the follow-up that
covers it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
imkarrer and others added 2 commits September 29, 2026 08:37
Replace the placeholder with the hard prerequisites: host platform, the Nix
features catalog population needs, the database you supply, the external URL,
network egress, the systems your catalog serves, and the identity provider.

Three of these are easy to discover too late. The external URL is baked into
every token the deployment issues, so it is a decision made before installing
rather than after. The front door speaks plain HTTP and holds no certificate,
so TLS termination belongs to your own edge. And the Nix feature requirement
is retained rather than install-time — every periodic catalog refresh needs it.

Sizing is deliberately absent. Stating that it is being measured is more useful
than guidance this page cannot yet support, and the follow-up scoped to it will
fill it in.

`loopback` joins the Vale vocabulary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replace the placeholder with the overview: a diagram of where the components
sit, a table of what each one is responsible for, and the boundary between what
runs on the host and what you supply.

States plainly that one topology is documented — a single host — so the
question is whether it fits rather than which to choose. Says what it does not
give you: no component runs redundantly and an upgrade stops services, so
availability rests on the host and database you supply and on your ability to
restore from backup.

Separated, redundant, and managed-component arrangements are named as
undocumented rather than left for a reader to assume either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replace the placeholder with the per-catalog store configuration: why a new
deployment publishes metadata only, the store types and which one on-prem uses,
the API call that attaches a store, and the difference between the ingress URI
publishing writes through and the egress URL installs read from.

Two points get emphasis because they diverge from the hosted service. FloxHub
hands a publishing client the bucket URI and nothing else, so whoever publishes
must already hold write credentials and there is no per-user scoping inside
FloxHub. And consumers substitute nothing unsigned, so a signing key and the
trust that clients place in its name are part of making a store useful rather
than an afterthought.

The verification section insists on two machines. A single-machine publish and
install succeeds with no object store at all, because the local store already
holds the build, so it proves nothing about the round trip.

`substituter` and `unconfigured` join the Vale vocabulary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Comparing the new pages against `customer/signing-keys` and the Flox install
instructions turned up three things worth correcting.

The object-storage page reimplemented, thinly, what `customer/signing-keys`
already covers properly: where trusted keys live for standalone Nix, NixOS,
nix-darwin, and home-manager, restarting the daemon on each platform, and
verifying a key took effect. It now defers to that page and keeps only what is
specific to running your own store. It also states that the private key path
must be absolute, which that page warns about and this one had omitted.

The requirements page gained the trusted-key prerequisite for the host. It is
usually already satisfied, because the Flox installer configures those keys — so
it only bites a host where Flox arrived through Nix, and then it presents as the
deployment's own packages failing to install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
staging — dbcbf034 Deployed Sep 29, 2026 by mintlify[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant