Skip to content
OleBindersPublic

About

A self hosted library of information. Made for AI-Human collaboration. Fully vibe coded.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Urd

A knowledge library that's both human-browsable and AI-readable via MCP. Named after Urðr, the Norn who tends the well of fate at the root of Yggdrasil, keeper of what has come to pass. See the full project definition (goals, architecture rationale, roadmap) in project-definition.md — this README only covers running the code.

This is the Phase 0–3 build: the format, schema, and services are in place and runnable end-to-end. search is hybrid full-text + semantic (local embeddings via fastembed). The web UI is a Wikipedia-style browse/read experience: entry pages get a table of contents, inline citation markers rendered as numbered references, live search-as-you-type, category browsing, and per-entry git history with rendered word-level diffs. Quality control is in place too: a draft→reviewed toggle (human-only — not exposed to AI agents), a library-wide "recent changes" view for reviewing AI-written edits, and an audit (MCP tool, CLI, and web page) that flags stale entries, dead citation links, and pairs of suspiciously similar, non-cross-linked entries worth a human look.

Structure

content/entries/       markdown+frontmatter entries — the source of truth,
                        its own internal git repo (see below)
content/SCHEMA.md       what the frontmatter fields mean
db/init.sql             Postgres schema (entries table, pgvector extension)
services/mcp_server/    FastMCP server: library_map resource, search/read/write tools
services/web_api/       FastAPI app: browse/search/categories/history the library in a browser

content/entries/ is its own git repo, separate from this app-code repo (and gitignored here) — the MCP server initializes it on first run and write() commits to it directly, so every AI-written entry is versioned. It has no remote by default; it's local history, durable on disk via the bind mount but not pushed anywhere unless you add a remote yourself.

Running it

cp .env.example .env        # edit POSTGRES_PASSWORD before deploying anywhere real
docker compose up --build -d db
docker compose run --rm mcp_server python -m app.sync   # loads content/entries/*.md into Postgres
docker compose up --build -d mcp_server web_api

Also verified working with podman-compose in place of docker compose (rootless Podman + SELinux enforcing, e.g. Fedora Atomic/Bazzite) — no separate instructions needed, the compose file already carries what that needs (fully-qualified image names, :z-labeled bind mounts).

.env must exist with a real POSTGRES_PASSWORD before the first up — compose refuses to start otherwise (POSTGRES_PASSWORD must be set), on purpose: Postgres only applies that variable on first-ever init of an empty data directory, so if you ever changed it in .env after already running up once, that change has no effect on the existing volume. To pick up a changed password (or recover from ever having started once without a real .env), wipe the volume and let it reinitialize: docker compose down -v (or podman-compose down -v) — this only deletes the rebuildable Postgres search index (db_data), not content/entries/, which is bind-mounted and untouched; re-run the sync command afterward to rebuild the index.

Point an MCP client at the server's URL to get library_map, search, read, and write. write() commits to the content repo and updates the search index in the same call; after hand-editing a markdown file directly, re-run the sync command above to pick it up (hand edits aren't git-committed automatically).

A Claude Code session working in this repo gets AI-usage guidance for Urd's own MCP tools automatically — see .claude/skills/urd-usage/.

Authentication

mcp_server requires a bearer token by default (MCP_AUTH_MODE=token, the default in .env.example). Generate one with openssl rand -hex 32 and set it as MCP_AUTH_TOKEN in .env before bringing the stack up — both mcp_server (which enforces it) and web_api (whose "Mark reviewed" button and /audit page call mcp_server's internal routes) need the same value. Point an MCP client at the server with an Authorization: Bearer <token> header; the same header is required on the plain-HTTP /internal/* routes.

Plaintext risk: in token mode, the bearer token travels in plaintext on the wire unless mcp_server sits behind a TLS-terminating reverse proxy — anyone who can observe the network can read the token. Set MCP_BEHIND_TLS_PROXY=true once TLS is actually in place in front of it to silence the startup warning; leaving it false is the honest default for a plain local/LAN deployment.

Set MCP_AUTH_MODE=open to disable auth entirely (e.g. for a fully trusted local network) — mcp_server logs a loud startup warning in this mode, and every response carries an X-Urd-Auth-Mode: open header so you can verify from outside which mode is actually active (curl -i any route and check the header). Not recommended outside a trusted network: in open mode anyone who can reach the port can read and write the library, including via write().

Connecting a client that can't set a bearer token

Some MCP clients only support spawning a local stdio server (a command + args), with no field for custom headers on a remote HTTP connection — e.g. Odysseus. For these, bridge through mcp-remote, a small proxy that runs as that local stdio process and injects the header on your behalf when it connects out to mcp_server:

  • command: npx
  • args: ["-y", "mcp-remote", "https://<host>:8000/mcp", "--header", "Authorization:${AUTH_HEADER}"]
  • env: AUTH_HEADER=Bearer <your-mcp-auth-token>

Keep the header's colon tight against ${AUTH_HEADER} (no space before it) — some clients mangle spaces inside args when invoking npx, so the space belongs only inside the env var's value, not the args string itself. If the client has no per-server env field, the header can go directly in args instead ("--header", "Authorization: Bearer <token>") — functional, just less clean since the token then sits in the visible command rather than a separate field.

If mcp_server isn't behind TLS (this project's own default — see "Running behind a reverse proxy" below), point mcp-remote at http:// instead of https:// and add --allow-http to args: mcp-remote refuses plain-HTTP URLs to a non-localhost host by default and exits immediately, which looks like an instant "Connection Closed" from the client's side rather than a clear error. This is the same plaintext-token tradeoff as always — reasonable on a trusted LAN/VM network, not once mcp_server is reachable beyond one.

If the client can only reach mcp_server over a network you already trust (e.g. a private LAN, no public exposure), running that instance with MCP_AUTH_MODE=open instead is the simpler option — see above for what that trades away.

Running behind a reverse proxy

Neither mcp_server nor web_api terminates TLS — both speak plain HTTP only, on their published ports (8000 and 8080). That's fine for a local/LAN deployment, but for anything reachable over the open internet you should put a TLS-terminating reverse proxy in front of them. This project doesn't assume any particular proxy — pick whichever you already run.

The routes each service needs upstream:

  • web_api:8080 — the entire web UI, route all traffic here for your web hostname.
  • mcp_server:8000 — the MCP protocol itself lives at /mcp (FastMCP's streamable-http transport default), plus two /internal/* routes (/internal/entries/{id}/review, /internal/audit) that web_api calls server-to-server. If your proxy supports path-based routing you can restrict external access to /mcp only and leave /internal/* unreachable from outside, but it's not required — auth (see "Authentication" above) already covers the whole app, /internal/* included.

Traefik: see docker-compose.override.yml.example in the repo root. Copy it to docker-compose.override.yml (Compose auto-merges override files with no extra flags needed) and edit the Host() rules and PROXY_NETWORK_NAME for your setup. Note its !reset trick — which drops the direct host-port publish so the services are reachable only via the proxy, not both through the proxy and in plaintext on the raw port at the same time — needs Docker Compose v2.24+; support under podman-compose varies by version (tested working on podman-compose 1.6.0), so if yours rejects the tag, the file's header comments explain the workaround (comment out ports: directly in your own copy of docker-compose.yml instead).

Caddy: join Caddy's container to the same external proxy network (the networks: block from docker-compose.override.yml.example, without its Traefik labels:), then:

urd.example.com {
    reverse_proxy web_api:8080
}
mcp.example.com {
    reverse_proxy mcp_server:8000
}

nginx: same network prerequisite, then an equivalent proxy_pass in your server blocks:

server {
    server_name urd.example.com;
    location / {
        proxy_pass http://web_api:8080;
    }
}
server {
    server_name mcp.example.com;
    location / {
        proxy_pass http://mcp_server:8000;
    }
}

No proxy at all (today's default, still fully supported): connect directly to http://<host>:8000 and http://<host>:8080 over plain HTTP — e.g. a trusted LAN, or over an SSH tunnel. See the plaintext-token warning in "Authentication" above: with no proxy there's usually no TLS either, so the bearer token travels unencrypted — fine on a network you trust, not otherwise.

Sharing

web_api has no authentication at all by default — anyone who can reach port 8080 sees the whole library, drafts included. WEB_SHARING_MODE (default off) optionally makes specific entries, or the whole library, readable by an unauthenticated web visitor. It's a separate knob from MCP_AUTH_MODE above: that one gates the AI-facing MCP protocol, WEB_SHARING_MODE gates the human-facing web UI. Set either, both, or neither, independently.

  • off (default) — today's behavior, unchanged: the web UI shows every entry regardless of its visibility field.
  • entries — only entries with visibility: public in their frontmatter (mirrored into the entries.visibility Postgres column) are served: the recent-entries listing, search results, and category pages simply omit private entries, and a direct request for a private entry's own page or its history/diff views 404s exactly like a nonexistent id, so its existence isn't leaked either way.
  • library — the whole read surface is open regardless of visibility, for a deployer who wants everything public. Behaves the same as off today (no filtering); kept as a distinct value so intent is explicit in .env rather than inferred from off.

Known gap: /recent-changes and /audit (the Phase 3 human-audit surfaces) are not filtered by WEB_SHARING_MODE — both can still surface a private entry's id/title (e.g. /audit's staleness and broken-citation checks run over every entry regardless of visibility). They're human-audit surfaces, not part of the public browse/search path, but if you're relying on entries mode to actually hide a private entry's existence, don't expose /recent-changes or /audit on the same public-facing vhost.

Every entry defaults to visibility: private (see content/SCHEMA.md). Making one public is a manual, human-only action — an entry page's "Make public" button (mirroring the existing "Mark reviewed" button next to it), which calls a plain-HTTP /internal/entries/{id}/visibility route on mcp_server, the same shape as the review-toggle route. There is no write() MCP parameter for this — an AI agent can never make an entry public, by design, same reasoning as status.

Current limitation: this web UI has no login/session concept yet (see project-definition.md's Non-goals — multi-user auth is Phase 5), so entries mode filters every request the same way; there's no "trusted, authenticated viewer who sees the full library" through this same UI. If you want an admin view of everything while the public sees only shared entries, the recommended approach today is proxy-level Basic Auth on an internal-only vhost/path pointed at a WEB_SHARING_MODE=off (or library) instance, with your public-facing vhost pointed at a WEB_SHARING_MODE=entries instance — see "Running behind a reverse proxy" above for the Traefik/ Caddy/nginx patterns to attach Basic Auth to.

Content backup

content/entries/ is its own git repo (see "Structure" above), local to this host's bind-mounted content/ directory — durable across container restarts, but not against loss of the host itself. CONTENT_BACKUP_REMOTE adds an optional off-box backup: a second git remote (named backup) that mcp_server pushes main to.

Unset by default — no backup, and mcp_server logs a startup warning saying so. Set it to a full git remote URL and it works with any git host reachable over HTTPS with embedded credentials: GitHub, GitLab, Gitea/Forgejo, Bitbucket, anything. Nothing here is hardcoded to a particular provider.

Recommended format: an HTTPS URL with an embedded, scoped token — not a full-account personal access token — restricted to push access on just the one backup repo. For example, against a self-hosted Gitea/Forgejo instance:

CONTENT_BACKUP_REMOTE=https://urd-backup:TOKEN@git.example.com/yourname/urd-content-backup.git

Create an empty repo for this on your git host first (urd-content-backup in the example above), and generate a token scoped to just that repo's push access rather than reusing a broader account token.

SSH remotes also work, as an advanced alternative, if you mount an SSH key and known_hosts into the mcp_server container yourself via docker-compose.override.yml — this isn't wired up by default since it needs an extra volume mount most deployments won't need.

Two push points:

  • Opportunistic — mcp_server pushes to backup right after every write() call and every "Mark reviewed" action. Best-effort: a failed push (remote unreachable, network blip) is logged as a warning and never blocks the write or fails the caller.
  • Timer, as a safety net — for anything the opportunistic push misses: a transient network failure, or entries added by hand-editing files and picked up by sync.py without ever going through write(). Templates for a systemd service + timer are in deploy/systemd/, running python -m app.backup (a thin CLI wrapper around the same push logic) every 15 minutes.

To install the timer:

  1. Copy both files to /etc/systemd/system/: sudo cp deploy/systemd/urd-content-backup.{service,timer} /etc/systemd/system/
  2. Edit WorkingDirectory in urd-content-backup.service to your actual repo path.
  3. sudo systemctl enable --now urd-content-backup.timer

Requires CONTENT_BACKUP_REMOTE set in .env (see .env.example) — it's already passed through to the mcp_server container in docker-compose.yml.

If your deployment was brought up with a non-default compose project name (e.g. podman-compose -p somename up or COMPOSE_PROJECT_NAME set when starting the stack), also set COMPOSE_PROJECT_NAME in .env so the timer's podman-compose exec call targets the right containers — see the EnvironmentFile= line in urd-content-backup.service.

Database backup

Postgres (the db service) is a rebuildable search index — everything in it can be reconstructed from content/entries/ by re-running python -m app.sync (see "Running it" above). The git-remote content backup in the previous section is what actually matters most for durability. This section's Postgres backup is a fast-restore convenience on top of that: skip the sync rebuild (which re-embeds every entry) by restoring a recent dump instead, if one happens to be around and fresh. If it's ever missing or stale, nothing is lost — just re-sync from content/entries/.

scripts/backup-pg.sh is a host-side script (not a container, not part of Compose) that runs pg_dump -Fc — Postgres's custom/compressed dump format, which supports selective and parallel restore via pg_restore, a better fit here than a plain .sql dump — against the already-running db service via podman-compose exec (or docker compose exec, see the comment at the top of the script).

Why a separate mechanism from the git-remote content backup (previous section) instead of reusing it? Postgres dumps are large, opaque, binary blobs produced on a recurring schedule forever — a poor fit for git (unbounded repo growth, no meaningful diffs). rsync/rclone are purpose-built for "copy files to remote storage, prune independently," which is what a dump rotation actually needs.

Configuration (set in .env — this script runs on the host, outside Compose, and sources .env directly itself, so these are not passed through docker-compose.yml):

If your deployment uses a non-default compose project name, also set COMPOSE_PROJECT_NAME in .env — same reasoning as "Content backup" above, and the script auto-exports whatever it finds in .env so this just works.

  • PG_BACKUP_DIR — where dumps land locally. Default ./backups.
  • PG_BACKUP_RETENTION_DAYS — local dumps older than this are pruned on every run. Default 14. Local prune only — doesn't touch anything already copied off-box.
  • PG_BACKUP_METHOD — none (default), rsync, or rclone.
    • none — dump stays local only; the script logs a warning to stderr every run as a reminder.
    • rsync — copies the dump to PG_BACKUP_REMOTE, e.g. PG_BACKUP_REMOTE=user@host:/path/ — any box you can SSH into.
    • rclone — copies the dump to PG_BACKUP_REMOTE, e.g. PG_BACKUP_REMOTE=remotename:path/ — any rclone-configured remote (S3, B2, etc), for users who want object storage instead of a box they manage.

Run it manually with ./scripts/backup-pg.sh, or install the systemd timer (same pattern as the content backup timer above):

  1. Copy both files to /etc/systemd/system/: sudo cp deploy/systemd/urd-pg-backup.{service,timer} /etc/systemd/system/
  2. Edit WorkingDirectory and ExecStart in urd-pg-backup.service to your actual repo path.
  3. sudo systemctl enable --now urd-pg-backup.timer

The timer runs daily. As with the content backup timer, systemd's restricted default PATH may not include podman-compose — see the comment in urd-pg-backup.service if the unit fails to find it.

Adding entries by hand

Copy content/entries/example-entry.md, follow the frontmatter fields described in content/SCHEMA.md, then run the sync command.

Embeddings

EMBEDDING_PROVIDER (docker-compose.yml, default fastembed) picks the provider — see services/mcp_server/app/embeddings.py for the full list (fastembed local, ollama local, voyage/openai cloud). Default is fully local: no API key, no network call at request time, no cost, content never leaves the machine — the model's weights are baked into the mcp_server image at build time. write() embeds synchronously on every create/update; sync backfills any entry missing one (e.g. hand-edited files). Switching providers is a config change, but always requires resizing the embedding column in db/init.sql to the new provider's dimension and re-running sync to re-embed everything — different models' vectors are never comparable to each other.

Known gaps (tracked as open design questions in the project doc)

  • Auth is a single shared bearer token, not per-user/per-client identity — fine for single-user/local, revisit before opening this up to multiple people.

About

A self hosted library of information. Made for AI-Human collaboration. Fully vibe coded.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages