A knowledge library that's both human-browsable and AI-readable via MCP.
Named after Urðr, the Norn who tends the well of fate at the root of
Yggdrasil, keeper of what has come to pass.
See the full project definition (goals, architecture rationale, roadmap) in
project-definition.md — this README only covers running the code.
This is the Phase 0–3 build: the format, schema, and services are in
place and runnable end-to-end. search is hybrid full-text + semantic
(local embeddings via fastembed). The web UI is a Wikipedia-style
browse/read experience: entry pages get a table of contents, inline
citation markers rendered as numbered references, live search-as-you-type,
category browsing, and per-entry git history with rendered word-level
diffs. Quality control is in place too: a draft→reviewed toggle
(human-only — not exposed to AI agents), a library-wide "recent changes"
view for reviewing AI-written edits, and an audit (MCP tool, CLI, and web
page) that flags stale entries, dead citation links, and pairs of
suspiciously similar, non-cross-linked entries worth a human look.
content/entries/ markdown+frontmatter entries — the source of truth,
its own internal git repo (see below)
content/SCHEMA.md what the frontmatter fields mean
db/init.sql Postgres schema (entries table, pgvector extension)
services/mcp_server/ FastMCP server: library_map resource, search/read/write tools
services/web_api/ FastAPI app: browse/search/categories/history the library in a browser
content/entries/ is its own git repo, separate from this app-code repo
(and gitignored here) — the MCP server initializes it on first run and
write() commits to it directly, so every AI-written entry is versioned.
It has no remote by default; it's local history, durable on disk via the
bind mount but not pushed anywhere unless you add a remote yourself.
cp .env.example .env # edit POSTGRES_PASSWORD before deploying anywhere real
docker compose up --build -d db
docker compose run --rm mcp_server python -m app.sync # loads content/entries/*.md into Postgres
docker compose up --build -d mcp_server web_apiAlso verified working with podman-compose in place of docker compose
(rootless Podman + SELinux enforcing, e.g. Fedora Atomic/Bazzite) — no
separate instructions needed, the compose file already carries what that
needs (fully-qualified image names, :z-labeled bind mounts).
.env must exist with a real POSTGRES_PASSWORD before the first up —
compose refuses to start otherwise (POSTGRES_PASSWORD must be set), on
purpose: Postgres only applies that variable on first-ever init of an empty
data directory, so if you ever changed it in .env after already running
up once, that change has no effect on the existing volume. To pick up a
changed password (or recover from ever having started once without a real
.env), wipe the volume and let it reinitialize: docker compose down -v
(or podman-compose down -v) — this only deletes the rebuildable Postgres
search index (db_data), not content/entries/, which is bind-mounted and
untouched; re-run the sync command afterward to rebuild the index.
- Web UI: http://localhost:8080
- MCP server (streamable-http transport): http://localhost:8000
Point an MCP client at the server's URL to get library_map, search,
read, and write. write() commits to the content repo and updates the
search index in the same call; after hand-editing a markdown file directly,
re-run the sync command above to pick it up (hand edits aren't
git-committed automatically).
A Claude Code session working in this repo gets AI-usage guidance for Urd's
own MCP tools automatically — see .claude/skills/urd-usage/.
mcp_server requires a bearer token by default (MCP_AUTH_MODE=token, the
default in .env.example). Generate one with openssl rand -hex 32 and set
it as MCP_AUTH_TOKEN in .env before bringing the stack up — both
mcp_server (which enforces it) and web_api (whose "Mark reviewed" button
and /audit page call mcp_server's internal routes) need the same value.
Point an MCP client at the server with an Authorization: Bearer <token>
header; the same header is required on the plain-HTTP /internal/* routes.
Plaintext risk: in token mode, the bearer token travels in plaintext
on the wire unless mcp_server sits behind a TLS-terminating reverse proxy
— anyone who can observe the network can read the token. Set
MCP_BEHIND_TLS_PROXY=true once TLS is actually in place in front of it to
silence the startup warning; leaving it false is the honest default for a
plain local/LAN deployment.
Set MCP_AUTH_MODE=open to disable auth entirely (e.g. for a fully trusted
local network) — mcp_server logs a loud startup warning in this mode, and
every response carries an X-Urd-Auth-Mode: open header so you can verify
from outside which mode is actually active (curl -i any route and check
the header). Not recommended outside a trusted network: in open mode
anyone who can reach the port can read and write the library, including
via write().
Some MCP clients only support spawning a local stdio server (a command +
args), with no field for custom headers on a remote HTTP connection — e.g.
Odysseus. For these,
bridge through mcp-remote, a
small proxy that runs as that local stdio process and injects the header on
your behalf when it connects out to mcp_server:
- command:
npx - args:
["-y", "mcp-remote", "https://<host>:8000/mcp", "--header", "Authorization:${AUTH_HEADER}"] - env:
AUTH_HEADER=Bearer <your-mcp-auth-token>
Keep the header's colon tight against ${AUTH_HEADER} (no space before
it) — some clients mangle spaces inside args when invoking npx, so the
space belongs only inside the env var's value, not the args string
itself. If the client has no per-server env field, the header can go
directly in args instead ("--header", "Authorization: Bearer <token>")
— functional, just less clean since the token then sits in the visible
command rather than a separate field.
If mcp_server isn't behind TLS (this project's own default — see
"Running behind a reverse proxy" below), point mcp-remote at http://
instead of https:// and add --allow-http to args: mcp-remote
refuses plain-HTTP URLs to a non-localhost host by default and exits
immediately, which looks like an instant "Connection Closed" from the
client's side rather than a clear error. This is the same plaintext-token
tradeoff as always — reasonable on a trusted LAN/VM network, not once
mcp_server is reachable beyond one.
If the client can only reach mcp_server over a network you already trust
(e.g. a private LAN, no public exposure), running that instance with
MCP_AUTH_MODE=open instead is the simpler option — see above for what
that trades away.
Neither mcp_server nor web_api terminates TLS — both speak plain HTTP
only, on their published ports (8000 and 8080). That's fine for a
local/LAN deployment, but for anything reachable over the open internet you
should put a TLS-terminating reverse proxy in front of them. This project
doesn't assume any particular proxy — pick whichever you already run.
The routes each service needs upstream:
web_api:8080— the entire web UI, route all traffic here for your web hostname.mcp_server:8000— the MCP protocol itself lives at/mcp(FastMCP's streamable-http transport default), plus two/internal/*routes (/internal/entries/{id}/review,/internal/audit) thatweb_apicalls server-to-server. If your proxy supports path-based routing you can restrict external access to/mcponly and leave/internal/*unreachable from outside, but it's not required — auth (see "Authentication" above) already covers the whole app,/internal/*included.
Traefik: see docker-compose.override.yml.example in the repo root.
Copy it to docker-compose.override.yml (Compose auto-merges override
files with no extra flags needed) and edit the Host() rules and
PROXY_NETWORK_NAME for your setup. Note its !reset trick — which drops
the direct host-port publish so the services are reachable only via the
proxy, not both through the proxy and in plaintext on the raw port at the
same time — needs Docker Compose v2.24+; support under podman-compose
varies by version (tested working on podman-compose 1.6.0), so if yours
rejects the tag, the file's header comments explain the workaround
(comment out ports: directly in your own copy of docker-compose.yml
instead).
Caddy: join Caddy's container to the same external proxy network
(the networks: block from docker-compose.override.yml.example, without
its Traefik labels:), then:
urd.example.com {
reverse_proxy web_api:8080
}
mcp.example.com {
reverse_proxy mcp_server:8000
}
nginx: same network prerequisite, then an equivalent proxy_pass in
your server blocks:
server {
server_name urd.example.com;
location / {
proxy_pass http://web_api:8080;
}
}
server {
server_name mcp.example.com;
location / {
proxy_pass http://mcp_server:8000;
}
}
No proxy at all (today's default, still fully supported): connect
directly to http://<host>:8000 and http://<host>:8080 over plain HTTP —
e.g. a trusted LAN, or over an SSH tunnel. See the plaintext-token warning
in "Authentication" above: with no proxy there's usually no TLS either, so
the bearer token travels unencrypted — fine on a network you trust, not
otherwise.
web_api has no authentication at all by default — anyone who can reach
port 8080 sees the whole library, drafts included. WEB_SHARING_MODE
(default off) optionally makes specific entries, or the whole library,
readable by an unauthenticated web visitor. It's a separate knob from
MCP_AUTH_MODE above: that one gates the AI-facing MCP protocol,
WEB_SHARING_MODE gates the human-facing web UI. Set either, both, or
neither, independently.
off(default) — today's behavior, unchanged: the web UI shows every entry regardless of itsvisibilityfield.entries— only entries withvisibility: publicin their frontmatter (mirrored into theentries.visibilityPostgres column) are served: the recent-entries listing, search results, and category pages simply omit private entries, and a direct request for a private entry's own page or its history/diff views 404s exactly like a nonexistent id, so its existence isn't leaked either way.library— the whole read surface is open regardless ofvisibility, for a deployer who wants everything public. Behaves the same asofftoday (no filtering); kept as a distinct value so intent is explicit in.envrather than inferred fromoff.
Known gap: /recent-changes and /audit (the Phase 3 human-audit
surfaces) are not filtered by WEB_SHARING_MODE — both can still surface
a private entry's id/title (e.g. /audit's staleness and broken-citation
checks run over every entry regardless of visibility). They're human-audit
surfaces, not part of the public browse/search path, but if you're relying
on entries mode to actually hide a private entry's existence, don't expose
/recent-changes or /audit on the same public-facing vhost.
Every entry defaults to visibility: private (see content/SCHEMA.md).
Making one public is a manual, human-only action — an entry page's "Make
public" button (mirroring the existing "Mark reviewed" button next to it),
which calls a plain-HTTP /internal/entries/{id}/visibility route on
mcp_server, the same shape as the review-toggle route. There is no
write() MCP parameter for this — an AI agent can never make an entry
public, by design, same reasoning as status.
Current limitation: this web UI has no login/session concept yet (see
project-definition.md's Non-goals — multi-user auth is Phase 5), so
entries mode filters every request the same way; there's no "trusted,
authenticated viewer who sees the full library" through this same UI. If
you want an admin view of everything while the public sees only shared
entries, the recommended approach today is proxy-level Basic Auth on an
internal-only vhost/path pointed at a WEB_SHARING_MODE=off (or library)
instance, with your public-facing vhost pointed at a WEB_SHARING_MODE=entries
instance — see "Running behind a reverse proxy" above for the Traefik/
Caddy/nginx patterns to attach Basic Auth to.
content/entries/ is its own git repo (see "Structure" above), local to
this host's bind-mounted content/ directory — durable across container
restarts, but not against loss of the host itself. CONTENT_BACKUP_REMOTE
adds an optional off-box backup: a second git remote (named backup) that
mcp_server pushes main to.
Unset by default — no backup, and mcp_server logs a startup warning
saying so. Set it to a full git remote URL and it works with any git host
reachable over HTTPS with embedded credentials: GitHub, GitLab,
Gitea/Forgejo, Bitbucket, anything. Nothing here is hardcoded to a
particular provider.
Recommended format: an HTTPS URL with an embedded, scoped token — not a full-account personal access token — restricted to push access on just the one backup repo. For example, against a self-hosted Gitea/Forgejo instance:
CONTENT_BACKUP_REMOTE=https://urd-backup:TOKEN@git.example.com/yourname/urd-content-backup.git
Create an empty repo for this on your git host first (urd-content-backup
in the example above), and generate a token scoped to just that repo's
push access rather than reusing a broader account token.
SSH remotes also work, as an advanced alternative, if you mount an SSH
key and known_hosts into the mcp_server container yourself via
docker-compose.override.yml — this isn't wired up by default since it
needs an extra volume mount most deployments won't need.
Two push points:
- Opportunistic —
mcp_serverpushes tobackupright after everywrite()call and every "Mark reviewed" action. Best-effort: a failed push (remote unreachable, network blip) is logged as a warning and never blocks the write or fails the caller. - Timer, as a safety net — for anything the opportunistic push misses:
a transient network failure, or entries added by hand-editing files and
picked up by
sync.pywithout ever going throughwrite(). Templates for a systemd service + timer are indeploy/systemd/, runningpython -m app.backup(a thin CLI wrapper around the same push logic) every 15 minutes.
To install the timer:
- Copy both files to
/etc/systemd/system/:sudo cp deploy/systemd/urd-content-backup.{service,timer} /etc/systemd/system/ - Edit
WorkingDirectoryinurd-content-backup.serviceto your actual repo path. sudo systemctl enable --now urd-content-backup.timer
Requires CONTENT_BACKUP_REMOTE set in .env (see .env.example) — it's
already passed through to the mcp_server container in
docker-compose.yml.
If your deployment was brought up with a non-default compose project name
(e.g. podman-compose -p somename up or COMPOSE_PROJECT_NAME set when
starting the stack), also set COMPOSE_PROJECT_NAME in .env so the
timer's podman-compose exec call targets the right containers — see the
EnvironmentFile= line in urd-content-backup.service.
Postgres (the db service) is a rebuildable search index — everything
in it can be reconstructed from content/entries/ by re-running
python -m app.sync (see "Running it" above). The git-remote content
backup in the previous section is what actually matters most for durability.
This section's Postgres backup is a fast-restore convenience on top of
that: skip the sync rebuild (which re-embeds every entry) by restoring a
recent dump instead, if one happens to be around and fresh. If it's ever
missing or stale, nothing is lost — just re-sync from content/entries/.
scripts/backup-pg.sh is a host-side script (not a container, not part of
Compose) that runs pg_dump -Fc — Postgres's custom/compressed dump
format, which supports selective and parallel restore via pg_restore,
a better fit here than a plain .sql dump — against the already-running
db service via podman-compose exec (or docker compose exec, see the
comment at the top of the script).
Why a separate mechanism from the git-remote content backup (previous
section) instead of reusing it? Postgres dumps are large, opaque, binary
blobs produced on a recurring schedule forever — a poor fit for git
(unbounded repo growth, no meaningful diffs). rsync/rclone are
purpose-built for "copy files to remote storage, prune independently,"
which is what a dump rotation actually needs.
Configuration (set in .env — this script runs on the host, outside
Compose, and sources .env directly itself, so these are not passed
through docker-compose.yml):
If your deployment uses a non-default compose project name, also set
COMPOSE_PROJECT_NAME in .env — same reasoning as "Content backup"
above, and the script auto-exports whatever it finds in .env so this
just works.
PG_BACKUP_DIR— where dumps land locally. Default./backups.PG_BACKUP_RETENTION_DAYS— local dumps older than this are pruned on every run. Default14. Local prune only — doesn't touch anything already copied off-box.PG_BACKUP_METHOD—none(default),rsync, orrclone.none— dump stays local only; the script logs a warning to stderr every run as a reminder.rsync— copies the dump toPG_BACKUP_REMOTE, e.g.PG_BACKUP_REMOTE=user@host:/path/— any box you can SSH into.rclone— copies the dump toPG_BACKUP_REMOTE, e.g.PG_BACKUP_REMOTE=remotename:path/— any rclone-configured remote (S3, B2, etc), for users who want object storage instead of a box they manage.
Run it manually with ./scripts/backup-pg.sh, or install the systemd
timer (same pattern as the content backup timer above):
- Copy both files to
/etc/systemd/system/:sudo cp deploy/systemd/urd-pg-backup.{service,timer} /etc/systemd/system/ - Edit
WorkingDirectoryandExecStartinurd-pg-backup.serviceto your actual repo path. sudo systemctl enable --now urd-pg-backup.timer
The timer runs daily. As with the content backup timer, systemd's
restricted default PATH may not include podman-compose — see the
comment in urd-pg-backup.service if the unit fails to find it.
Copy content/entries/example-entry.md, follow the frontmatter fields
described in content/SCHEMA.md, then run the sync command.
EMBEDDING_PROVIDER (docker-compose.yml, default fastembed) picks the
provider — see services/mcp_server/app/embeddings.py for the full list
(fastembed local, ollama local, voyage/openai cloud). Default is
fully local: no API key, no network call at request time, no cost, content
never leaves the machine — the model's weights are baked into the
mcp_server image at build time. write() embeds synchronously on every
create/update; sync backfills any entry missing one (e.g. hand-edited
files). Switching providers is a config change, but always requires
resizing the embedding column in db/init.sql to the new provider's
dimension and re-running sync to re-embed everything — different models'
vectors are never comparable to each other.
- Auth is a single shared bearer token, not per-user/per-client identity — fine for single-user/local, revisit before opening this up to multiple people.