Русская версия: README.ru.md
The short version. 35 modules for running a self-hosted LLM agent in production — written for Hermes Agent: cron reliability, cost guards, context compaction, harness probes and skill scanners. Everyone can show you an agent that books a flight. This repository is about the part that comes after: keeping that agent alive, honest and affordable when it runs around the clock, on its own schedule, with its own wallet.
Who it is for. Anyone running an LLM agent on a schedule — not in a notebook, in production — and anyone who is about to. It is written against Hermes Agent, but most of the tooling applies to any gateway with a scheduler, a state store and provider APIs.
What you get. Thirty-five self-contained folders. Each one is the fix for a failure that actually happened, shipped as a runnable script or a configuration template, together with the incident that caused it — so you can take one folder without adopting the rest: read the incident, copy the script, keep the habit. The numbers behind every claim are in BENCHMARKS.md, one deployment with the measurement window next to each figure, so a claim can be re-derived rather than believed.
Four failure modes account for most agent downtime and spend:
- one session grows until it dominates spend and response quality,
- the gateway process stays alive while refusing new writes,
- scheduled jobs drift or die silently,
- the documented schedule diverges from the real one.
Each one is handled by a component with a defined detection rule and response path:
- spend is anchored to the provider balance and attributed per session,
- process health is judged by heartbeat and fatal log markers, from outside the gateway,
- every scheduled prompt is checked for self-containment before it ships,
- the ops documentation is generated from the job list rather than maintained by hand.
Thirty-five modules under numbered directories, each readable on its own, cross-referenced in SERIES.md.
┌──────────────────────────────┐
human (DM/topics) ──▶│ gateway + role profiles │◀── SOUL.md per role
│ coordinator · chef · doctor │ model pinning
│ operator (3rd model) │ topic walls
└──────────────┬───────────────┘
│ state DB + logs + heartbeat file
┌──────────────┬───────────────┼────────────────┬────────────────┐
▼ ▼ ▼ ▼ ▼
scheduler gateway wallet guard usage DB backups
(cron jobs) supervisor (every 30 min) (sqlite) (daily)
│ (external) │
│ + guard cron ▼
▼ cost dashboard
no-agent watchers (attribution)
LLM digests (gated)
one-shot reminders
Three properties hold at every layer:
- Watchers - watched by a different layer than the one they watch. The gateway is supervised by an external daemon; the daemon is re-spawned by a scheduler-owned guard cron; the schedule itself is audited by a daily operator sweep running on a separate model.
- Silence - valid state. Watchers print nothing when healthy; the scheduler delivers nothing on empty stdout.
- Nothing heals from the inside. Restart, restore, and escalation paths live outside the gateway, and state is never auto-restored by a robot.
| Module | Responsibility |
|---|---|
01-agent-wallet-guard |
Measures the provider balance directly; alerts on floor breaches and daily burn, naming the top session by cache-read tokens. Silent when healthy. |
02-agent-gateway-supervisor |
External daemon as a systemd user unit (not a cron guard): heartbeat staleness + fresh fatal log markers; restarts only via systemctl (polkit grant); cooldown + hourly cap; maintenance pause; escalation instead of auto-restore. |
03-agent-ops-playbook |
Sanitized incident autopsies and decision trees for the recurring failure classes. |
04-role-profiles |
SOUL.md master-prompt template, per-role model/config pinning, file-based role bridges. |
05-cron-of-crons |
Operator sweep brief, who-watches-whom matrix, delivery policy for digests, alerts, and one-shots. |
06-hermes-plugins-skills |
Packaged skill (one-shot reminder) and cron recipes: monitor-gated digests, watchdog chains. |
07-fresh-prompt-linter |
Deterministic checker for prompt self-containment; gates scheduled prompts before deploy. |
08-session-housekeeping |
Session/memory/skill lifecycle rules and consolidation jobs. |
09-ops-as-data |
Exporter that renders the watch table from the scheduler job list. |
10-cost-dashboard |
One-file HTML panel: spend by day, provider, and top sessions. |
11-domain-persona-packs |
Role cartridges for 04: chef with inventory, doctor with a limits file, operator. |
12-topic-routing |
Topic-as-domain isolation, ignore-list walls, alert routing to the DM. |
13-voice-input-hypotheses |
Transcription-garble defense: names from voice - hypotheses until verified against ground truth. |
14-agent-data-intake |
Reliable device-to-agent intake: multi-threaded receiver as a systemd user unit (never a gateway cron), device-side spool that retries until delivered, 24/7 scanning with no time windows. |
15-agent-tool-guardrails |
Judge from outside, applied to commands and worktrees: an AST-based pre_tool_call gate (not a regex on the string), a skill scanner for injection/exfil before an install is trusted, and git-worktree lanes whose cards close with orchestrator-produced evidence. |
16-gepa-skill-tuner |
Reflective prompt evolution (GEPA) pointed at the artifacts an agent actually ships — the prompt and the SKILL.md — with a declared metric budget, a deduplicated dataset, a held-out split, and a diff as the output. |
17-model-slot-bakeoff |
Picks a model for a slot by measuring the slot's real tasks: billed cost per task, latency, the upstream that actually served, hidden-test grade, and tool-call support. |
18-harness-probes |
Deterministic acceptance probes for the harness itself: JSON-declared checks over artifacts and guard scripts, a baseline that only accept moves, and a check mode that fails on regression. |
19-cost-governance |
Budget cap with teeth — pauses the N most expensive jobs once a day and resumes them — cost per successful task per role, and a prompt-vs-toolset audit. |
20-task-evals-and-autopsy |
Did the job do its work: output freshness, size and shape, failure streaks, stuck queues — plus an error classifier that reports each new failure once and a compiled operator brief. |
21-blind-spot-audit |
Measures the share of accepted output a strong judge finds wrong: weekly random sample, redaction, rubric, one recommended change per report. |
22-difficulty-router |
Routes scheduled jobs to a model tier by measured difficulty, and refuses a repin that makes mechanical work more expensive per token. |
23-research-intake |
OAI-PMH harvest of a paper corpus, local term counting in equal windows, and the journal schema — first step, metric, kill date, rejection reason — that turns reading into adopted changes. |
24-llm-to-script |
The inversion that removes the largest cost line: a deterministic collector plus a small formatter prompt, the shortlist tool that finds the next candidate, and measured before/after numbers. |
25-querylog-domain-scout |
Aggregates an AdGuard query log per client, so I can see which device talks to which domains and what is new. |
26-ecosystem-map |
Builds a live map of the agent estate from the job list and profiles, so the documentation cannot drift from reality. |
27-research-scout |
Harvests arXiv over OAI-PMH, scores papers by theme and reception, writes a brief for the model and a stable fingerprint that keeps quiet weeks silent. |
28-digest-delivery-health |
Keeps a journal of digest deliveries (sent, silent, failed) and surfaces the channels that stopped working. |
29-thinking-layer-cost |
Reports what the reasoning layer costs per day against a cap, so the expensive part stays visible. |
30-schedule-audit |
Reviews the schedule for jobs worth moving off-peak and for the few that actually cost real money. |
31-compaction-effect-check |
Measures a context-compaction policy change against its own before-and-after windows instead of assuming it helped. |
32-routing-outcomes |
Counts how each model performs on real scheduled work, which is what decides a repin. |
33-multi-model-consilium |
Runs a decision past a second model: one proposes, a different family attacks the proposal, the author answers the attack and settles the plan. The transcript, the token bill and the rejected objections are kept. |
34-skill-library-janitor |
Walks a skill library and reports what is checkably broken — missing paths, scripts that stopped compiling, drifted frontmatter. Silent when clean, cheap enough for a monthly cron. |
35-context-compaction-engine |
Compacts a session into a working state — task, decisions, exact values, next step — and carries that state across cycles instead of re-summarising it; accepted by whether the task can still be continued, not by how many tokens were freed. |
Clone the repository and start with the module that matches your current pain:
git clone <this-repository> hermes-agent-ops
cd hermes-agent-ops
# spend watch: point the guard at your provider balance endpoint
cd 01-agent-wallet-guard
export WALLET_API_KEY=sk-...
python3 agent_wallet_guard.py # silent when healthy15 ships with tests (tests/) and needs one pure-python wheel (bashlex, vendored by its installer). The guard, supervisor, linter, exporter, and dashboard - Python 3.10+ standard library only — no dependency install. Templates and prompt briefs in the remaining modules - used as-is.
- Python: 3.10+; no third-party packages for any shipped tool (network calls use
urllib, rendering is plain HTML, state is JSON). - Optional datastore: sqlite usage table for per-session attribution (schema below). Without it, the guard still alerts; it simply cannot name the culprit session.
- Transport: watchers write to stdout (scheduler delivers), the supervisor talks to Telegram via Bot API using a token from the environment. Delivery is the scheduler's concern.
Scheduled components - meant to run as:
every 30 min → 01 wallet guard (no-agent job, stdout → alert topic)
every 5 min → 02 supervisor guard cron (re-spawn daemon if dead)
continuous → 02 supervisor daemon (setsid, detached)
0 5 * * * → 05 operator sweep (LLM job, third model)
daily 11:00 → 09 exporter → commit (docs generation)
Lint a scheduled prompt before it ships:
python3 07-fresh-prompt-linter/fresh_prompt_linter.py --prompt-file brief.md
echo $? # 0 = deploy, 1 = make it self-contained firstGenerate the ops watch table from the job list:
python3 09-ops-as-data/cron_exporter.py \
--jobs jobs.json --out WATCH-TABLE.mdRender the cost panel:
python3 10-cost-dashboard/cost_dashboard.py \
--db /var/lib/agent/state.db --days 14 --out cost-dashboard.html| Artifact | Producer | Consumer |
|---|---|---|
| Alert lines (stdout) | 01 guard, watchdog scripts |
scheduler → alert topic |
| Telegram alerts | 02 supervisor |
operator DM / ops topic |
| Watch table (markdown) | 09 exporter |
committed docs, operator sweep diff |
| Cost dashboard (HTML) | 10 dashboard |
browser / static host |
| Incident reports (markdown) | 03 |
humans, decision trees |
| Guard state (JSON) | 01 |
the guard itself (throttle windows, day anchor) |
Shared sqlite usage table (written by the gateway's usage accounting, read by 01 and 10):
CREATE TABLE session_model_usage (
session_id TEXT,
billing_provider TEXT,
first_seen INTEGER, -- unix epoch
cache_read_tokens INTEGER,
input_tokens INTEGER,
output_tokens INTEGER,
reasoning_tokens INTEGER,
api_call_count INTEGER
);Scheduler job list consumed by the exporter (09-ops-as-data/jobs.example.json):
{
"id": "ab12cd34",
"name": "morning train digest",
"schedule": "0 6 * * 1-5",
"agent": true,
"deliver": "topic: commutes",
"notes": "Mon-Fri only; deviations only, silence when on schedule",
"watcher": "operator"
}Guard state file (JSON, one per host): day anchor balance, alert throttle timestamps, last observed balance. A top-up re-anchors the day so refills - not counted as burn.
- Time anchors - UTC. Cron expressions and one-shot timestamps - written in UTC; the operator brief and examples assume a deployment in Europe/Berlin.
- Compaction before cost. Sessions shrink at a threshold low enough that a failing compression cannot rack up hours of paid retries (the
03/incidents/530k-token-session.mdwrite-up is the reference case). - Aux calls belong to the cheap provider. Compression, titling, and review calls dominate token counts; pinning them off the primary provider - budget decision, not an optimization.
- Backups precede restores. State DB backups - daily and kept ~1–2 days; the recovery path in
02quarantines before restoring and never deletes the damaged copy. - Exclusions - documented per deployment (e.g. a chef role exempt from the operator sweep). They - a configuration choice, not an oversight.
In scope: observability and spend accounting for agent traffic, external supervision of gateway processes, role and session lifecycle management, scheduler documentation, prompt hygiene, and incident runbooks. Out of scope: model training, prompt content marketplaces, and the Hermes Agent core itself — this repository operates around a gateway rather than modifying one.
- Single-gateway deployments with several professional roles and strict budget ceilings.
- Teams running scheduled LLM jobs who need the docs, the gates, and the watch chain before trusting them unattended.
- Operators inheriting an agent deployment who need the incident record and decision trees more than another architecture diagram.
- The prompt linter is heuristic: it flags self-containment violations with high recall, but passing it is not proof a prompt will succeed.
- The cost dashboard estimates from list prices for attribution; the wallet guard measures actual balances for fact. Never bill from the estimate.
- Delivered examples use Telegram topics and Bot API as the reference transport; the guard and watchdog contract is stdout, so other transports require no code change in the watchers.
- Incident reports - sanitized: details that would identify the operator or the infrastructure - removed.
hermes-agent-ops/
├── 01-agent-wallet-guard/ agent_wallet_guard.py, guard_config.example.env
├── 02-agent-gateway-supervisor/ gateway_supervisor.py, systemd/, polkit/, scripts/
├── 03-agent-ops-playbook/ incidents/, decision-trees.md
├── 04-role-profiles/ template/, docs/
├── 05-cron-of-crons/ operator_prompt.example.md, watch-table.md, delivery-policy.md
├── 06-hermes-plugins-skills/ skills/, cron-recipes/
├── 07-fresh-prompt-linter/ fresh_prompt_linter.py
├── 08-session-housekeeping/
├── 09-ops-as-data/ cron_exporter.py, jobs.example.json
├── 10-cost-dashboard/ cost_dashboard.py
├── 11-domain-persona-packs/
├── 12-topic-routing/
├── 13-voice-input-hypotheses/
├── 14-agent-data-intake/
├── 15-agent-tool-guardrails/ hooks/, lanes.py, tests/, install.sh
├── 16-gepa-skill-tuner/ tuner.py, tests/, examples/
├── 17-model-slot-bakeoff/ bakeoff.py, tests/, examples/
├── 18-harness-probes/ probes.py, probes_check.sh, examples/, tests/
├── 19-cost-governance/ budget_guard.py, cost_per_outcome.py, toolsets_audit.py, examples/, tests/
├── 20-task-evals-and-autopsy/ task_evals.py, postmortem.py, briefing.py, examples/, tests/
├── 21-blind-spot-audit/ sample_outputs.py, judge_prompt.md, examples/, tests/
├── 22-difficulty-router/ difficulty_router.py, examples/, tests/
├── 23-research-intake/ oai_harvest.py, direction_digest.py, backlog.md, examples/, tests/
├── 24-llm-to-script/ collector.py, token_audit.py, examples/, tests/
├── 25-querylog-domain-scout/ ag_domain_scout.py, examples/, tests/
├── 26-ecosystem-map/ ecosystem_map.py, examples/, tests/
├── 27-research-scout/ research_scout.py, examples/, tests/
├── 28-digest-delivery-health/ digest_health.py, examples/, tests/
├── 29-thinking-layer-cost/ thinking_layer_cost.py, examples/, tests/
├── 30-schedule-audit/ schedule_audit.py, examples/, tests/
├── 31-compaction-effect-check/ compaction_effect_check.py, examples/, tests/
├── 32-routing-outcomes/ routing_outcomes.py, examples/, tests/
├── 33-multi-model-consilium/ consilium.py, examples/, tests/
├── 34-skill-library-janitor/ janitor.py, examples/, tests/
├── 35-context-compaction-engine/ autocompact.py, resume_eval.py, examples/, tests/
├── BENCHMARKS.md before/after numbers from the deployment
├── SERIES.md how the modules extend each other
├── ROADMAP.md
└── site/ banner assets
Reference layout for a single host:
- Gateway + role profiles per
04; state and logs under one directory the supervisor can read. - Scheduler jobs per
05: watchers on short intervals, digests on their own cadence, the operator sweep once daily. - Supervisor daemon started detached (
setsid), restarted by its guard cron — its PID is expected to die with gateway restarts. - Backup job for the state DB, daily, kept ~1–2 days.
- Wallet guard every 30 minutes, key from the environment, never committed.
Minimum viable deployment is two components: the wallet guard and the supervisor + guard pair. The playbook is read-only; the remaining modules add roles, documentation, and attribution as the deployment grows.
Code is MIT (see LICENSE). Documentation is CC-BY-4.0.
The incident reports in this repository - sanitized reconstructions; tooling is provided as-is for operators to adapt to their own infrastructure. This is an independent project and is not affiliated with or endorsed by Nous Research.