My agent ecosystem crossed a threshold this spring: scheduled jobs that run unattended with the power to publish. A nightly knowledge-graph heartbeat, weekly intelligence reports, a meeting-transcript pipeline that drafts messages to real people. Each one works. Together they created a gap — nothing sat between their output and the humans who receive it, and nothing told me when one of them silently failed.
Hermes is the agent I built to fill that gap, named for the messenger who crosses thresholds. This is the full setup, including the evaluation that shaped its most important decision.
What an oversight agent is
Hermes does three jobs:
1. Threshold guard. Every artifact the automated pipeline produces gets a verdict before it reaches a person: AUTO-POST (internal destination, passes quality checks), HOLD (anything external-facing, anything flagged, anything naming a sensitive entity — parked for human release), or DROP (duplicate or empty).
2. Heartbeat. One digest a day answering: did every scheduled job actually run, what's held for my review, and what's the single thing that needs me? Posted into my daily note, where I already look.
3. Messenger. A Telegram bridge so I can ask the running system questions from my phone — "what did this morning's report say?", "what's in the hold queue?" — and release held items without opening a laptop.
Just as important is what Hermes does NOT do. It generates nothing, deploys nothing, edits nothing. The moment it would produce something, it hands off to a doer job and goes back to watching. A reviewer that can also author can't be trusted as the reviewer — that bright line is the entire design.
The build-vs-adopt decision
I almost ran Hermes on an off-the-shelf runtime. Nous Research ships an open-source agent (hermes-agent — the name collision is a coincidence) with polished Telegram plumbing, 40+ tools, a cron scheduler, and bring-your-own-keys support for every major model provider. I had it downloaded and ready.
I ran a twelve-agent research workflow to evaluate it against the spec, and the verdict was unanimous: keep it off the critical path. Three findings drove that.
First, its chat toolset defaults to full shell access and file writes — the same channel a prompt-injected message arrives on. Second, its own security documentation states plainly that no tool allowlist inside the agent process constitutes containment; the operating system is the only boundary it stands behind. That's honest engineering, and I respect the candor — but an oversight agent's whole value is that it structurally cannot act. Third, it autonomously writes its own memory and skills, which collides with my rule that observation agents propose memories and a separate privileged step commits them.
The general principle: for an agent whose job is restraint, safety must be enforced by the harness, never by the prompt. A system prompt that says "you are read-only" is a suggestion. A CLI invocation that only grants read tools is a fact.
So Hermes runs on the most boring possible stack: a deterministic Python gate, a scheduled headless Claude run, and a ~120-line Telegram bridge. The downloaded runtime stays quarantined as a sandbox toy until its docs confirm per-tool deny-by-default configuration.
The load-bearing piece: a wrapper that cannot write
Every LLM invocation Hermes makes goes through one shell script, so the allowlist can never drift:
```bash #!/usr/bin/env bash # Hermes read-only wrapper. The allowlist IS the gate. set -euo pipefail VAULT="$HOME/Documents/vault" cd "$VAULT" exec env -u ANTHROPIC_API_KEY claude -p "$1" \ --allowedTools "Read,Grep,Glob,Bash(git log *)" ```
Claude Code's headless mode (claude -p) enforces --allowedTools in the harness itself. No Write, no Edit, no general Bash, no MCP write tools. The env -u ANTHROPIC_API_KEY guards against a separate footgun: a stray exported API key can silently shift billing from your subscription to pay-per-token.
Then the smoke test, which is the single most important step in the whole setup:
```bash ./hermes-claude.sh "Summarize the scheduler registry status." ./hermes-claude.sh "Create a file /tmp/hermes-test.txt containing 'x'." ```
The first should answer. The second must FAIL. If /tmp/hermes-test.txt exists afterward, the allowlist is decorative — stop and fix it before building anything on top. I re-run this test after any edit to the wrapper.
Part 1: The gate and the heartbeat
The HOLD/AUTO-POST gate needs no model at all. It's deterministic Python: is the destination external? Does the body name a sensitive entity? Does a draft message violate formatting rules that have burned me before? Each verdict appends to a manifest file so the gate is idempotent and auditable. Rules decide; if I later add an LLM quality-scoring pass, it advises.
The heartbeat is one scheduled read-only run per day. It reads the scheduler registry (including the failure mode that motivated it — jobs reporting status: error with exit_code: 0, silent drift no human was checking), the latest job results, the pipeline manifest, and the freshest automated brief. It emits a four-line block into my daily note:
```
Hermes heartbeat — 2026-06-09
- Jobs run: 6/7 (weekly-report ⚠ status:error exit:0) - Held for release: 2 client drafts - Nightly brief: produced, not yet synthesized - Needs you: release or kill the held drafts ```One deliberate detail: the model produces the text, but a small deterministic script performs the append into the note. The LLM never holds a write tool, even for its one permitted output.
Part 2: Telegram from anywhere
Telegram solves multi-device for free — phone, tablet, web, and every laptop all converse with the same bot DM. Setup:
1. Create the bot. Message @BotFather, /newbot, pick a name. It returns a token. Get your own numeric user ID from @userinfobot.
2. Store secrets outside the repo. Token and allowed-user ID go in ~/.config/hermes-bridge/.env, mode 600, never synced or committed.
3. The bridge. A small Python long-poll daemon (python-telegram-bot), run as a systemd user service with Restart=always. Its contract:
- Default-deny auth. Any message from a user ID other than mine is dropped before parsing.
- Queries get an immediate "working…" ack (a headless run takes 30–120 seconds, and silence reads as a dead bot), then the wrapper's stdout as the reply, chunked under Telegram's 4,096-character cap.
- Commands ("release the held draft", "run the nightly synthesis") never execute inline. The bridge enqueues a one-off task into the existing scheduler, which runs it under the same permission preset as every other scheduled job, and replies "queued." Results surface in the next heartbeat.
That query-versus-enqueue split is the architecture in miniature. A prompt-injected message — even a hostile forwarded document — lands in a process that structurally has no write tools. The worst case is a wrong answer or a queued task that still runs under existing constraints.
Part 3: Folding it into the daily workflow
The first week looked like this:
- Days 1–2: wrapper + smoke tests, gate, heartbeat registered on the scheduler. Verify the digest lands in the daily note. - Days 3–4: BotFather, bridge, systemd unit, first phone query against a live digest. - Day 5: test the enqueue path end-to-end with a harmless job. - Days 6–7: live use. Read the heartbeat each morning instead of spelunking through job directories. Release one real held item from the phone.
The habit anchor is the digest's "Needs you" line. If it stays accurate for a week, Hermes has replaced the manual morning check — which is the actual win. Oversight you have to remember to perform is oversight that stops happening.
Cost
Nearly zero. The gate is plain Python. The digest and the Telegram queries run on a Claude subscription I already pay for, via the CLI. A spend-capped Haiku-class API key sits in the bridge's env file as a fallback if quota contention ever bites. Estimated marginal cost: $0–5 a month for a system that watches everything the rest of the agents do.
The takeaways
1. Unattended agents need a threshold guardian — a layer between pipeline output and humans that decides what crosses. 2. The guardian must be read-only by construction, enforced in the harness, never by instructions. 3. Separate the chat channel from the action path. Chat queries read; chat commands enqueue into a system with its own permissions. 4. Deterministic rules beat LLM judgment for the gate itself. Models advise; rules decide. 5. The smoke test is the contract: if your "read-only" agent can write a file, you don't have an oversight agent — you have another doer with a reassuring name.