My Freebo platform has a Hermes agent that lives on an always-on desktop and does nothing but run cron jobs. Twenty of them. Every half hour it checks whether production SHA changed. Every ten minutes it looks for urgent operator signals. Hourly it checks that Railway, Supabase, PostHog, Novu, and Stripe are all breathing. Every morning it posts an ops digest to a private Slack channel I own.
Through most of September, roughly four in every ten of those posts contained the word “unavailable.”
Not “the API is down.” Not “I hit a rate limit.” Just “PostHog unavailable” or “Railway unavailable” or “Supabase unavailable (approval-pending),” reported with the same confidence as a real finding. Which meant I had to read every post, decide whether to believe it, and most of the time just shrug and wait for the next tick.
That is not an agent. That is a very expensive carrier pigeon with a drinking problem.
0
cron jobs
0
MCP servers removed
0%
posts reporting 'unavailable'
Two bugs wearing the same costume
I spent a weekend tracing it before the shape clicked. The “unavailable” line was coming from two completely different failure modes that happened to look identical in Slack.
Bug one: MCP sessions wedge on network hiccups. The box is on wifi. Wifi has moments. A long-lived MCP server opens a websocket at gateway startup and keeps it open forever. When wifi burps for ten seconds, the socket dies in a way that nothing inside the server notices. The next tool call hangs, times out, and the agent reports the provider is down. The provider is fine. The socket is dead. Nothing short of restarting the gateway brings it back.
Bug two: the agent kept writing one-line shell pipelines that failed its own security guard. When the agent wanted a Railway log line, it would construct something like railway variables | python3 -c "..." or curl ... | python3 -c "..." on the fly. Hermes has a terminal guard that flags those for human approval, which is fine when I’m sitting at a laptop. In a cron job firing at 3 a.m., there is no human to approve. The job sits “approval-pending” forever and reports the data source as unavailable.
Two completely different problems, both printing the same message. For a month I’d been looking at the Slack posts, assuming a provider outage, and ignoring most of them.
The fix was fewer moving parts
The cleanest fix I could draw was to delete the long-lived MCP sessions entirely and replace them with a single-file Python CLI. One tool. Stateless. HTTPS requests only. The agent calls it like fq posthog sql '...' or fq railway deployments --service api and gets JSON back on stdout.
That one change kills both bugs at once. There is no socket to wedge, because every invocation is a fresh process making one HTTPS call. There is no shell pipeline for the terminal guard to flag, because fq fetches and filters inside itself — the agent writes fq ..., not fq ... | python3 -c "...".
The script’s own docstring is the clearest explanation of what the design is doing:

The piece I’m proudest of in that is the exit code contract. Three outcomes, three codes, and only one of them is allowed to show up in Slack as “unavailable.”
- Exit 2 means the query was wrong. The agent asked for a column that doesn’t exist, a service that isn’t in production, a time range in the future. That’s the agent’s problem to fix, not a provider outage. The agent sees exit 2 and rewrites the query.
- Exit 3 means the API key was rejected. That’s an auth problem, worth saying out loud.
- Exit 4 means
fqretried, got nothing, and is giving up. Only exit 4 is allowed to produce the word “unavailable” in a Slack post. Everything else is the agent’s job to handle silently.
Before this contract existed, the agent would see any failure and label it an outage. Now it has to earn the word.
The jobs don’t know they moved
Here’s the part I like most about where this landed. The 20 cron jobs didn’t change at all. Their prompts still say “check Railway status” and “query PostHog for the last hour of signups” and “list recent Stripe charges.” I just rewrote the agent’s reference doc to say “the way you talk to Railway is now fq railway status” and let the agent re-read it.
The cron job for production-SHA watch still fires every 30 minutes, still classifies commits as user_visible or operator_fix or internal_only, still posts a one-line release observation to Slack. From the outside, nothing moved. From inside, five MCP servers are gone, and the gateway’s attack surface went down with them.
- Sept 2026Cron agent reports 'unavailable' on ~42% of postsI ignore most Slack digests. The agent is noise.
- WeekendTrace the message to two unrelated bugsWedged MCP sockets + terminal guard eating shell pipelines.
- Oct 1Write fq.py, delete 5 MCP servers, rewrite reference docThe gateway now holds two MCP sessions (Todoist + Linear), down from seven.
- Oct 2Morning digest reads cleanSix Railway services green, no 'unavailable' lines, one real finding surfaced.
The lesson I didn’t want
I built a lot of MCP into this agent because MCP is the fashionable surface for giving an LLM access to external systems. It’s structured, it’s typed, it’s discoverable. All of that is real. None of it survives contact with a wifi connection that drops a packet once an hour and a security guard that doesn’t know the difference between a cron job and an interactive terminal.
The thing I keep learning, in different forms, is that durability beats elegance for anything that has to work when I’m not watching. A one-file Python script that does an HTTPS GET and prints JSON is embarrassingly less sophisticated than a stateful MCP server. It is also incapable of wedging, incapable of asking me to approve anything at 3 a.m., and incapable of lying about whether it’s up.
If the agent can report “unavailable” without knowing whether that’s true, you don’t have an agent. You have a very confident mail carrier.
— Me, to myself, on Oct 1
I’m not retiring MCP everywhere. Todoist and Linear still ride on it in this gateway, because they’re interactive-feeling and the wedging never cost me much there. But for anything a cron job needs — anything that has to work in the middle of the night, on wifi, with no human to referee — I’m biasing hard toward the dumbest possible transport that can tell the truth about its own state.
The next thing I write for this agent will probably not be an MCP server. It will be another small CLI.