✳weblog

Supercharging agents: Hermes skills that earn their keep

Most agent setups we've seen die the same death: every lesson, correction, and hard-won procedure gets stuffed into the system prompt or a giant memory blob until the context is 90% ritual and 10% task. Hermes Agent — the runtime that drives most of our automation on a pair of self-hosted boxes — takes a different approach, and it's the reason our agent still feels sharp after months of daily use.

Skills are procedural memory

A Hermes skill is a markdown file: frontmatter plus a written procedure. The agent loads one only when a task needs it, via \1. Nothing sits in context between tasks. The distinction matters: facts about the world go in memory; \1 goes in skills. When the agent learns that a particular API returns \1 headers around plain-JSON bodies, or that a cron scheduler stores times in UTC while the user reads Dhaka time, that knowledge belongs in a skill where it loads exactly when relevant — and never taxes the other four hundred tasks that don't care.

How we organize them

Skills live in categories that mirror the boxes they run on: \1, \1, \1, \1, \1, \1, and a few project-specific ones. Two operations cover the whole lifecycle:

the pain is fresh, not in some retrospective that never happens

The honest caveat: skills are living documents, not trophies. A \1 can change a command's flags or rename a tool, and any skill that quotes exact commands goes stale silently. We treat a skill failure the same way we treat a flaky test — as a prompt to re-read the skill against reality and patch it. A skill that hasn't been touched since it was written is usually wrong about something.

The ones that earn their keep

We run a few dozen skills; these are the ones that get loaded constantly:

gate before side effects, verify before reporting. It roughly halves output noise and stops the agent from narrating every step it's about to take.

worker lanes while a stronger model orchestrates and reviews. This one paid for itself the first week.

itself: config, providers, cron, MCP servers. When the agent fixes its own runtime, it should follow its own documentation.

branch to merge, plus review discipline (diff first, comment inline, no drive-by refactors). Together they made the agent usable as a second reviewer.

root cause, verify. The value is mostly negative: it stops the agent from patching symptoms because a plausible fix appeared first.

implementation plans with file paths and acceptance criteria before any code gets written. Cheap insurance against hour-long wrong directions.

agent writes the failing test first instead of "verifying" with a successful build it wanted to succeed.

arXiv and RSS/blog watching, which is how new posts on topics we follow find us instead of the other way around.

the "it's not just X, it's Y" constructions. This blog's drafts go through it.

That's more than eight, but the boundary is fuzzy — half of these are two skills that only make sense as a pair.

A skill is procedural memory: the agent doesn't get smarter by knowing

more at all times, it gets smarter by loading exactly what the task needs.

Everything else stays on disk, where it costs nothing.

The shape of a skill is deliberately boring — frontmatter that says \1, markdown that says \1:


---
name: selfhost-update
description: Use when updating self-hosted services (9router, hive, dashboard).
---

## Steps

1. Snapshot current state (versions, config paths) before touching anything.
2. Update one service at a time; verify health before moving on.
3. Record quirks discovered during the update back into this skill.

No schema, no runtime, no registration step. The description is the routing table: it's what the agent matches against when deciding a skill is relevant.

Why this beats a bigger context

The failure mode we were avoiding is drift toward generic answers. An agent with a huge static context answers everything the same way — from the center of its training distribution — because nothing in the context says \1. Skills are how the agent gets better at \1 setup instead: your deployment quirks, your naming conventions, your verified-safe command sequences, written down once and loaded on demand.

None of this is exotic. It's markdown files and a loader. That's the point — the mechanism is simple enough that the content can be where the effort goes.