hermes-validate
Test, observe, and improve Hermes Agent behavior — send test messages, read session traces, identify routing failures, and fix SOUL.md/SKILL.md
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Test, observe, and improve Hermes Agent behavior — send test messages, read session traces, identify routing failures, and fix SOUL.md/SKILL.md
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Hand repo work to a bounded Claude Code episode via the `hermes-cc.sh` verb dispatcher, then answer with its verdict. Use when triage needs the actual source — "why is X failing, look in the repo", "what changed in Y", "warum ist Z rot, schau ins repo", "read the code and tell me", "check the repo for", a red monitor whose cause is code-shaped, a stale GitHub issue, or any question you can only answer by guessing otherwise. Also the path for "file an issue about what you find" (author tier) and, only with Johannes's explicit confirmation, "make that change" (implement tier — isolated worktree, draft PR), then "merge it" to land that PR.
Update Hermes Agent to latest version, resolve conflicts in locally-modified source files, and restart the gateway
Read, search, and write Johannes's Obsidian vault (the PARA second brain at ~/SourceRoot/brain) via the Obsidian CLI (metadata-aware — backlinks, tags, Dataview) with a filesystem fallback. Capture notes to the inbox, create resource/inspiration notes with the right frontmatter, search by text/tag/backlink.
Diagnose and remediate homelab / VPS / Mac Mini infrastructure through the bounded `hermes-ops.sh` verb dispatcher — monitors down, unhealthy or restart-looping containers, cron failures, push-monitor heartbeat gaps, uptime-kuma config corruption, deploy drift. Read-only triage is free; every mutation is a named verb needing --confirm and --why. Use for "#alerts" messages, "what's down", "is everything up", "why is X red", "restart Y", "redeploy Z".
Call the argo REST API (https://argo.jkrumm.com/api) for TickTick tasks, Gmail, Calendar, Docker (homelab + VPS), UptimeKuma, Slack, weather, Garmin Health, Strength tracking, WalkingPad treadmill stats, AI usage/cost, user profile, and read-only SQL — use curl with Bearer $HOMELAB_API_KEY
TTS audio synthesis via the VPS audio-gateway (https://audio-gateway.jkrumm.com/v1) — Gemini 3.1 Flash TTS, voice "Charon", OpenAI-compatible /v1/audio/speech. Use for any spoken output: morning/evening briefings, voice memos, ad-hoc TTS.
| name | hermes-validate |
| description | Test, observe, and improve Hermes Agent behavior — send test messages, read session traces, identify routing failures, and fix SOUL.md/SKILL.md |
Iterative workflow for validating and improving Hermes skill routing and response quality. Run this when adding a new skill, after changing SOUL.md/SKILL.md, or when Hermes gives a bad response.
Two ways to drive Hermes. Pick by what you're testing, not by convenience:
| Testing | Use | Why |
|---|---|---|
| Routing, skills, SOUL, answer quality | A — gateway API | Synchronous, no Slack-auth dependency, no channel noise |
| Threading, session keying, mrkdwn, TTS/media delivery | B — Slack API | The only path that exercises the Slack adapter at all |
Method A runs do not write
~/.hermes/sessions/*.jsonl— that store is the Slack path. Verify method-A routing from the response content plusgateway.log/agent.log, not the sessions dir.
Where secrets come from (v0.19.0+): there is no
~/.hermes/.envany more — it was retired when Hermes moved to nativesecrets.commandresolution. Read values directly from the cache withsecrets-run read op://…(the drop-inopshim; works headless on the mini). This also sidesteps the oldcut -d= -f2-footgun, since nothing is being parsed out of a dotenv line:secrets-run readreturns the raw value,=chars and all.
HOST=$(secrets-run read op://hermes/gateway/host)
KEY=$(secrets-run read op://hermes/gateway/api-server-key)
curl -s -X POST "http://$HOST:8642/v1/chat/completions" \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" --max-time 180 \
-d '{"model":"hermes-agent","messages":[{"role":"user","content":"your test prompt here"}]}' \
| python3 -c "import json,sys; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
HK=$(secrets-run read op://common/api/SECRET)
CH=$(secrets-run read op://hermes/slack/channel-hermes)
# New thread (fresh context window) — returns {"ts": "...", "channel": "..."}
curl -s -X POST -H "Authorization: Bearer $HK" -H "Content-Type: application/json" \
-d '{"text":"your test prompt here"}' \
"https://argo.jkrumm.com/api/slack/channels/$CH/messages"
# Continue that thread (same session) — TS is the ts returned above
curl -s -X POST -H "Authorization: Bearer $HK" -H "Content-Type: application/json" \
-d '{"text":"your follow-up here"}' \
"https://argo.jkrumm.com/api/slack/channels/$CH/messages/$TS/reply"
Caveat: the HomeLab synthetic sender can be rejected with
Unauthorized user: U… (HomeLab) on slackinagent.log(theallow_bots/SLACK_ALLOW_ALL_USERSpath isn't always honoured for it). If that line appears, no session is created — use method A. (Worked fine at v0.19.0 withallow_bots: all.)Method B is the only way to exercise the real Slack path: threading,
format_message()mrkdwn normalization, media/TTS attachment delivery, and session keying. Method A bypasses all of it. The response POST returns the messagets— keep it; it is the thread root you look for instate.dband the anchor for a continuation reply.
Wait for a response (method B):
until tail -1 ~/.hermes/logs/agent.log | grep -q "response ready"; do sleep 5; done
tail -3 ~/.hermes/logs/agent.log
platforms.slack.reply_in_thread is true, so one Slack thread == one session ==
one context window. build_session_key() appends the thread ts whenever
source.thread_id is set, and the Slack adapter sets it to the message's own ts for
top-level messages. Consequences when validating:
thread_ts == ts on your own message. thread_ts: null now means something
went wrong; before the v0.19.0 flip it meant the opposite.isolate_user is forced off
whenever a thread is present, so group_sessions_per_user: true doesn't apply here.Confirm the session your message actually created:
sqlite3 -header -column ~/.hermes/state.db \
"SELECT substr(session_key,1,66) k, thread_id, message_count mc,
datetime(started_at,'unixepoch','localtime') started
FROM sessions WHERE session_key LIKE '%slack%'
ORDER BY started_at DESC LIMIT 3;"
A healthy post-flip key looks like agent:main:slack:group:<team>:<C…>:<ts> with
thread_id populated. A key ending at the channel ID is a pre-flip (or misrouted)
channel-wide session.
Traces live in ~/.hermes/state.db (messages joined to sessions).
~/.hermes/sessions/*.jsonlis dead — do not read it. It stopped being written on 2026-06-02 and the 155 files there are frozen history. Sorting that directory and taking the newest file silently hands you a seven-week-old conversation that looks current. Verify withls -t ~/.hermes/sessions/*.jsonl | head -1before ever trusting it.
By thread (the useful form — pass the thread root ts you got back from method B):
sqlite3 ~/.hermes/state.db "
SELECT m.role, COALESCE(m.tool_name,''), substr(REPLACE(COALESCE(m.content,''),char(10),' '),1,150)
FROM messages m JOIN sessions s ON s.id = m.session_id
WHERE s.session_key LIKE '%<thread_ts>%'
ORDER BY m.timestamp;"
Most recent turn regardless of thread:
sqlite3 ~/.hermes/state.db "
SELECT m.role, COALESCE(m.tool_name,''), substr(REPLACE(COALESCE(m.content,''),char(10),' '),1,150)
FROM messages m
WHERE m.session_id = (SELECT id FROM sessions ORDER BY started_at DESC LIMIT 1)
ORDER BY m.timestamp;"
Reasoning is in the reasoning / reasoning_content columns; tool arguments in
tool_calls. A continuation reply arrives prefixed [Replying to: "…"], which is how you
confirm the parent message was threaded onto the right session.
From agent.log response line:
response ready: platform=slack chat=... time=51.7s api_calls=3 response=751 chars
| Metric | Good | Investigate |
|---|---|---|
api_calls | ≤5 | >8 |
time | <90s | >150s |
Healthy trace pattern (verified 2026-07-24, 3 api_calls, 21.3s):
user | | [U… | Slack user <@U…>] wie viele offene Watchdog-Items …
assistant | |
tool | skill_view | {"success": true, "name": "argo-api", …} ← skill hit directly
tool | skill_view | {"file": "references/infrastructure.md", …} ← drilled to the reference
tool | terminal | {"output": "{\"generatedAt\":…}"} ← one curl, not many
assistant | | Es gibt aktuell 12 offene Watchdog-Items … ← clean answer
Two skill_view calls in a row is normal now, not a red flag: argo-api is a hub
skill and the second call opens the right references/*.md. What matters is that neither
is a skills_list.
Red flags in the trace:
skills_list appearing twice before skill_view → skill not mentioned in SOUL.md by nameexecute_code with Python requests → SOUL.md needs to say "use terminal, not execute_code"find skills/argo-api/reference.md → dead file path in SOUL.md (rename to skill name)not found (title generation, compression) → non-blocking, but
check the model name in config.yaml still exists on the endpoint. Auxiliaries are
DeepSeek-V4-Flash; any log line naming a gpt-* auxiliary is stale config, not a bug.Command Approval Required / blocked on a routine argo curl → a tirith patch didn't
re-apply; see the hermes-update skill.Symptom: session shows search_files, read_file with a path like skills/argo-api/reference.md
Cause: SOUL.md had a dead file path reference
Fix: Replace file paths in SOUL.md with skill names: skill_view('argo-api')
execute_code instead of terminal for curlSymptom: session shows Python requests code, often with import errors
Fix: Add explicit instruction to SOUL.md:
"use
terminalwith curl — neverexecute_code"
skills_list calls)Symptom: skills_list appears twice in trace before skill_view
Cause: the model lists all skills to verify, rather than calling skill_view directly
Fix: In SOUL.md name the exact skill and tool call: call skill_view('argo-api')
Symptom: Hermes reports wrong status (e.g., UptimeKuma status: 1 called "down")
Fix: Add field semantics to the relevant SKILL.md. Example:
"
status: 1= UP,status: 0= DOWN"
Symptom: api_calls very high, session reasoning repeats the same question
Cause: Usually a dead reference in SOUL.md causing the model to search and give up repeatedly
Fix: Find and remove the dead reference, point to skill name instead
Skills are symlinked so SKILL.md changes are live immediately.
SOUL.md, config.yaml, and any patched Python module require a gateway restart —
Python modules are imported once at startup, so an edited tirith_security.py (or any
other patched file) is inert in the running process until you restart, no matter what a
direct in-process test reports.
hermes gateway restart # launchd-supervised; drains in-flight runs (up to 180s)
# Wait for Slack to reconnect — poll the state file, not the log
until [[ "$(jq -r '.platforms.slack.state' ~/.hermes/gateway_state.json)" == "connected" ]]; do sleep 2; done
jq -r '"pid=\(.pid) gateway=\(.gateway_state) slack=\(.platforms.slack.state)"' ~/.hermes/gateway_state.json
Don't wait on
grep -q "Bolt app is running" ~/.hermes/logs/agent.log— the log is append-only, so it matches a previous startup instantly and the loop exits before the new process is up. Pollgateway_state.jsoninstead.
A restart posts a ⚠️ Gateway shutting down notice into Slack. That's expected, not a
fault — but it means restarts are user-visible, so batch them.
Then re-send the same test message and compare api_calls and time in agent.log.
| Query type | Skill used | Calls | Time | Status |
|---|---|---|---|---|
| Infra status (all services + containers) | infrastructure | 3 | ~132s | Working — uses /summary, concise |
| Weather forecast (weekend) | weather | 2 | ~66s | Working |
| Weather UV query (sunscreen) | weather | 2 | ~60s | Working |
| TickTick overdue/due tasks | tasks | 3 | ~152s | Working — uses /summary .ticktick, grouped by project |
| Calendar + unread emails | schedule | 3 | ~77s | Working — parallel calendar + emails in one pass |
| Slack unreads | slack | 3 | ~31s | Working — uses /slack/unreads, very fast |
| Cross-domain (meetings + tasks + weather) | schedule + tasks + weather | 4 | ~121s | Working — loads 3 skills, 3 curl calls, clean combined response |
| Capture: clear GitHub repo (slack patch → dotfiles) | capture | 2 | ~12s | Working — 95% confidence, correct title + body shape |
| Capture: TickTick Personal (water plants tomorrow) | capture | 2 | ~11s | Working — date math correct, project ID resolved |
| Capture: IU/Work override (EP-1234 engineering task) | capture | 2 | ~9s | Working — hard rule applied, no GitHub even for code work |
| Capture: Shopping (buy oat milk and bananas) | capture | 2 | ~9s | Working — split into 2 tasks per multi-item rule |
| Capture: GitHub repo by name (jkrumm.dev OG tags) | capture | 2 | ~13s | Working — 99% confidence, correct repo ID |
| Capture: Ambiguous (look into morning briefing) | capture | 2 | ~9s | Working — asks clarifying question, no write |
| Capture: Multi-item (Tailscale cert + watchdog issue) | capture | 2 | ~10s | Working — TickTick Personal + GitHub dotfiles in one response |
| Capture: Cache miss + refresh (snow-finder) | capture | 6 | ~27s | Working — read empty cache, ran gh repo list jkrumm, confirmed repo, routed |
| KaraKeep: keep a link (karakeep repo) | karakeep | ~2 | — | Working — POST /bookmarks, returns ID, flags async crawl/AI-tag (verified: tags applied) |
| Obsidian: search vault (north-star) | obsidian | ~2 | — | Working — obsidian CLI search+read, summarised the real note content |
| Reading: "was soll ich lesen" | reading | — | ~34s | Working — pulled /api/reading shelf, German rec grounded in the real shelf (Tress / Bad Karma / Iron Flame) |
| WalkingPad: weekly distance | argo-api (walking-pad) | — | ~26s | Working — /walking-pad/sessions/summary + heroes; 33.7 km / 12 sessions / streak / +30.5% wk, all correct |
| Usage: monthly AI spend | argo-api (usage) | — | ~64s | Working — /usage/headline + breakdown, per-source + per-model, honest about fixed 7/30/90d windows |
| Research: latest Bun version | research-gateway | — | ~54s | Working — submit+poll to research-gateway, cited answer (Bun v1.3.14); tirith allowlist let curl research… | jq through with no approval gate |
| Post-update (v0.16.0→v0.18.2) TickTick overdue check | tasks (method A) | 4 | ~26s | Working — general agent loop intact after platform-plugin rewrite |
| Post-update infra status (homelab+vps) | infrastructure (method A) | 4 | ~16s | Working — 2× terminal/curl tool calls, zero tirith approval gates (confirms tirith-hermes-guards.patch re-applied correctly) |
Post-update Slack formatting check (* bullet) | — (method B, real Slack path) | 1 | 5.2s | Working — response posted with - bullets (confirms format_message() pre-steps ported to plugins/platforms/slack/adapter.py). Recorded thread_ts: null — correct then (reply_in_thread: false), and the opposite of correct now; see the threading note below. |
| Post-update TTS title check ("say out loud") | — (method B, real Slack path) | 2 | 16.2s | Working — tools.tts_tool: TTS audio saved: .../Update Verification Successful.mp3 (confirms tts-tool-audio-title.patch re-applied correctly post-conflict, not tts_<timestamp>.mp3); media send via patched base.py completed with no errors |
| Post-audit infra status (method A) | argo-api (infrastructure) | 3 | 15.4s | Working — Homelab 36/36 + VPS 29/29, zero approval gates (argo allowlist intact after the tirith rewrite) |
| Post-audit download-guard block (method A) | tirith_security hardening | — | — | Working — wget -qO /tmp/f URL && chmod +x … && /tmp/f blocked end-to-end, no file created. This shape was a bypass before the restart, so it proves the running process carries the hardened guard, not just the file on disk |
| Post-audit Slack round-trip (method B) | reply_in_thread: true | 3 | 21.3s | Working — session key agent:main:slack:group:<team>:<C…>:<ts> carries the message ts, reply threaded under it. First live confirmation that one thread == one session == one context window |
Post-audit thread continuation (method B, /reply) | reply_in_thread: true | 1 | 3.9s | Working — reply into the thread reused the same session (message_count 8→10, no new session) and correctly recalled the first question. Confirms context carries within a thread; inbound arrives prefixed [Replying to: "…"] |
| v0.19.1 Slack bullets + threading (method B) | — | 2 | 10.9s | Working — session key …:C0ASRUD7K1U:1785673823.484699 with thread_id set. Note the model happened to emit - bullets itself, so the round-trip alone proved nothing about the patch — see the in-process recipe below |
v0.19.1 format_message() pre-steps (in-process) | slack-cannot-reply-to-message.patch | — | — | Working — * → - at all nesting levels, and `:white_check_mark: …` backticks stripped. This is the check that actually pins the patch |
v0.19.1 TTS + thread continuation (method B, /reply) | tts-tool-audio-title + base.py anchor | — | ~29s | Working — TTS audio saved: …/Gateway und Dienste prüfen.mp3 (human title, not tts_<timestamp>), [Slack] Delivering 1 non-image MEDIA attachment(s) then a clean send, no errors. Same session reused (mc 5→9) |
| v0.19.1 STT round-trip (in-process) | stt.openai → audio-gateway | — | — | Working — fed the TTS mp3 back to transcribe_audio(): gpt-4o-transcribe returned accurate German with umlauts intact ("Prüfe Gateway und Dienste… keine Fehler in der Gateway-Logdatei"). Closes the loop Gemini Charon → gpt-4o-transcribe |
| v0.19.1 download-guard (method A + in-process) | tirith-hermes-guards | 1 | — | Working — the agent's wget -qO … && chmod +x … && exec chain was cut: the file landed but was never +x'd and never ran. In-process verdicts confirm why (chain block, bare download allow), and the gateway (pid started 14:22:42) postdates the module write (14:17:16) |
| v0.19.1 cron delivery target (in-process) | scheduler-skip-resolver-for-slack-ids | — | — | Working — slack:C0ASRUD7K1U resolved to {'chat_id': 'C0ASRUD7K1U', 'thread_id': None}, un-mangled by the channel directory |
In-process recipe for format_message() — worth keeping, because two things block the
obvious approach: the module uses a relative import (from .block_kit import …) that fails
outside its package, and SlackAdapter's base is abstract. Add the adapter's own directory
to sys.path so the module's except ImportError fallback resolves, then subclass to
satisfy the abstract methods:
import importlib.util, sys
sys.path.insert(0, 'plugins/platforms/slack') # lets `from block_kit import …` resolve
spec = importlib.util.spec_from_file_location('slack_probe', 'plugins/platforms/slack/adapter.py')
mod = importlib.util.module_from_spec(spec); sys.modules['slack_probe'] = mod
spec.loader.exec_module(mod)
class Stub(mod.SlackAdapter): # base is abstract — stub the 4 methods
def __init__(self): pass
async def connect(self): ...
async def disconnect(self): ...
async def get_chat_info(self, *a, **k): ...
async def send(self, *a, **k): ...
print(mod.SlackAdapter.__dict__['format_message'](Stub(), '* eins\n* zwei'))
A Slack round-trip cannot prove a formatting patch. The model writes whatever markdown
it likes; if it already emits - bullets, the output is identical with or without the
patch. Drive format_message() directly with input that must be transformed, and keep
the round-trip for what only it exercises — threading, session keying, media delivery.
Update this table after each validation run. Rows are historical, not current spec —
they record what was true on the day. The Skill used column on pre-2026-06 rows names
domain skills (infrastructure, tasks, schedule, weather, slack) that no longer
exist as directories; those endpoints now live in argo-api/references/*.md and route
through the argo-api skill. Don't "fix" an old row to match today — add a new one.
Verifying a security control needs a shape that only the NEW code catches. A block that the old code also blocked proves nothing about which version is loaded. Pick a case from the regression suite's fix list, and confirm the side effect too (here: the downloaded file must not exist). Python modules are imported once at gateway start — an in-process test result says nothing about the running process.
All paths are relative to ~/SourceRoot/hermes-agent/ unless noted. The skill set is
whatever HERMES_SKILLS in the Makefile lists — check there rather than trusting this
table, which has gone stale before.
| File | Purpose |
|---|---|
SOUL.md | System prompt — skill routing table lives here |
config.yaml | Model, platform (reply_in_thread), secrets, tirith settings |
skills/argo-api/SKILL.md | Full endpoint reference + references/*.md per domain (infrastructure, tasks, schedule, weather, slack, garmin-health, strength, walking-pad, usage) |
skills/capture/SKILL.md | Router: KaraKeep vs Obsidian vs TickTick vs GitHub |
skills/work/SKILL.md | IU surface — M365, Jira, Confluence, GitLab |
skills/karakeep/SKILL.md · skills/obsidian/SKILL.md | Read-later bucket · second brain |
skills/reading/SKILL.md · skills/research-gateway/SKILL.md | Book shelf · deep cited research |
skills/image-delivery/SKILL.md | imgcli share/publish — private by default |
tests/test_download_guard.py | Regression suite for the local tirith hardening |
~/.hermes/logs/agent.log | Structured run log (api_calls, time, inbound messages) |
~/.hermes/logs/gateway.log | Gateway stdout (startup, tool progress bars) |
~/.hermes/sessions/*.jsonl | Turn-by-turn traces — Slack path only (see method-A note) |
~/.hermes/state.db | sessions table — session keys, thread ids, message counts |
~/.hermes/gateway_state.json | Live pid + gateway/Slack connection state |