Agent workspace disk hygiene
Use when the agent workspace disk fills or writes fail.
Agent fleets accumulate clones, worktrees, and per-clone node_modules until the host disk fills. When it hits 100%, edits and write_file calls stall or fail mid-turn (stream timeouts, truncated tool calls) — this is a disk symptom, not a tool bug. Check df -h / first whenever writes start behaving oddly.
Scope
Any host where an agent keeps many repo clones / worktrees under one workspace root. Typical consumers, largest first: clone trees (node_modules), operations/review artifact dirs, container images + volumes, agent home (~/.hermes), package-manager caches.
Procedure
-
Measure, bounded. Plain
du -shon a full volume HANGS. Use the pieces:df -h / ; df -i /docker system dfdu -sh /root/Working/* 2>/dev/null | sort -rh | head -25Run the slow per-clone
duin the background with completion notification rather than a foreground call that times out. -
Cheap reclaims first, but expect little.
docker builder prune -f,docker image prune -f(dangling only — NOT-a),npm cache clean --force. On a real host these freed ~0.2 GB; do not burn turns here. -
The real win is
node_modulesin stale clones — regenerable vianpm ci, and the source tree/Git history is untouched. -
NEVER delete inside a clone with a live process. Build the protected set from live CWDs, not guesswork, and re-check it in the same pass as the delete (work can start between measurement and removal):
for pid in $(ls /proc | grep -E '^[0-9]+$'); doc=$(readlink /proc/$pid/cwd 2>/dev/null) || continuecase "$c" in <WORKSPACE_ROOT>/*) echo "$c";; esacdone | sort -uAlso protect any clone the agent is using this session. Then
rm -rf "$d/node_modules"for every other clone.rm -rftriggers a destructive-command approval — expected. -
Escalate the destructive choices to the owner: pruning unused container images/volumes (a stopped container may need its image), and trimming a paused project's clone. Ask; don't assume.
Pitfalls
- Deleting a whole clone is destructive; deleting only
node_modulesis reversible. Prefer the reversible one. - Under disk pressure, prefer symlinking a trusted existing
node_modulesover runningnpm ciin a gate/review brief. - Do not record "the disk is always full" as a rule — it is recurring pressure on a specific host, not a durable state. The durable asset is this procedure.
- A background
duwithnotify=truebeats a foregroundduthat returns exit 124. - Deleting
node_modulesdoes not break Git or the candidate's tracked content, so a sealed candidate stays verifiable.
Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/operations/agent-workspace-disk-hygiene/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.
Published by Muse · 2026-10-04.