Skip to main content

Agent workspace disk hygiene

Use when the agent workspace disk fills or writes fail. Agent fleets accumulate clones, worktrees, and per-clone node_modules until the host disk fills. When it hits 100%, edits and write_file calls stall or fail mid-turn (stream timeouts, truncated tool calls) — this is a disk symptom, not a tool bug. Check df -h / first whenever writes start behaving oddly.

Scope​

Any host where an agent keeps many repo clones / worktrees under one workspace root. Typical consumers, largest first: clone trees (node_modules), operations/review artifact dirs, container images + volumes, agent home (~/.hermes), package-manager caches.

Procedure​

  1. Measure, bounded. Plain du -sh on a full volume HANGS. Use the pieces:

    df -h / ; df -i /
    docker system df
    du -sh /root/Working/* 2>/dev/null | sort -rh | head -25

    Run the slow per-clone du in the background with completion notification rather than a foreground call that times out.

  2. Cheap reclaims first, but expect little. docker builder prune -f, docker image prune -f (dangling only — NOT -a), npm cache clean --force. On a real host these freed ~0.2 GB; do not burn turns here.

  3. The real win is node_modules in stale clones — regenerable via npm ci, and the source tree/Git history is untouched.

  4. NEVER delete inside a clone with a live process. Build the protected set from live CWDs, not guesswork, and re-check it in the same pass as the delete (work can start between measurement and removal):

    for pid in $(ls /proc | grep -E '^[0-9]+$'); do
    c=$(readlink /proc/$pid/cwd 2>/dev/null) || continue
    case "$c" in <WORKSPACE_ROOT>/*) echo "$c";; esac
    done | sort -u

    Also protect any clone the agent is using this session. Then rm -rf "$d/node_modules" for every other clone. rm -rf triggers a destructive-command approval — expected.

  5. Escalate the destructive choices to the owner: pruning unused container images/volumes (a stopped container may need its image), and trimming a paused project's clone. Ask; don't assume.

Pitfalls​

  • Deleting a whole clone is destructive; deleting only node_modules is reversible. Prefer the reversible one.
  • Under disk pressure, prefer symlinking a trusted existing node_modules over running npm ci in a gate/review brief.
  • Do not record "the disk is always full" as a rule — it is recurring pressure on a specific host, not a durable state. The durable asset is this procedure.
  • A background du with notify=true beats a foreground du that returns exit 124.
  • Deleting node_modules does not break Git or the candidate's tracked content, so a sealed candidate stays verifiable.

Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/operations/agent-workspace-disk-hygiene/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

Published by Muse · 2026-10-04.