Skip to main content

Diagnosing a failure that will not reproduce

Use when a failure will not reproduce under probes. A recurring crash, hang, or corruption is observed in production but every probe you construct comes back clean. The deliverable for this class of task is not a cause — it is a rigorously bounded negative: what you eliminated, what survived elimination, and the failure's unchanged signature.

When to Use​

  • A service crashes intermittently (SIGSEGV, OOM, silent exit) and controlled runs pass.
  • A failure appears in host logs but no reduced workload triggers it.
  • You are tempted to name a co-occurring anomaly as "the cause".
  • Someone asks you to explain a fault you could not reproduce.

The rule​

Report the eliminated set, never a plausible cause you did not isolate. An unreproduced hypothesis stated as a finding is worse than an honest negative: it closes the investigation on a guess, and the next session inherits it as fact.

Equally, do not end with nothing. "Could not reproduce" alone wastes the probe runs you already paid for. The negative is only useful when it is bounded.

Procedure​

1. Capture the signature before probing​

Extract the invariant parts of every historical occurrence: instruction pointer vs fault address, offset, stack pointer, error code, exit status, the log line immediately preceding. A signature that is identical across all occurrences is positive evidence about the failure's nature even while the trigger is unknown — e.g. ip == fault address with error 15 is a jump into an unmapped page, which says a native library was unmapped before teardown completed. Say what the signature implies; that is a real finding.

Count the historical occurrences and check whether they predate a recent change (a model rebind, a dependency bump, a config edit). Occurrences on both sides of a change eliminate that change without running anything.

2. Build arms, one per suspected cause​

Design the probe set so each arm removes exactly one variable, and include at least one arm that reproduces the production configuration exactly — same toolset, same flags, same environment. Without that arm the whole negative is unanchored.

Run each arm enough times to be meaningful against the observed base rate, and run the cheap arms first. Typical shape:

  • A — minimal path, no external calls, high run count.
  • B — real workload, minimal dependencies/toolset.
  • C — real workload, exact production configuration.
  • D..n — real workload plus one suspected environmental cause each.

Also probe the bare primitive when a native stack is implicated (import the libraries and exit, many iterations); a clean result there moves the fault from load/teardown of the libraries to the workload that drives them.

3. Report as a probe table​

ProbeConfigurationRunsFaults
Aminimal path, no inference200
Breal workload, minimal toolset60
Creal workload, exact production toolset60
Dreal workload, against the suspected corrupt state50

Then state the eliminated set explicitly and, separately, that it reproduces only under the full workload. Reducing the workload is what gives the negative its meaning — name which reductions you tried.

4. Bind every observed fault to a probe window​

Host-global logs (dmesg -T, journal, syslog) collect faults from every process on the machine. A concurrent investigation — another agent session, a cron job, a sibling lane — running its own probes will drop faults into the same log, and reading them as yours inverts the conclusion.

Record each arm's start and end time, then compare every fault timestamp against those windows. Faults outside your windows are contamination: disclose them explicitly and count them in neither direction. Say who else was probing if you can identify them.

5. Do not promote a co-occurring anomaly to the cause​

Investigations routinely surface an unrelated broken thing in the same window — a corrupt datastore, a failing sibling service, an alarming log line. It is a separate finding, reported separately, until an arm running against it actually faults. Probe it as arm D rather than asserting it.

6. Report the bigger finding when the crash is not it​

A crash investigation frequently turns up a more consequential defect than the crash. Counters that never reset, backoff pinned at its ceiling, work delivered successfully by every "failed" pass — these change what the owner should fix first. Check whether the crashed runs actually lost work (look for their output artifacts, receipts, and downstream side effects); a pass that crashed after delivering is a cosmetic failure, and saying so redirects the effort.

When you find it, lead with it. The owner asked about the crash, but they want the system working.

Pitfalls​

  • Naming a cause to look conclusive. State "not isolated" plainly and put the eliminated set next to it; a bounded negative is a real result.
  • A single clean arm read as elimination. One run proves nothing against an intermittent fault — match run counts to the observed rate.
  • Missing the exact-production arm. Reduced arms that all pass tell you nothing without an arm that matches production and also passes.
  • Counting host-global faults without windowing. See step 4 — this silently fabricates evidence.
  • Treating an exit code as the classification. A runtime that sleeps on a rate limit surfaces capacity events as timeouts; classify from the logs and the session record, not the exit status.
  • Reporting probe counts without saying what they covered. "32 runs, 0 faults" is meaningless until the reader knows which variables those runs held fixed.

Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/engineering/non-reproducing-failure-diagnosis/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

version 1.0.0 · author Hermes Curator · license MIT.

Published by Muse · 2026-10-04.