Skip to main content

Orchestrating Research

Use when a user asks for research, deep research, benchmark or standards analysis, competitive investigation, literature review, evidence-backed recommendations, or CIS benchmark automation research.

Purpose​

Own the research mission while keeping collection, compilation, and review as separate roles. A final answer is not ready until material claims have primary-source citations, the research artifact passes deterministic Level 1 validation, and a fresh rigor-reviewer session has issued a verdict.

Capability Selection​

Load grounded-citations for every mission. Then load only what fits:

  • arxiv for academic literature.
  • competitor-news-monitor for company or market monitoring.

For CIS, Microsoft 365, Azure, AWS, or other benchmark automation, prioritize the licensed benchmark, official platform documentation, source code, API schemas, and controlled experiments. Distinguish benchmark interpretation from platform-state collection and from evaluator logic.

Mission Contract​

Before dispatch, write a short contract containing:

  • research question and decision the result must support;
  • scope, exclusions, current-version cutoff, and jurisdiction or tenant assumptions;
  • required deliverables and output directory;
  • primary-source hierarchy;
  • claims that require experiments rather than documentary support;
  • acceptance criteria, including unresolved questions that must remain explicit.

For substantial work, create /root/Working/Research/<task-slug>/ and keep sources, working notes, the compiled ara/, validation receipts, and review reports there.

Workflow​

  1. Decompose by evidence boundary. Use independent lanes when questions can be researched without shared mutable state. Examples: benchmark interpretation, platform API feasibility, licensing, competing offerings, implementation architecture, and security/privacy.
  2. Collect primary evidence. Cite the source that owns each fact. Preserve exact title, URL or document ID, section, retrieval date, and relevant quote or extracted field. Secondary sources may identify leads but do not close material claims.
  3. Keep an evidence ledger. Separate observed facts, source-backed interpretation, assumptions, hypotheses, and unresolved conflicts. Never convert missing access, missing licensing, or collection errors into a negative finding.
  4. Synthesize. Reconcile contradictions, state confidence, identify failed or blocked approaches only when evidence supports that classification, and retain counterevidence.
  5. Preflight role profiles. Run hermes profile show artifact-compiler and hermes profile show rigor-reviewer. If either is unavailable, stop or use explicitly separate fresh delegated roles; never collapse authoring and review. Treat nonzero subprocess exits, timeouts, missing files, or malformed receipts as infrastructure failures, not research results.
  6. Compile through the role profile. Start a new artifact-compiler session:
hermes -p artifact-compiler chat -q "Compile the research packet at <absolute-input-path> into an ARA at <absolute-output-path>/ara. Select and declare artifact_type documentary|benchmark|experimental|software. Run deterministic Level 1 validation and write the receipt to <absolute-output-path>/ara/level1_receipt.json."
  1. Require Level 1. Read the receipt. Require valid: true, an empty errors list, and a non-empty artifact_digest. Do not invoke semantic review on an invalid artifact. Route structural failures back to the artifact-compiler.
  2. Review in a fresh session. Never resume the compiler session. Start the rigor-reviewer profile with only the validated artifact path and bounded source packet:
hermes -p rigor-reviewer chat -q "Perform a fresh-context, artifact-only Level 2 review of <absolute-output-path>/ara. Level 1 receipt: <absolute-output-path>/ara/level1_receipt.json. Write <absolute-output-path>/ara/level2_report.json and repeat the Level 1 artifact_digest in the report."
  1. Remediate through the coordinator. Apply specific fixes or dispatch them to the compiler. After two failed remediation attempts on one issue, obtain an independent escalation-review plan before another change.
  2. Record the epilogue before sealing. Use recording-research-epilogues to preserve decisions, experiments, dead ends, pivots, open threads, and provenance. Because this mutates the ARA, rerun Level 1 afterward.
  3. Seal the final bytes. Invoke a new rigor-reviewer session against the post-epilogue Level 1 receipt. Then run the coordinator-bundled validator:
python3 <orchestrating-research-skill-directory>/scripts/validate_level2_report.py \
<absolute-output-path>/ara/level2_report.json \
--receipt <absolute-output-path>/ara/level2_validation_receipt.json

Require both receipts to be valid and to report the same artifact digest. Any later artifact mutation invalidates the seal and requires both gates again. 12. Deliver. Lead with conclusions and decisions. Include artifact path, digest, both validation statuses, review grade, material limitations, and the next action.

CIS Benchmark Automation Contract​

A CIS automation research result must distinguish:

  • exact benchmark version and profile;
  • CIS licensing and redistribution constraints;
  • control interpretation and applicability;
  • collection interface and least privilege;
  • deterministic evaluator predicate;
  • pass, fail, manual, not_applicable, unsupported, error, and inconclusive states;
  • known-good and known-bad fixtures;
  • raw evidence provenance and content hashes;
  • false-pass prevention;
  • benchmark/API/schema update strategy;
  • read-only assessment from separately authorized remediation.

Never claim a control is automatable merely because an API returns a related field. The returned state must be sufficient to evaluate the exact benchmark condition under documented semantics.

Completion Gate​

A substantial research mission is complete only when:

  • material claims map to primary evidence;
  • uncertainties and contradictory evidence remain visible;
  • an ARA exists and Level 1 is valid;
  • a fresh-context Level 2 review exists;
  • critical and major findings are fixed or explicitly accepted by the owner;
  • the final artifact and receipts are preserved at stable paths.

Supporting files: this skill's supporting files are held in the docsite at docs/15-skills/_support/research/orchestrating-research/ — fetch them fresh from jknash/docsite main alongside this page. Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/research/orchestrating-research/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

version 1.0.0 · author jknash · license MIT.

Published by Muse · 2026-10-04.