Skip to main content

Compiling Research Artifacts

Use when papers, repositories, experiment logs, benchmark mappings, notes, or source packets must become a structured, falsifiable, machine-traversable Agent-Native Research Artifact.

Contract​

Compile supplied research inputs into a complete Agent-Native Research Artifact (ARA). Use Hermes file, document, web, repository, and terminal tools as appropriate. Do not expose private reasoning; make the result auditable through claims, evidence, provenance, and validation receipts.

Never invent missing content. Write “Not available from provided input” when a required field cannot be grounded, and list the gap in the final report.

Inputs​

Accept PDFs, arXiv links, repositories, code, notebooks, configuration, experiment logs, benchmark mappings, meeting notes, source ledgers, or a directory combining them. Identify the inputs before generating files. Read appendices and supporting references when they affect reproduction, limitations, or claims.

Extraction Protocol​

  1. Inventory sources. Record stable path or URL, title, publisher/author, version or commit, retrieval date, and content hash where possible.
  2. Declare artifact type. Put one type in PAPER.md frontmatter:
    • documentary — market, competitive, literature, policy, or standards research supported directly by source evidence;
    • benchmark — benchmark/control interpretation, collectors, evaluator predicates, fixtures, and automation feasibility;
    • experimental — hypotheses and empirical experiments, without necessarily shipping implementation code;
    • software — research whose claims depend on a concrete implementation or repository.
  3. Extract knowledge atoms. Capture formulations, architecture, exact configurations, exact numerical results, citations, negative findings, limitations, and implementation constraints.
  4. Capture evidence before claims. Preserve original table/figure identifiers and captions. Store faithful raw transcriptions separately from derived subsets. Never label a filtered or merged view as the original table.
  5. Build the common cognitive layer. Every type has observations, gaps, assumptions, falsifiable claims, concepts, related work, an exploration graph, and grounded evidence.
  6. Build type-specific layers. benchmark, experimental, and software artifacts include experiments and solution files. benchmark and software artifacts include typed execution stubs. A documentary artifact binds claim Proof fields directly to evidence/*.md; never fabricate experiments or code to satisfy a template.
  7. Build the exploration graph. Include questions, decisions, and only source-supported experiments, pivots, or dead ends. Every node has support_level: explicit|inferred; explicit nodes have source_refs. Inferred reconstruction must not masquerade as session history.
  8. Coverage loop. Re-read the source packet and patch omissions, up to three rounds. Stop early when no gap remains.
  9. Validate. Run the bundled deterministic validator:
python3 <skill-directory>/scripts/validate_ara.py <artifact-dir> \
--receipt <artifact-dir>/level1_receipt.json

Fix all structural failures and rerun. A prose assertion that validation passed is not a receipt.

Required Layout​

Every artifact has:

PAPER.md # frontmatter includes artifact_type
logic/{problem,claims,concepts,related_work}.md
trace/exploration_tree.yaml
evidence/README.md
evidence/tables/*.md and/or evidence/figures/*.md and/or evidence/sources/*.md
level1_receipt.json

benchmark, experimental, and software add:

logic/experiments.md
logic/solution/{architecture,algorithm,constraints,heuristics}.md
src/configs/{training,model}.md
src/environment.md

benchmark and software also require src/execution/*.py. Omit non-applicable type-specific files from documentary artifacts rather than filling them with invented content.

Binding Rules​

  • Claims use Cxx; experiments use Exx; heuristics use Hxx; trace nodes use Nxx.
  • Every claim states status, falsification criteria, direct evidence basis, interpretation, dependencies, and tags. benchmark, experimental, and software Proof fields bind to experiment IDs; documentary Proof fields bind directly to contained evidence/*.md paths.
  • Every experiment identifies verified claims, setup, procedure, metrics, directional expected outcome, baselines, and dependencies.
  • Exact observed numbers belong in evidence, not expected outcomes.
  • Every heuristic has rationale, sensitivity, bounds, a resolving code reference, and source.
  • Claim scope must not exceed evidence scope.
  • Collection failure, denied access, null, empty, not applicable, and negative findings remain distinct.

Evidence Fidelity​

A quantitative raw evidence file in tables/ or figures/ must reproduce one source object faithfully and include Source, caption or identifier, and a Markdown table. A documentary source file in sources/ includes Source, the exact quoted or extracted passage, and its location. A derived view must be named as derived/subset and point to its raw parent. Preserve exact values; do not round unless the source itself does.

Report​

Return:

  • artifact path;
  • source count and source boundaries;
  • file and entity counts from the Level 1 receipt;
  • unresolved or unavailable fields;
  • receipt path and validation status;
  • exact artifact hash manifest if generated.

Do not perform the independent Level 2 verdict. The coordinator sends the validated artifact to a fresh rigor-reviewer session.


Supporting files: this skill's supporting files are held in the docsite at docs/15-skills/_support/research/compiling-research-artifacts/ — fetch them fresh from jknash/docsite main alongside this page. Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/research/compiling-research-artifacts/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

version 1.0.0 · license MIT.

Published by Muse · 2026-10-04.