Skip to main content

Web Research via Curl

Fetch/verify web facts via curl when no web_search tool. A working pattern for gathering verifiable, primary-source web facts using terminal + curl when the session has no web_search / web_extract tool configured (this happens in subagent and some headless environments). It also improves any citation workflow by favoring direct primary sources and checking that every cited URL actually resolves before it is trusted.

Companion skills: pair with grounded-citations for the citation/ledger management once you have fetched real sources. This skill is about the retrieval and verification mechanism; grounded-citations owns the numbering and evidence chain.

When to Use​

  • The task asks for current (2024–20XX) verifiable facts and your toolset has no web_search/web_extract/browser tools — only terminal.
  • You must confirm a source URL actually exists (HTTP 200) before citing it.
  • You want primary/vendor/.gov docs instead of a search-snippet summary.

When the main session has NO web tools at all: delegate parallel research​

If your own toolset has no web_search/web_extract/browser AND you'd rather not hand-curl every page, the fastest path to a broad, well-sourced report is to fan out parallel research subagents via delegate_task. Each subagent gets its own toolset (it can curl / render / search itself), and you orchestrate the synthesis. This pattern produced a 37-source, 10-offering report in ~11 minutes of wall-clock research.

  • Issue one delegate_task(goal=..., context=...) call per research stream, in a single message so they run concurrently. Do NOT pass a single-element tasks array — batch mode rejects it with "Batch mode requires at least 2 tasks." For N independent streams, make N separate goal-based calls in one turn (up to the delegation.max_concurrent_children limit).
  • Give each subagent a self-contained brief: the exact questions to answer, the output persona/format, a required per-claim source format (Claim — [Title](URL), publisher, date), and an explicit "do not fabricate URLs or numbers; mark unverifiable items [unverified]" instruction. Children know nothing of your conversation — pass everything they need.
  • Split by concern, not by source. A proven split for a market/offerings report: (1) market landscape & statistics, (2) competitor offerings, (3) emerging solution patterns. Each returns a structured markdown report with its own source list.
  • Have subagents write their report to a file (/root/.../report.md) so you can read_file it after they finish, rather than relying on the summary alone.
  • Register every returned URL in the citation ledger (grounded-citations sources.py add) before writing prose, then synthesize with inline [n] citations and render the Sources block mechanically. See references/research-report-pipeline.md for the full end-to-end worked example (citation-integration pitfalls + PDF export).

Search engines (DuckDuckGo HTML, Bing) are unreliable for scripting: they rate-limit rapid consecutive requests and serve anti-bot "challenge" pages. When a search does return real results, the snippet only supports what it literally says — go fetch the underlying page and extract from its text.

For a report the fastest, most trustworthy reads come from the authoritative page itself: vendor product pages, official docs (.docs./docs. subdomains), GitHub repos, and .gov / official-EC pages. Go there directly when you know the canonical URL; use search only to discover URLs you don't yet know.

Procedure​

① Probe connectivity once so a 0-byte / network-blocked curl isn't misread as "page is empty":

curl -s -m 15 -o /dev/null -w "%{http_code}" https://www.google.com

② Fetch each page to a file with a browser UA string (many sites 403 or serve JS-shells to curl's default UA):

ua="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0 Safari/537.36"
curl -s -m 25 -A "$ua" -L "https://example.com/page" -o research/page.html
wc -c research/page.html # sanity-check size

③ Extract readable text with scripts/html_to_text.py (handles the <script>/<style> noise, tag stripping, and HTML-entity unescaping):

python3 ~/.hermes/skills/research/web-research-via-curl/scripts/html_to_text.py research/page.html

④ Verify each URL you plan to cite returns a real page before including it:

code=$(curl -s -o /dev/null -w "%{http_code}" -m 20 -A "$ua" -L "https://example.com/page")
echo "$code <- https://example.com/page"

Treat anything other than 2xx/the odd 3xx as unverified. Two URLs in this session's source came back 404 on direct fetch even though they appeared in live search results — search-hit ≠ valid link. Always check before citing.

⑤ Extract the exact supporting sentence from the fetched text and copy it verbatim into your notes so a grounded-citations quote step can match it later, or so your report quotes the page accurately.

⑥ Label honestly. A claim you could not retrieve stays [unverified] in the report. JS-gated (no body in the HTML) and 403'd pages count as unverified for anything you could not actually read. Say which sources were JS-gated/403 rather than silently citing them.

Escalation ladder when curl is blocked (403 / JS-shell / HTTP2 errors)​

Many authoritative .gov/.com pages 403 curl's UA outright (Gartner → 403, McKinsey → net::ERR_HTTP2_PROTOCOL_ERROR), and in-browser/headless-rendered searches are needed to discover URLs. Take this order — only escalate when the current step fails:

  1. Direct primary fetch via curl (Procedure above). Many .gov/vendor pages work fine this way (Federal Reserve, JPMorganChase Institute responded to curl). Sanity-check wc -c (a ~5–30KB body is often a JS shell, not the article).
  2. Headless browser render — for JS-gated pages and to run searches in a real browser engine. Install Playwright + a Chromium engine, and do not skip the OS-library step (it resolves missing libnspr4/NSS shared libs and yields TargetClosedError on launch otherwise):
    pip install playwright
    python3 -m playwright install chromium
    python3 -m playwright install-deps chromium # "error while loading shared libraries: libnspr4.so..." → run this
    Then render + extract with scripts/fetch_playwright.py (launch args --no-sandbox --disable-blink-features=AutomationControlled, set a full Chrome UA, wait_until="domcontentloaded", wait a few seconds, then inner_text("body")).
  3. Search via Startpage in the browser — when Bing/DDG/Mojeek/Brave all serve anti-bot challenge shells to curl and to raw HTML parsing, Startpage (https://www.startpage.com/sp/search?query=...) proxies Google results and renders them as plain anchor tags in a headless browser. It was the only search engine that returned real result links out of DDG/Bing/Mojeek/Brave tested. Use scripts/search_startpage.py "<query>" to run it and print result links; then fetch the primary page directly for the facts. Extract with: a elements where href starts with http and filters out the engine's own domains.
  4. Reader proxy for hard-blocked sites — prefix any URL with https://r.jina.ai/<full-url> to get a Markdown/plain-text version of pages that 403/bounce even the browser (used successfully on McKinsey). It returns Title / URL Source / Published Time / Markdown Content — ideal verbatim source capture. Note it may still truncate very long pages.

When primary pages AND search are both blocked (vendor 403 / anti-bot)​

Enterprise/vendor sites (Infosys, TCS, McKinsey, Wipro press, Accenture newsroom) frequently 403 or 404 automated curl clients behind Akamai/CDN bot walls — even with a browser UA. Don't stop; there are three reliable fallback paths that worked together:

  • Bing RSS as a discovery search. When Bing regular HTML and DDG (both html.duckduckgo.com and lite.duckduckgo.com) return empty anti-bot shells, use Bing's RSS endpoint — it returned real, dated results while the HTML search returned nothing:
    curl -s -m 30 -A "$ua" \
    "https://www.bing.com/search?format=rss&q=<urlencoded>&count=8"
    Parse <item><title>/<link>/<pubDate>/<description> from the feed. Caveat: it often surfaces only navigational hits (homepages, Wikipedia, LinkedIn) for wordy queries — use it to discover a canonical URL, then fetch that page directly. grep b_algo still finds nothing on RSS; parse the <item>s.
  • Wayback CDX API to recover 403'd vendor pages. The "closest snapshot" convenience API (https://archive.org/wayback/available?url=...) returns None even when archives exist, and live vendor URLs 403. The authoritative lookup is the CDX API with prefix match + status filter:
    curl -s -m 40 \
    "https://web.archive.org/cdx/search/cdx?url=example.com/path&matchType=prefix&output=json&limit=3&filter=statuscode:200&collapse=timestamp:6"
    → returns [["urlkey","timestamp","original",...]] rows; reconstruct the snapshot as https://web.archive.org/web/<timestamp>/<original>. Add &collapse=digest to dedupe near-identical captures, or &collapse=timestamp:6 for one-per-day. If CDX returns [], it genuinely has no 200 capture — don't loop retrying.
  • Wikipedia as a fallback verifier for blocked vendor claims. When the vendor page 403s, a current Wikipedia article (e.g. on the parent company or product) often independently confirms the service/offering exists. Cite Wikipedia explicitly and downgrade vendor-specific numbers (investment $, headcount, launch dates) that only appeared in un-fetchable press to [unverified].
  • Verify entity identity, not just existence. For competitors/boutiques in a rumor mill, fetch the named company's page and confirm it actually does what was claimed. In one competitive pass "Andium" turned out to be an industrial-IoT/emissions hardware vendor, and "Latent Space" an AI-engineer newsletter — neither was the AI consultancy they were rumored to be. Calling that out beats copying a mis-categorized competitor.

Pitfalls​

  • Vendor 403 ≠ claim dead. The page is blocked, not necessarily gone. Try the Wayback CDX path and Wikipedia before marking the claim unverified.

  • wayback/available under-reports. Prefer the CDX API; the convenience endpoint returned None in this session for URLs CDX had captures of.

  • Search engines may choke on multi-word queries that contain a common substring. E.g. "mid-market genAI adoption" → Bing/DDG keyed on the slang word "mid" and returned dictionary/irrelevant results. If a query yields off-topic hits, reword (drop the ambiguous token, or add an unambiguous qualifier) rather than assuming "no results."

  • Startpage via curl won't parse. startpage.com served a 403 to curl with a default UA and a bare 363-byte body; it only works through a real browser engine (use the Playwright search script, not curl).

  • Rapid consecutive searches → anti-bot page. DDG/Bing return ~14KB of no-results (a challenge/consent shell) after a few fast requests. If you must batch, space calls 4–5s apart and check output size, or give up on search quickly and go direct to primary pages. Don't keep hammering.

  • Bing raw HTML often has no results at all (JS-rendered list) — grep b_algo finds nothing. Even in a headless browser, Bing's organic result anchors may not surface as <a href> elements. Prefer Startpage in-browser, or direct primary fetching.

  • A fetched page with ~5–30KB of text is often a shell/landing page, not the article. Compare len(text) against expectation; a 10KB page may only prove the title, not the claim. Cite only what you can actually read in-body.

  • Redirect pages. .io docs and some vendor URLs 302 to a JS redirect or a moved docs host. Follow with -L and re-read the final URL; a 757-byte "Redirecting to…" page is not content.

  • Search-snippet citations. A web_search description supports only what it literally says. Fetch the page and extract from the body when the claim needs more than the headline.

  • Verifying URLs into a report. Always do the %{http_code} check; search results frequently surface URLs that 404 on direct fetch, and citing them embeds an untrusted link.

  • Remember the successful path, not the dead ends. The durable lesson is "fetch primary pages directly + verify + extract"; search-engine friction is a side note, not a "search doesn't work" constraint.

Verification​

A citation-ready output means: every URL cited returned 2xx on a direct curl -L (or, for JS-gated/403'd pages, was actually read in-body via the scripts/fetch_playwright.py render or the r.jina.ai proxy), and every quoted figure/date was copied verbatim from the fetched page body (a [unverified] marker covers anything not actually read). Prefer the strongest verification available: verbatim copy from the primary page body beats a secondary synthesizer quoting it; flag partial captures (e.g. a reader proxy that only returns the title + opening lines) rather than presenting them as fully read.

Concrete session evidence (HTTP2/403 blocks, the Startpage-vs-other-engines outcome, scripting steps, 404-despite-search-hit cases) lives in references/session-evidence.md. Note: this skill covers retrieval and verification; the numbered citation ledger belongs to grounded-citations. A full worked example of researching competitor/enterprise AI offerings where several vendors 403 curl (URLs, CDX queries, entity-mistake catches) is in references/competitive-analyst-ai-pass.md.


Supporting files: this skill's supporting files are held in the docsite at docs/15-skills/_support/research/web-research-via-curl/ — fetch them fresh from jknash/docsite main alongside this page. Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/research/web-research-via-curl/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

version 1.0.0 · author Hermes Agent · license MIT.

Published by Muse · 2026-10-04.