Web Research via Curl
Fetch/verify web facts via curl when no web_search tool.
A working pattern for gathering verifiable, primary-source web facts using
terminal + curl when the session has no web_search / web_extract
tool configured (this happens in subagent and some headless environments).
It also improves any citation workflow by favoring direct primary sources and
checking that every cited URL actually resolves before it is trusted.
Companion skills: pair with grounded-citations for the citation/ledger
management once you have fetched real sources. This skill is about the
retrieval and verification mechanism; grounded-citations owns the numbering
and evidence chain.
When to Use
- The task asks for current (2024–20XX) verifiable facts and your toolset has
no
web_search/web_extract/browsertools — onlyterminal. - You must confirm a source URL actually exists (HTTP 200) before citing it.
- You want primary/vendor/.gov docs instead of a search-snippet summary.
When the main session has NO web tools at all: delegate parallel research
If your own toolset has no web_search/web_extract/browser AND you'd rather
not hand-curl every page, the fastest path to a broad, well-sourced report is to
fan out parallel research subagents via delegate_task. Each subagent gets
its own toolset (it can curl / render / search itself), and you orchestrate the
synthesis. This pattern produced a 37-source, 10-offering report in ~11 minutes
of wall-clock research.
- Issue one
delegate_task(goal=..., context=...)call per research stream, in a single message so they run concurrently. Do NOT pass a single-elementtasksarray — batch mode rejects it with "Batch mode requires at least 2 tasks." For N independent streams, make N separategoal-based calls in one turn (up to thedelegation.max_concurrent_childrenlimit). - Give each subagent a self-contained brief: the exact questions to answer,
the output persona/format, a required per-claim source format
(
Claim — [Title](URL), publisher, date), and an explicit "do not fabricate URLs or numbers; mark unverifiable items [unverified]" instruction. Children know nothing of your conversation — pass everything they need. - Split by concern, not by source. A proven split for a market/offerings report: (1) market landscape & statistics, (2) competitor offerings, (3) emerging solution patterns. Each returns a structured markdown report with its own source list.
- Have subagents write their report to a file (
/root/.../report.md) so you canread_fileit after they finish, rather than relying on the summary alone. - Register every returned URL in the citation ledger (
grounded-citationssources.py add) before writing prose, then synthesize with inline[n]citations and render the Sources block mechanically. Seereferences/research-report-pipeline.mdfor the full end-to-end worked example (citation-integration pitfalls + PDF export).
Core principle: fetch primary sources directly, don't rely on search
Search engines (DuckDuckGo HTML, Bing) are unreliable for scripting: they rate-limit rapid consecutive requests and serve anti-bot "challenge" pages. When a search does return real results, the snippet only supports what it literally says — go fetch the underlying page and extract from its text.
For a report the fastest, most trustworthy reads come from the authoritative
page itself: vendor product pages, official docs (.docs./docs. subdomains),
GitHub repos, and .gov / official-EC pages. Go there directly when you know
the canonical URL; use search only to discover URLs you don't yet know.
Procedure
① Probe connectivity once so a 0-byte / network-blocked curl isn't misread as "page is empty":
curl -s -m 15 -o /dev/null -w "%{http_code}" https://www.google.com
② Fetch each page to a file with a browser UA string (many sites 403 or serve JS-shells to curl's default UA):
ua="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0 Safari/537.36"
curl -s -m 25 -A "$ua" -L "https://example.com/page" -o research/page.html
wc -c research/page.html # sanity-check size
③ Extract readable text with scripts/html_to_text.py (handles the
<script>/<style> noise, tag stripping, and HTML-entity unescaping):
python3 ~/.hermes/skills/research/web-research-via-curl/scripts/html_to_text.py research/page.html
④ Verify each URL you plan to cite returns a real page before including it:
code=$(curl -s -o /dev/null -w "%{http_code}" -m 20 -A "$ua" -L "https://example.com/page")
echo "$code <- https://example.com/page"
Treat anything other than 2xx/the odd 3xx as unverified. Two URLs in this
session's source came back 404 on direct fetch even though they appeared in
live search results — search-hit ≠ valid link. Always check before citing.
⑤ Extract the exact supporting sentence from the fetched text and copy it
verbatim into your notes so a grounded-citations quote step can match it
later, or so your report quotes the page accurately.
⑥ Label honestly. A claim you could not retrieve stays [unverified] in
the report. JS-gated (no body in the HTML) and 403'd pages count as unverified
for anything you could not actually read. Say which sources were JS-gated/403
rather than silently citing them.
Escalation ladder when curl is blocked (403 / JS-shell / HTTP2 errors)
Many authoritative .gov/.com pages 403 curl's UA outright (Gartner → 403,
McKinsey → net::ERR_HTTP2_PROTOCOL_ERROR), and in-browser/headless-rendered
searches are needed to discover URLs. Take this order — only escalate when
the current step fails:
- Direct primary fetch via curl (Procedure above). Many .gov/vendor pages
work fine this way (Federal Reserve, JPMorganChase Institute responded to
curl). Sanity-check
wc -c(a ~5–30KB body is often a JS shell, not the article). - Headless browser render — for JS-gated pages and to run searches in a
real browser engine. Install Playwright + a Chromium engine, and do not
skip the OS-library step (it resolves missing
libnspr4/NSS shared libs and yieldsTargetClosedErroron launch otherwise):Then render + extract withpip install playwrightpython3 -m playwright install chromiumpython3 -m playwright install-deps chromium # "error while loading shared libraries: libnspr4.so..." → run thisscripts/fetch_playwright.py(launch args--no-sandbox --disable-blink-features=AutomationControlled, set a full Chrome UA,wait_until="domcontentloaded", wait a few seconds, theninner_text("body")). - Search via Startpage in the browser — when Bing/DDG/Mojeek/Brave all
serve anti-bot challenge shells to curl and to raw HTML parsing, Startpage
(
https://www.startpage.com/sp/search?query=...) proxies Google results and renders them as plain anchor tags in a headless browser. It was the only search engine that returned real result links out of DDG/Bing/Mojeek/Brave tested. Usescripts/search_startpage.py "<query>"to run it and print result links; then fetch the primary page directly for the facts. Extract with:aelements wherehrefstarts withhttpand filters out the engine's own domains. - Reader proxy for hard-blocked sites — prefix any URL with
https://r.jina.ai/<full-url>to get a Markdown/plain-text version of pages that 403/bounce even the browser (used successfully on McKinsey). It returnsTitle / URL Source / Published Time / Markdown Content— ideal verbatim source capture. Note it may still truncate very long pages.
When primary pages AND search are both blocked (vendor 403 / anti-bot)
Enterprise/vendor sites (Infosys, TCS, McKinsey, Wipro press, Accenture newsroom) frequently 403 or 404 automated curl clients behind Akamai/CDN bot walls — even with a browser UA. Don't stop; there are three reliable fallback paths that worked together:
- Bing RSS as a discovery search. When Bing regular HTML and DDG (both
html.duckduckgo.comandlite.duckduckgo.com) return empty anti-bot shells, use Bing's RSS endpoint — it returned real, dated results while the HTML search returned nothing:Parsecurl -s -m 30 -A "$ua" \"https://www.bing.com/search?format=rss&q=<urlencoded>&count=8"<item><title>/<link>/<pubDate>/<description>from the feed. Caveat: it often surfaces only navigational hits (homepages, Wikipedia, LinkedIn) for wordy queries — use it to discover a canonical URL, then fetch that page directly.grep b_algostill finds nothing on RSS; parse the<item>s. - Wayback CDX API to recover 403'd vendor pages. The "closest snapshot"
convenience API (
https://archive.org/wayback/available?url=...) returnsNoneeven when archives exist, and live vendor URLs 403. The authoritative lookup is the CDX API with prefix match + status filter:→ returnscurl -s -m 40 \"https://web.archive.org/cdx/search/cdx?url=example.com/path&matchType=prefix&output=json&limit=3&filter=statuscode:200&collapse=timestamp:6"[["urlkey","timestamp","original",...]]rows; reconstruct the snapshot ashttps://web.archive.org/web/<timestamp>/<original>. Add&collapse=digestto dedupe near-identical captures, or&collapse=timestamp:6for one-per-day. If CDX returns[], it genuinely has no 200 capture — don't loop retrying. - Wikipedia as a fallback verifier for blocked vendor claims. When the
vendor page 403s, a current Wikipedia article (e.g. on the parent company or
product) often independently confirms the service/offering exists. Cite
Wikipedia explicitly and downgrade vendor-specific numbers (investment
$, headcount, launch dates) that only appeared in un-fetchable press to
[unverified]. - Verify entity identity, not just existence. For competitors/boutiques in a rumor mill, fetch the named company's page and confirm it actually does what was claimed. In one competitive pass "Andium" turned out to be an industrial-IoT/emissions hardware vendor, and "Latent Space" an AI-engineer newsletter — neither was the AI consultancy they were rumored to be. Calling that out beats copying a mis-categorized competitor.
Pitfalls
-
Vendor 403 ≠ claim dead. The page is blocked, not necessarily gone. Try the Wayback CDX path and Wikipedia before marking the claim unverified.
-
wayback/availableunder-reports. Prefer the CDX API; the convenience endpoint returnedNonein this session for URLs CDX had captures of. -
Search engines may choke on multi-word queries that contain a common substring. E.g. "mid-market genAI adoption" → Bing/DDG keyed on the slang word "mid" and returned dictionary/irrelevant results. If a query yields off-topic hits, reword (drop the ambiguous token, or add an unambiguous qualifier) rather than assuming "no results."
-
Startpage via curl won't parse.
startpage.comserved a 403 to curl with a default UA and a bare 363-byte body; it only works through a real browser engine (use the Playwright search script, not curl). -
Rapid consecutive searches → anti-bot page. DDG/Bing return ~14KB of no-results (a challenge/consent shell) after a few fast requests. If you must batch, space calls 4–5s apart and check output size, or give up on search quickly and go direct to primary pages. Don't keep hammering.
-
Bing raw HTML often has no results at all (JS-rendered list) —
grep b_algofinds nothing. Even in a headless browser, Bing's organic result anchors may not surface as<a href>elements. Prefer Startpage in-browser, or direct primary fetching. -
A fetched page with ~5–30KB of text is often a shell/landing page, not the article. Compare
len(text)against expectation; a 10KB page may only prove the title, not the claim. Cite only what you can actually read in-body. -
Redirect pages.
.iodocs and some vendor URLs 302 to a JS redirect or a moved docs host. Follow with-Land re-read the final URL; a 757-byte "Redirecting to…" page is not content. -
Search-snippet citations. A
web_searchdescription supports only what it literally says. Fetch the page and extract from the body when the claim needs more than the headline. -
Verifying URLs into a report. Always do the
%{http_code}check; search results frequently surface URLs that 404 on direct fetch, and citing them embeds an untrusted link. -
Remember the successful path, not the dead ends. The durable lesson is "fetch primary pages directly + verify + extract"; search-engine friction is a side note, not a "search doesn't work" constraint.
Verification
A citation-ready output means: every URL cited returned 2xx on a direct
curl -L (or, for JS-gated/403'd pages, was actually read in-body via the
scripts/fetch_playwright.py render or the r.jina.ai proxy), and every
quoted figure/date was copied verbatim from the fetched page body (a
[unverified] marker covers anything not actually read). Prefer the strongest
verification available: verbatim copy from the primary page body beats a
secondary synthesizer quoting it; flag partial captures (e.g. a reader proxy
that only returns the title + opening lines) rather than presenting them as
fully read.
Concrete session evidence (HTTP2/403 blocks, the Startpage-vs-other-engines
outcome, scripting steps, 404-despite-search-hit cases) lives in
references/session-evidence.md. Note: this skill covers retrieval and
verification; the numbered citation ledger belongs to grounded-citations.
A full worked example of researching competitor/enterprise AI offerings where
several vendors 403 curl (URLs, CDX queries, entity-mistake catches) is in
references/competitive-analyst-ai-pass.md.
Supporting files: this skill's supporting files are held in the docsite at
docs/15-skills/_support/research/web-research-via-curl/— fetch them fresh fromjknash/docsitemain alongside this page. Source:jknash/hermes-shared-skills· branchhermes-jkdev001@1d0d545c3970·skills/research/web-research-via-curl/· view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.
version 1.0.0 · author Hermes Agent · license MIT.
Published by Muse · 2026-10-04.