Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Crawl websites to extract text content from multiple pages and internal links. Use when the user wants to scrape a website, gather content across a section, archive docs text, understand a site more thoroughly by following internal links, or follow internal links beyond a single page. Prefer this general skill for bounded crawl workflows, route to Crawl4AI or Katana when the task is specifically AI-ready ingestion or site mapping, and escalate blocked pages to camofox before any residential-prox
.claude/skills/valtterimelkko-web-crawling/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-08 | ✓→✓ | = Same ✓ | -18% | 0% |
| case-10 | ✓→✓ | = Same ✓ | -17% | 0% |
Use this skill for general multi-page crawling when the user wants more than a one-page fetch but does not yet need a highly specialised workflow.
crawl4ai when:katana when:scrapegraph-ai when:web-crawling skill directly when:When the crawl path hits blocking or rendering problems:
web_fetchcamofox if the problem is:residential-proxy only if all of these are true:camofox was actually tried and still failed, or the blocked target is a narrow public HTTP endpoint where a browser step is not sufficientDo not automatically turn a whole crawl into a proxy-backed crawl.
crawl4ai.katana.Save crawled content in a file with clear source markers, for example:
markdown# Website Crawl - example.com **Seed URL:** https://example.com **Depth:** 2 **Crawl Date:** YYYY-MM-DD --- ## Page: https://example.com/page1 **Depth:** 0 [content] --- ## Page: https://example.com/page2 **Depth:** 1 [content]
markdown## Web crawl result - Seed URL: ... - Depth / page cap: ... - Primary method: native fetch | crawl4ai | katana | custom script - Escalated to camofox: yes/no - Escalated to residential proxy: yes/no - Output file: ... - Notes: ...
This public extraction keeps optional Browserless reference material and example scripts for cases where a local or paid browser-automation escalation path is useful. Those extras are optional and require your own BROWSERLESSIO_API_KEY if you choose to use them. They are not required for the core skill.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | fail→fail | 11,675 | 28,515 | +144% | 1 | 1 | 0% | 1,130 | 2,515 | +123% | 0 | 0 | — |
case-01 | fail→fail | 12,660 | 6,393 | -50% | 1 | 1 | 0% | 2,453 | 1,139 | -54% | 0 | 0 | — |
case-02 | fail→fail | 12,663 | 8,291 | -35% | 1 | 1 | 0% | 2,474 | 1,348 | -46% | 0 | 0 | — |
case-03 | fail→fail | 16,615 | 10,228 | -38% | 1 | 1 | 0% | 501 | 1,495 | +198% | 0 | 0 | — |
case-05 | fail→pass | 12,144 | 4,444 | -63% | 1 | 1 | 0% | 2,043 | 1,561 | -24% | 0 | 0 | — |
case-06 | fail→pass | 13,139 | 9,073 | -31% | 1 | 1 | 0% | 1,523 | 1,728 | +13% | 0 | 0 | — |
case-07 | fail→pass | 13,713 | 5,517 | -60% | 1 | 1 | 0% | 2,136 | 1,751 | -18% | 0 | 0 | — |
case-08 | pass→pass | 12,489 | 5,836 | -53% | 1 | 1 | 0% | 2,266 | 1,860 | -18% | 0 | 0 | — |
case-09 | fail→fail | 5,772 | 3,871 | -33% | 1 | 1 | 0% | 948 | 1,399 | +48% | 0 | 0 | — |
case-10 | pass→pass | 17,980 | 6,639 | -63% | 1 | 1 | 0% | 2,258 | 1,877 | -17% | 0 | 0 | — |
case-11 | pass→pass | 14,070 | 4,406 | -69% | 1 | 1 | 0% | 2,208 | 1,648 | -25% | 0 | 0 | — |
case-12 | pass→pass | 10,343 | 3,940 | -62% | 1 | 1 | 0% | 1,880 | 1,531 | -19% | 0 | 0 | — |
case-13 | fail→fail | 8,626 | 2,953 | -66% | 1 | 1 | 0% | 1,666 | 1,349 | -19% | 0 | 0 | — |
case-14 | pass→pass | 10,361 | 4,305 | -58% | 1 | 1 | 0% | 1,851 | 1,525 | -18% | 0 | 0 | — |
case-15 | pass→pass | 11,587 | 6,978 | -40% | 1 | 1 | 0% | 1,911 | 1,986 | +4% | 0 | 0 | — |
case-16 | pass→pass | 12,147 | 3,802 | -69% | 1 | 1 | 0% | 2,000 | 1,437 | -28% | 0 | 0 | — |
case-17 | pass→pass | 9,364 | 2,810 | -70% | 1 | 1 | 0% | 1,626 | 1,320 | -19% | 0 | 0 | — |
case-18 | pass→pass | 14,923 | 7,105 | -52% | 1 | 1 | 0% | 2,212 | 1,902 | -14% | 0 | 0 | — |
case-19 | pass→pass | 8,873 | 3,795 | -57% | 1 | 1 | 0% | 1,449 | 1,415 | -2% | 0 | 0 | — |
case-20 | pass→pass | 9,097 | 4,221 | -54% | 1 | 1 | 0% | 1,274 | 1,509 | +18% | 0 | 0 | — |
case-21 | pass→pass | 7,814 | 10,096 | +29% | 1 | 1 | 0% | 1,684 | 2,135 | +27% | 0 | 0 | — |
case-22 | pass→pass | 6,551 | 7,163 | +9% | 1 | 1 | 0% | 1,306 | 2,248 | +72% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.