Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Scrape sites with stealth browsing and Cloudflare bypass.
.claude/skills/nousresearch-scrapling/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 101% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 101% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 50% | 0% |
Scrapling is a web scraping framework with anti-bot bypass, stealth browser automation, and a spider framework. It provides three fetching strategies (HTTP, dynamic JS, stealth/Cloudflare) and a full CLI.
This skill is for educational and research purposes only. Users must comply with local/international data scraping laws and respect website Terms of Service.
web_extract tool does not return the data you needbashpip install "scrapling[all]" scrapling install
Minimal install (HTTP only, no browser):
bashpip install scrapling
With browser automation only:
bashpip install "scrapling[fetchers]" scrapling install
| Approach | Class | Use When | |----------|-------|----------| | HTTP | Fetcher / FetcherSession | Static pages, APIs, fast bulk requests | | Dynamic | DynamicFetcher / DynamicSession | JS-rendered content, SPAs | | Stealth | StealthyFetcher / StealthySession | Cloudflare, anti-bot protected sites | | Spider | Spider | Multi-page crawling with link following |
bashscrapling extract get 'https://example.com' output.md
With CSS selector and browser impersonation:
bashscrapling extract get 'https://example.com' output.md \ --css-selector '.content' \ --impersonate 'chrome'
bashscrapling extract fetch 'https://example.com' output.md \ --css-selector '.dynamic-content' \ --disable-resources \ --network-idle
bashscrapling extract stealthy-fetch 'https://protected-site.com' output.html \ --solve-cloudflare \ --block-webrtc \ --hide-canvas
bashscrapling extract post 'https://example.com/api' output.json \ --json '{"query": "search term"}'
The output format is determined by the file extension:
.html -- raw HTML.md -- converted to Markdown.txt -- plain text.json / .jsonl -- JSONpythonfrom scrapling.fetchers import Fetcher page = Fetcher.get('https://quotes.toscrape.com/') quotes = page.css('.quote .text::text').getall() for q in quotes: print(q)
pythonfrom scrapling.fetchers import FetcherSession with FetcherSession(impersonate='chrome') as session: page = session.get('https://example.com/', stealthy_headers=True) links = page.css('a::attr(href)').getall() for link in links[:5]: sub = session.get(link) print(sub.css('h1::text').get())
pythonpage = Fetcher.post('https://api.example.com/data', json={"key": "value"}) page = Fetcher.put('https://api.example.com/item/1', data={"name": "updated"}) page = Fetcher.delete('https://api.example.com/item/1')
pythonpage = Fetcher.get('https://example.com', proxy='http://user:pass@proxy:8080')
For pages that require JavaScript execution (SPAs, lazy-loaded content):
pythonfrom scrapling.fetchers import DynamicFetcher page = DynamicFetcher.fetch('https://example.com', headless=True) data = page.css('.js-loaded-content::text').getall()
pythonpage = DynamicFetcher.fetch( 'https://example.com', wait_selector=('.results', 'visible'), network_idle=True, )
Blocks fonts, images, media, stylesheets (~25% faster):
pythonfrom scrapling.fetchers import DynamicSession with DynamicSession(headless=True, disable_resources=True, network_idle=True) as session: page = session.fetch('https://example.com') items = page.css('.item::text').getall()
pythonfrom playwright.sync_api import Page from scrapling.fetchers import DynamicFetcher def scroll_and_click(page: Page): page.mouse.wheel(0, 3000) page.wait_for_timeout(1000) page.click('button.load-more') page.wait_for_selector('.extra-results') page = DynamicFetcher.fetch('https://example.com', page_action=scroll_and_click) results = page.css('.extra-results .item::text').getall()
For Cloudflare-protected or heavily fingerprinted sites:
pythonfrom scrapling.fetchers import StealthyFetcher page = StealthyFetcher.fetch( 'https://protected-site.com', headless=True, solve_cloudflare=True, block_webrtc=True, hide_canvas=True, ) content = page.css('.protected-content::text').getall()
pythonfrom scrapling.fetchers import StealthySession with StealthySession(headless=True, solve_cloudflare=True) as session: page1 = session.fetch('https://protected-site.com/page1') page2 = session.fetch('https://protected-site.com/page2')
All fetchers return a Selector object with these methods:
pythonpage.css('h1::text').get() # First h1 text page.css('a::attr(href)').getall() # All link hrefs page.css('.quote .text::text').getall() # Nested selection
pythonpage.xpath('//div[@class="content"]/text()').getall() page.xpath('//a/@href').getall()
pythonpage.find_all('div', class_='quote') # By tag + attribute page.find_by_text('Read more', tag='a') # By text content page.find_by_regex(r'\$\d+\.\d{2}') # By regex pattern
Find elements with similar structure (useful for product listings, etc.):
pythonfirst_product = page.css('.product')[0] all_similar = first_product.find_similar()
pythonel = page.css('.target')[0] el.parent # Parent element el.children # Child elements el.next_sibling # Next sibling el.prev_sibling # Previous sibling
For multi-page crawling with link following:
pythonfrom scrapling.spiders import Spider, Request, Response class QuotesSpider(Spider): name = "quotes" start_urls = ["https://quotes.toscrape.com/"] concurrent_requests = 10 download_delay = 1 async def parse(self, response: Response): for quote in response.css('.quote'): yield { "text": quote.css('.text::text').get(), "author": quote.css('.author::text').get(), "tags": quote.css('.tag::text').getall(), } next_page = response.css('.next a::attr(href)').get() if next_page: yield response.follow(next_page) result = QuotesSpider().start() print(f"Scraped {len(result.items)} quotes") result.items.to_json("quotes.json")
Route requests to different fetcher types:
pythonfrom scrapling.fetchers import FetcherSession, AsyncStealthySession class SmartSpider(Spider): name = "smart" start_urls = ["https://example.com/"] def configure_sessions(self, manager): manager.add("fast", FetcherSession(impersonate="chrome")) manager.add("stealth", AsyncStealthySession(headless=True), lazy=True) async def parse(self, response: Response): for link in response.css('a::attr(href)').getall(): if "protected" in link: yield Request(link, sid="stealth") else: yield Request(link, sid="fast", callback=self.parse)
pythonspider = QuotesSpider(crawldir="./crawl_checkpoint") spider.start() # Ctrl+C to pause, re-run to resume from checkpoint
scrapling install after pip install -- without it, DynamicFetcher and StealthyFetcher will failsolve_cloudflare=True adds 5-15 seconds to fetch time -- only enable when needed| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | pass→pass | 4,199 | 3,593 | -14% | 1 | 1 | 0% | 1,053 | 3,321 | +215% | 0 | 0 | — |
case-21 | pass→pass | 4,432 | 3,689 | -17% | 1 | 1 | 0% | 1,076 | 3,273 | +204% | 0 | 0 | — |
case-22 | pass→pass | 5,082 | 2,955 | -42% | 1 | 1 | 0% | 1,235 | 3,098 | +151% | 0 | 0 | — |
case-01 | fail→pass | 13,052 | 8,535 | -35% | 1 | 1 | 0% | 1,546 | 3,107 | +101% | 0 | 0 | — |
case-02 | fail→pass | 8,162 | 5,400 | -34% | 1 | 1 | 0% | 1,820 | 3,663 | +101% | 0 | 0 | — |
case-03 | fail→pass | 11,091 | 6,646 | -40% | 1 | 1 | 0% | 2,840 | 3,913 | +38% | 0 | 0 | — |
case-04 | fail→pass | 20,856 | 1,804 | -91% | 1 | 1 | 0% | 4,108 | 2,804 | -32% | 0 | 0 | — |
case-05 | fail→pass | 8,633 | 1,735 | -80% | 1 | 1 | 0% | 1,838 | 2,751 | +50% | 0 | 0 | — |
case-06 | fail→pass | 9,224 | 2,911 | -68% | 1 | 1 | 0% | 1,841 | 3,032 | +65% | 0 | 0 | — |
case-07 | fail→pass | 4,463 | 1,784 | -60% | 1 | 1 | 0% | 1,068 | 2,728 | +155% | 0 | 0 | — |
case-08 | fail→pass | 14,933 | 2,119 | -86% | 1 | 1 | 0% | 3,718 | 2,846 | -23% | 0 | 0 | — |
case-09 | fail→pass | 13,122 | 5,822 | -56% | 1 | 1 | 0% | 3,198 | 3,818 | +19% | 0 | 0 | — |
case-10 | fail→pass | 11,686 | 5,234 | -55% | 1 | 1 | 0% | 2,967 | 3,292 | +11% | 0 | 0 | — |
case-11 | fail→pass | 8,578 | 2,367 | -72% | 1 | 1 | 0% | 1,643 | 2,939 | +79% | 0 | 0 | — |
case-12 | fail→pass | 9,258 | 3,435 | -63% | 1 | 1 | 0% | 1,686 | 3,110 | +84% | 0 | 0 | — |
case-13 | fail→pass | 7,108 | 5,040 | -29% | 1 | 1 | 0% | 1,712 | 3,620 | +111% | 0 | 0 | — |
case-14 | fail→pass | 10,092 | 6,783 | -33% | 1 | 1 | 0% | 1,160 | 3,042 | +162% | 0 | 0 | — |
case-15 | fail→pass | 12,644 | 9,209 | -27% | 1 | 1 | 0% | 1,887 | 3,544 | +88% | 0 | 0 | — |
case-16 | fail→pass | 5,477 | 3,902 | -29% | 1 | 1 | 0% | 1,260 | 3,343 | +165% | 0 | 0 | — |
case-17 | fail→pass | 10,672 | 7,397 | -31% | 1 | 1 | 0% | 2,289 | 4,122 | +80% | 0 | 0 | — |
case-18 | fail→pass | 5,409 | 1,816 | -66% | 1 | 1 | 0% | 1,153 | 2,781 | +141% | 0 | 0 | — |
case-19 | fail→pass | 14,471 | 6,652 | -54% | 1 | 1 | 0% | 3,478 | 4,103 | +18% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +86 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.