Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Fetch any URL and convert to markdown using Chrome CDP. Saves the rendered HTML snapshot alongside the markdown, uses an upgraded Defuddle pipeline with better web-component handling and YouTube transcript extraction, and automatically falls back to the pre-Defuddle HTML-to-Markdown pipeline when needed. If local browser capture fails entirely, it can fall back to the hosted defuddle.md API. Supports two modes - auto-capture on page load, or wait for user signal (for pages requiring login). Use
.claude/skills/leoyeai-baoyu-url-to-markdown/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 159% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 394% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 99% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 206% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 191% | 0% |
Fetches any URL via Chrome CDP, saves the rendered HTML snapshot, and converts it to clean markdown.
Important: All scripts are located in the scripts/ subdirectory of this skill.
Agent Execution Instructions:
{baseDir}{baseDir}/scripts/<script-name>.ts${BUN_X} runtime: if bun installed → bun; if npx available → npx -y bun; else suggest installing bun{baseDir} and ${BUN_X} in this document with actual valuesScript Reference: | Script | Purpose | |--------|---------| | scripts/main.ts | CLI entry point for URL fetching | | scripts/html-to-markdown.ts | Markdown conversion entry point and converter selection | | scripts/parsers/index.ts | Unified parser entry: dispatches URL-specific rules before generic converters | | scripts/parsers/types.ts | Unified parser interface shared by all rule files | | scripts/parsers/rules/*.ts | One file per URL rule, for example X status and X article | | scripts/defuddle-converter.ts | Defuddle-based conversion | | scripts/legacy-converter.ts | Pre-Defuddle legacy extraction and markdown conversion | | scripts/markdown-conversion-shared.ts | Shared metadata parsing and markdown document helpers |
Check EXTEND.md existence (priority order):
bash# macOS, Linux, WSL, Git Bash test -f .baoyu-skills/baoyu-url-to-markdown/EXTEND.md && echo "project" test -f "${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md" && echo "xdg" test -f "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md" && echo "user"
powershell# PowerShell (Windows) if (Test-Path .baoyu-skills/baoyu-url-to-markdown/EXTEND.md) { "project" } $xdg = if ($env:XDG_CONFIG_HOME) { $env:XDG_CONFIG_HOME } else { "$HOME/.config" } if (Test-Path "$xdg/baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "xdg" } if (Test-Path "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "user" }
┌────────────────────────────────────────────────────────┬───────────────────┐ │ Path │ Location │ ├────────────────────────────────────────────────────────┼───────────────────┤ │ .baoyu-skills/baoyu-url-to-markdown/EXTEND.md │ Project directory │ ├────────────────────────────────────────────────────────┼───────────────────┤ │ $HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md │ User home │ └────────────────────────────────────────────────────────┴───────────────────┘
┌───────────┬───────────────────────────────────────────────────────────────────────────┐ │ Result │ Action │ ├───────────┼───────────────────────────────────────────────────────────────────────────┤ │ Found │ Read, parse, apply settings │ ├───────────┼───────────────────────────────────────────────────────────────────────────┤ │ Not found │ MUST run first-time setup (see below) — do NOT silently create defaults │ └───────────┴───────────────────────────────────────────────────────────────────────────┘
EXTEND.md Supports: Download media by default | Default output directory | Default capture mode | Timeout settings
CRITICAL: When EXTEND.md is not found, you MUST use AskUserQuestion to ask the user for their preferences before creating EXTEND.md. NEVER create EXTEND.md with defaults without asking. This is a BLOCKING operation — do NOT proceed with any conversion until setup is complete.
Use AskUserQuestion with ALL questions in ONE call:
Question 1 — header: "Media", question: "How to handle images and videos in pages?"
Question 2 — header: "Output", question: "Default output directory?"
Question 3 — header: "Save", question: "Where to save preferences?"
After user answers, create EXTEND.md at the chosen location, confirm "Preferences saved to path]", then continue.
Full reference: references/config/first-time-setup.md
| Key | Default | Values | Description | |-----|---------|--------|-------------| | download_media | ask | ask / 1 / 0 | ask = prompt each time, 1 = always download, 0 = never | | default_output_dir | empty | path or empty | Default output directory (empty = ./url-to-markdown/) |
EXTEND.md → CLI mapping: | EXTEND.md key | CLI argument | Notes | |---------------|-------------|-------| | download_media: 1 | --download-media | | | default_output_dir: ./posts/ | --output-dir ./posts/ | Directory path. Do NOT pass to -o (which expects a file path) |
Value priority:
--download-media, -o, --output-dir)-captured.html filex.com / twitter.comarchive.ph / related archive mirrors can restore the original URL from input[name=q] and prefer #CONTENT before falling back to the page bodydefuddle.md/<url> and still save markdownbash# Auto mode (default) - capture when page loads ${BUN_X} {baseDir}/scripts/main.ts <url> # Force headless only ${BUN_X} {baseDir}/scripts/main.ts <url> --browser headless # Force visible browser ${BUN_X} {baseDir}/scripts/main.ts <url> --browser headed # Wait mode - wait for user signal before capture ${BUN_X} {baseDir}/scripts/main.ts <url> --wait # Save to specific file ${BUN_X} {baseDir}/scripts/main.ts <url> -o output.md # Save to a custom output directory (auto-generates filename) ${BUN_X} {baseDir}/scripts/main.ts <url> --output-dir ./posts/ # Download images and videos to local directories ${BUN_X} {baseDir}/scripts/main.ts <url> --download-media
| Option | Description | |--------|-------------| | <url> | URL to fetch | | -o <path> | Output file path — must be a file path, not directory (default: auto-generated) | | --output-dir <dir> | Base output directory — auto-generates {dir}/{domain}/{slug}.md (default: ./url-to-markdown/) | | --wait | Wait for user signal before capturing | | --browser <mode> | Browser strategy: auto (default), headless, or headed | | --headless | Shortcut for --browser headless | | --headed | Shortcut for --browser headed | | --timeout <ms> | Page load timeout (default: 30000) | | --download-media | Download image/video assets to local imgs/ and videos/, and rewrite markdown links to local relative paths |
| Mode | Behavior | Use When | |------|----------|----------| | Auto (default) | Try headless first, then retry in visible Chrome if needed | Public pages, static content, unknown pages | | Wait (--wait) | User signals when ready | Login-required, lazy loading, paywalls |
Wait mode workflow:
--wait → script outputs "Press Enter when ready"Default browser fallback:
--browser headed --waitCRITICAL: The agent must treat headless capture as provisional. Some sites render differently in headless mode and can silently return an error shell, partially hydrated page, or low-quality extraction without causing the CLI to fail.
After every run that used --browser auto or --browser headless, the agent MUST inspect the saved markdown first, and inspect the saved -captured.html when the markdown looks suspicious.
Application errorThis page could not be foundauto unless there is already a clear reason to use wait mode--browser headed for ordinary rendering issues--browser headed --wait when the page may need login, anti-bot interaction, cookie acceptance, or extra hydration time--wait is used, tell the user exactly what to do:defuddle.md after the local browser strategies have failed or are clearly lower fidelityEach run saves two files side by side:
url, title, description, author, published, optional coverImage, and captured_at, followed by converted markdown content*-captured.html, containing the rendered page HTML captured from ChromeWhen Defuddle or page metadata provides a language hint, the markdown front matter also includes language.
The HTML snapshot is saved before any markdown media localization, so it stays a faithful capture of the page DOM used for conversion. If the hosted defuddle.md API fallback is used, markdown is still saved, but there is no local -captured.html snapshot for that run.
Default: url-to-markdown/<domain>/<slug>.md With --output-dir ./posts/: ./posts/<domain>/<slug>.md
HTML snapshot path uses the same basename:
url-to-markdown/<domain>/<slug>-captured.html./posts/<domain>/<slug>-captured.html<slug>: From page title or URL path (kebab-case, 2-6 words)<slug>-YYYYMMDD-HHMMSS.mdWhen --download-media is enabled:
imgs/ next to the markdown filevideos/ next to the markdown fileConversion order:
--browser headed --wait and ask the user to complete access before capturehttps://defuddle.md/<url> API and save its markdown output directlyCLI output will show:
Converter: parser:... when a URL-specific parser succeededConverter: defuddle when Defuddle succeedsConverter: legacy:... plus Fallback used: ... when fallback was neededConverter: defuddle-api when local browser capture failed and the hosted API was used insteadBased on download_media setting in EXTEND.md:
| Setting | Behavior | |---------|----------| | 1 (always) | Run script with --download-media flag | | 0 (never) | Run script without --download-media flag | | ask (default) | Follow the ask-each-time flow below |
--download-media → markdown savedhttps:// in image/video links)AskUserQuestion:--download-media (overwrites markdown with localized links)| Variable | Description | |----------|-------------| | URL_CHROME_PATH | Custom Chrome executable path | | URL_DATA_DIR | Custom data directory | | URL_CHROME_PROFILE_DIR | Custom Chrome profile directory |
Troubleshooting: Chrome not found → set URL_CHROME_PATH. Timeout → increase --timeout. Complex pages → try --wait mode. If markdown quality is poor, inspect the saved -captured.html and check whether the run logged a legacy fallback.
--wait and capture after the watch page is fully hydrated.https://defuddle.md/<url>. In shell form: curl https://defuddle.md/stephango.comCustom configurations via EXTEND.md. See Preferences section for paths and supported options.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | pass→fail | 6,290 | 6,610 | +5% | 1 | 1 | 0% | 1,014 | 4,326 | +327% | 0 | 0 | — |
case-01 | fail→fail | 8,534 | 7,572 | -11% | 1 | 1 | 0% | 1,955 | 4,509 | +131% | 0 | 0 | — |
case-02 | fail→fail | 14,409 | 6,835 | -53% | 1 | 1 | 0% | 2,910 | 4,289 | +47% | 0 | 0 | — |
case-03 | fail→fail | 14,160 | 8,093 | -43% | 1 | 1 | 0% | 2,144 | 4,429 | +107% | 0 | 0 | — |
case-04 | fail→fail | 11,117 | 9,056 | -19% | 1 | 1 | 0% | 2,092 | 4,647 | +122% | 0 | 0 | — |
case-06 | pass→fail | 8,034 | 9,819 | +22% | 1 | 1 | 0% | 1,368 | 4,689 | +243% | 0 | 0 | — |
case-07 | fail→pass | 9,766 | 5,507 | -44% | 1 | 1 | 0% | 1,881 | 4,870 | +159% | 0 | 0 | — |
case-08 | fail→pass | 4,934 | 2,922 | -41% | 1 | 1 | 0% | 931 | 4,598 | +394% | 0 | 0 | — |
case-09 | fail→pass | 12,960 | 2,584 | -80% | 1 | 1 | 0% | 2,250 | 4,475 | +99% | 0 | 0 | — |
case-10 | fail→pass | 7,811 | 1,975 | -75% | 1 | 1 | 0% | 1,383 | 4,237 | +206% | 0 | 0 | — |
case-11 | fail→pass | 9,136 | 1,988 | -78% | 1 | 1 | 0% | 1,471 | 4,280 | +191% | 0 | 0 | — |
case-12 | fail→pass | 11,167 | 6,290 | -44% | 1 | 1 | 0% | 1,486 | 4,982 | +235% | 0 | 0 | — |
case-13 | fail→pass | 4,249 | 4,717 | +11% | 1 | 1 | 0% | 775 | 4,964 | +541% | 0 | 0 | — |
case-14 | pass→pass | 13,540 | 3,614 | -73% | 1 | 1 | 0% | 2,092 | 4,428 | +112% | 0 | 0 | — |
case-15 | fail→pass | 12,391 | 2,597 | -79% | 1 | 1 | 0% | 1,948 | 4,404 | +126% | 0 | 0 | — |
case-16 | pass→pass | 11,587 | 3,138 | -73% | 1 | 1 | 0% | 1,802 | 4,624 | +157% | 0 | 0 | — |
case-17 | fail→pass | 17,204 | 4,147 | -76% | 1 | 1 | 0% | 2,586 | 4,549 | +76% | 0 | 0 | — |
case-18 | fail→fail | 15,601 | 6,572 | -58% | 1 | 1 | 0% | 2,644 | 4,378 | +66% | 0 | 0 | — |
case-19 | pass→pass | 11,824 | 8,095 | -32% | 1 | 1 | 0% | 1,860 | 5,261 | +183% | 0 | 0 | — |
case-20 | fail→fail | 6,683 | 10,235 | +53% | 1 | 1 | 0% | 1,262 | 4,517 | +258% | 0 | 0 | — |
case-21 | fail→pass | 7,052 | 6,696 | -5% | 1 | 1 | 0% | 1,035 | 5,131 | +396% | 0 | 0 | — |
case-22 | fail→fail | 17,844 | 15,927 | -11% | 1 | 1 | 0% | 3,550 | 6,522 | +84% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 14 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 14 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.