Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.
.claude/skills/xberg-io-extracting-with-ocr/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 111% | 0% |
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:3ec8b7cf60f56cbe5cc15a5a0b29d0c2f8e3cf4c3823503eb1128fb7f9ee11db Source-Hash: blake3:ef1fa958e3b61fa61d2a2275a9c04dfab5c94186e3351a2216afb25cbb36ff1a Schema-Version: v1 -->
Use this when a document is image-based: scanned PDFs, photographed pages, screenshots, JPEG/PNG/TIFF with text. Xberg auto-OCRs raster images and auto-detects PDFs that lack a text layer. Force it on when extraction returned empty/garbled text from a PDF that "looks" textual.
content field, but the file opens visually.bashxberg extract scan.pdf --force-ocr=true xberg extract scan.pdf --ocr=true --ocr-language eng
If a page has an unreliable text layer, --force-ocr=true re-rasterizes and runs OCR on every page.
Tesseract is the default and ships with the CLI — no extra install. Other backends are opt-in:
| Backend | Flag | Install | Notes | | ------------- | ------------------------------------- | ------------------------------------------------ | -------------------------------------------------------------- | | Tesseract | --ocr-backend tesseract (default) | bundled | Best general-purpose, 100+ languages via tessdata. | | PaddleOCR | --ocr-backend paddle-ocr | bundled (ONNX Runtime) | Strong on Asian scripts. Not available on WASM or Windows. | | Candle VLM | --ocr-backend candle-trocr (and other candle-*) | bundled (Candle) | Local vision OCR models (candle-trocr, candle-paddleocr-vl, candle-glm-ocr, candle-deepseek-ocr). | | VLM (hosted) | --ocr-backend vlm + --vlm-model | liter-llm provider (--vlm-api-key) | Multimodal LLM via liter-llm. Use when OCR fails on dense or handwritten layouts. |
Pick Tesseract first. Switch only when accuracy is unacceptable.
Tesseract uses ISO 639-2 codes. Default is eng. Combine with +:
bashxberg extract menu.jpg --ocr=true --ocr-language "eng+deu" xberg extract bilingual.pdf --ocr-language "eng+jpn" xberg extract any.pdf --ocr-language all # all installed packs
Install missing packs at the OS level:
bash# macOS brew install tesseract-lang # Debian/Ubuntu sudo apt install tesseract-ocr-deu tesseract-ocr-jpn tesseract-ocr-fra # Specific lang only sudo apt install tesseract-ocr-<iso639-2>
Xberg fails fast with a helpful error if you request a language pack that is not installed. Read the error — it names the missing file.
--ocr=true — enable OCR (auto-enabled for images and scanned PDFs).--force-ocr=true — OCR every page even if a text layer exists.--disable-ocr=true — never OCR (extract embedded text only or fail).--ocr-language <lang> — single code or +-joined list, or all.--ocr-backend <tesseract|paddle-ocr|vlm|candle-trocr|candle-paddleocr-vl|candle-glm-ocr|candle-deepseek-ocr> — pick backend.--ocr-auto-rotate=true — pre-rotate via the auto-rotate model.--acceleration <cpu|coreml|cuda|tensorrt|auto> — ONNX accelerator forpaddle-ocr / auto-rotate / layout models.
instant. Do not pass --no-cache=true unless you have a reason.
xberg batch *.pdf --ocr=true — internal workerpool parallelizes across CPU cores. Cap with --max-concurrent N if memory is tight.
--target-dpi (default 300) only for low-resolution scans. HigherDPI is slower; 200 is usually enough for printed text.
--ocr-auto-rotate=true only when pages may be rotated; theclassifier adds latency.
--acceleration coreml typically beats CPU forpaddle-ocr and layout detection.
Long flag chains belong in xberg.toml — auto-discovered from cwd upward.
tomlforce_ocr = true output_format = "markdown" [ocr] backend = "tesseract" language = "eng+deu" auto_rotate = true
Then just run:
bashxberg extract document.pdf
--force-ocr — the file has abogus zero-width text layer. Re-run with --force-ocr=true.
--ocr-auto-rotate=true or pre-rotate.passed via --ocr-language; consider paddle-ocr for Chinese/Japanese.
See references/cli-reference.md and references/configuration.md in the sibling xberg skill for the full flag and config schema.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,702 | 3,377 | -80% | 1 | 1 | 0% | 2,713 | 1,942 | -28% | 0 | 0 | — |
case-02 | fail→pass | 13,313 | 4,045 | -70% | 1 | 1 | 0% | 2,136 | 2,014 | -6% | 0 | 0 | — |
case-03 | fail→pass | 14,172 | 5,830 | -59% | 1 | 1 | 0% | 1,922 | 2,299 | +20% | 0 | 0 | — |
case-04 | pass→pass | 6,929 | 6,272 | -9% | 1 | 1 | 0% | 1,129 | 2,443 | +116% | 0 | 0 | — |
case-05 | pass→pass | 10,186 | 7,248 | -29% | 1 | 1 | 0% | 1,615 | 2,484 | +54% | 0 | 0 | — |
case-06 | pass→pass | 4,247 | 2,777 | -35% | 1 | 1 | 0% | 543 | 1,833 | +238% | 0 | 0 | — |
case-07 | fail→pass | 27,227 | 4,712 | -83% | 1 | 1 | 0% | 4,357 | 2,194 | -50% | 0 | 0 | — |
case-08 | fail→pass | 16,756 | 3,064 | -82% | 1 | 1 | 0% | 870 | 1,835 | +111% | 0 | 0 | — |
case-09 | fail→pass | 8,729 | 3,936 | -55% | 1 | 1 | 0% | 1,325 | 2,005 | +51% | 0 | 0 | — |
case-10 | fail→pass | 16,981 | 3,768 | -78% | 1 | 1 | 0% | 2,357 | 1,923 | -18% | 0 | 0 | — |
case-11 | fail→pass | 8,472 | 4,634 | -45% | 1 | 1 | 0% | 1,264 | 2,087 | +65% | 0 | 0 | — |
case-12 | fail→pass | 11,673 | 4,000 | -66% | 1 | 1 | 0% | 1,645 | 1,848 | +12% | 0 | 0 | — |
case-13 | fail→pass | 12,983 | 3,178 | -76% | 1 | 1 | 0% | 2,129 | 1,874 | -12% | 0 | 0 | — |
case-14 | fail→pass | 16,974 | 2,723 | -84% | 1 | 1 | 0% | 2,348 | 1,779 | -24% | 0 | 0 | — |
case-15 | pass→pass | 8,740 | 5,038 | -42% | 1 | 1 | 0% | 1,361 | 1,998 | +47% | 0 | 0 | — |
case-16 | fail→pass | 24,639 | 4,332 | -82% | 1 | 1 | 0% | 2,388 | 2,103 | -12% | 0 | 0 | — |
case-17 | pass→pass | 4,805 | 5,410 | +13% | 1 | 1 | 0% | 574 | 1,824 | +218% | 0 | 0 | — |
case-18 | fail→pass | 18,835 | 3,479 | -82% | 1 | 1 | 0% | 2,723 | 1,881 | -31% | 0 | 0 | — |
case-19 | fail→pass | 8,834 | 3,289 | -63% | 1 | 1 | 0% | 1,287 | 1,947 | +51% | 0 | 0 | — |
case-20 | fail→pass | 10,103 | 3,141 | -69% | 1 | 1 | 0% | 1,310 | 1,741 | +33% | 0 | 0 | — |
case-21 | fail→pass | 14,937 | 3,076 | -79% | 1 | 1 | 0% | 2,442 | 1,736 | -29% | 0 | 0 | — |
case-22 | pass→pass | 8,816 | 3,520 | -60% | 1 | 1 | 0% | 958 | 1,926 | +101% | 0 | 0 | — |
case-23 | pass→pass | 9,694 | 2,490 | -74% | 1 | 1 | 0% | 1,728 | 1,683 | -3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +70 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/19/2026 | +55% |
| gemini-3.6-flash | verified | 9/9/2026 | +64% |
| gemini-3.6-flash | verified | 9/1/2026 | +60% |
| gemini-3.6-flash | verified | 8/17/2026 | +59% |
| gemini-3.6-flash | verified | 8/11/2026 | +36% |
Other measured skills in the registry, with their headline benchmark lift.