Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Document extraction pipeline architecture and patterns
.claude/skills/xberg-io-extraction-pipeline-patterns/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 7% | 0% |
Format detection → extractor routing → post-processing, across 106 formats / 140 file extensions
The full-registry counts are verified against published claims by scripts/sync_supported_counts.py verify. Runtime SUPPORTED_FORMAT_COUNT and SUPPORTED_EXTENSION_COUNT values are derived from the full static FORMATS registry.
crates/xberg/src/core/pipeline/ — orchestration (mod.rs, cache.rs, execution.rs,features.rs, format.rs, initialization.rs, page_markers.rs)
crates/xberg/src/core/mime.rs, core/formats.rs — detection and the FORMATS registrycrates/xberg/src/extractors/ — one module per format, each implementingInternalDocumentExtractor
crates/xberg/src/extraction/ — shared parsing/rendering helpers used by those extractorscrates/xberg/src/core/config/, core/config_validation/ — both directories, not filesEXT_TO_MIME, or from bytes viadetect_mime_type_from_bytes; validate against SUPPORTED_MIME_TYPES.
priority() extractor registered for that MIME.InternalDocument.core::pipeline::run_pipeline(doc, config) (async) orrun_pipeline_sync (WASM) runs validators, quality processing, chunking and hooks, and returns ExtractedDocument. Every extraction path goes through it.
extractors/{docx,pptx,ppt,doc,excel,odt,odp,hwp,hwpx,wordperfect}.rs and extractors/iwork/.
extractors/{markdown,text,rst,orgmode,asciidoc,typst}.rs and extractors/{rtf,djot_format}/.
extractors/{bibtex,jupyter,docbook,fictionbook}.rs and extractors/{latex,jats,epub}/.
extractors/pdf/.extractors/ and extraction/.extractors/ and extraction/html/.extractors/ and extraction/email.rs.extractors/archive.rs and extraction/archive/.extractors/.failure report is_encrypted in metadata rather than erroring out.
config.force_ocr and config.force_ocr_pages force it.
SecurityLimits.The cross-extractor fallback chain runs only for UnsupportedFormat and Plugin errors as defined by is_extractor_fallback_eligible. Parsing, IO, OCR, and validation errors abort the chain. A successful fallback records an extractor-fallback processing warning.
<cache_version_tag>-<content_hash>-<config_hash>, never path-based.The tag comes only from CARGO_PKG_VERSION and CACHE_SCHEMA_VERSION in cache/version.rs; it is not a build fingerprint.
can change without a crate version bump, bump CACHE_SCHEMA_VERSION. For A/B or revert checks, bump the schema or disable the cache so the experiment cannot replay the control.
extraction so a hit skips processing.
core/config/concurrency.rs::resolve_thread_budget.
core/io.rs::read_file_async currently reads the whole file with tokio::fs::read; thereis no AsyncRead extraction surface. Treat streaming as an open gap.
crates/xberg/src/plugins/. Registry selection is by priority(), highest wins — not by registration order. Register above 50 to override a built-in. See plugin-architecture-patterns.
All features live in crates/xberg/Cargo.toml; there is no FEATURE_MATRIX.md.
| Group | Features | | --- | --- | | OCR | ocr, ocr-wasm, paddle-ocr, paddle-ocr-tract, sceptre-ocr, candle-vlm-ocr | | Formats | pdf, office, excel, html, xml, email, archives, and format-specific flags | | AI/ML | embeddings, static-embeddings, layout, keywords, language detection, and NER flags | | Server | api (Axum), mcp, otel, prometheus, tokio-runtime | | Aggregates | formats, analysis, services, full, and platform target groups |
Bindings are separate crates, not features of crates/xberg. The one mutually-exclusive pair is ort-bundled / ort-dynamic. WASM excludes ORT-backed embeddings, but does carry ocr-wasm, keywords and static-embeddings. Full detail in feature-flag-policy.
run_pipeline / run_pipeline_sync.priority() for a MIMEtype.
SecurityLimits to user content — archive size, compression ratio, file count, nesting depth.Test the changed format categories and both success and failure paths. No coverage percentage is an enforced contract. The format headline test is the enforced count; update its constants and listed copy together with FORMATS. Use benchmark-workflow for performance or quality claims and test-corpus for bucket-backed fixtures.
FORMATS registry and how to add a format| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | pass→pass | 18,143 | 6,456 | -64% | 1 | 1 | 0% | 2,546 | 2,872 | +13% | 0 | 0 | — |
case-06 | fail→pass | 10,673 | 10,998 | +3% | 1 | 1 | 0% | 1,887 | 2,721 | +44% | 0 | 0 | — |
case-01 | fail→pass | 30,806 | 20,732 | -33% | 1 | 1 | 0% | 4,507 | 4,854 | +8% | 0 | 0 | — |
case-02 | fail→pass | 28,287 | 12,172 | -57% | 1 | 1 | 0% | 3,811 | 3,946 | +4% | 0 | 0 | — |
case-03 | fail→pass | 27,471 | 13,348 | -51% | 1 | 1 | 0% | 5,068 | 4,078 | -20% | 0 | 0 | — |
case-04 | pass→pass | 17,093 | 19,217 | +12% | 1 | 1 | 0% | 2,857 | 3,812 | +33% | 0 | 0 | — |
case-05 | fail→pass | 22,127 | 6,864 | -69% | 1 | 1 | 0% | 2,817 | 3,008 | +7% | 0 | 0 | — |
case-08 | fail→pass | 13,502 | 11,615 | -14% | 1 | 1 | 0% | 2,222 | 3,033 | +36% | 0 | 0 | — |
case-09 | fail→pass | 30,357 | 2,992 | -90% | 1 | 1 | 0% | 4,799 | 2,192 | -54% | 0 | 0 | — |
case-10 | fail→pass | 11,751 | 6,554 | -44% | 1 | 1 | 0% | 1,955 | 2,967 | +52% | 0 | 0 | — |
case-11 | pass→pass | 16,878 | 7,466 | -56% | 1 | 1 | 0% | 2,774 | 3,067 | +11% | 0 | 0 | — |
case-21 | pass→pass | 24,481 | 40,008 | +63% | 1 | 1 | 0% | 3,674 | 5,573 | +52% | 0 | 0 | — |
case-12 | pass→pass | 8,581 | 8,157 | -5% | 1 | 1 | 0% | 1,514 | 2,382 | +57% | 0 | 0 | — |
case-13 | fail→pass | 11,157 | 10,282 | -8% | 1 | 1 | 0% | 1,901 | 2,693 | +42% | 0 | 0 | — |
case-14 | pass→pass | 23,543 | 7,464 | -68% | 1 | 1 | 0% | 3,079 | 2,221 | -28% | 0 | 0 | — |
case-15 | fail→pass | 22,319 | 4,498 | -80% | 1 | 1 | 0% | 2,247 | 2,674 | +19% | 0 | 0 | — |
case-16 | fail→pass | 16,785 | 2,206 | -87% | 1 | 1 | 0% | 1,642 | 2,168 | +32% | 0 | 0 | — |
case-17 | fail→pass | 22,858 | 6,148 | -73% | 1 | 1 | 0% | 2,581 | 2,958 | +15% | 0 | 0 | — |
case-18 | pass→pass | 12,966 | 6,954 | -46% | 1 | 1 | 0% | 1,519 | 2,137 | +41% | 0 | 0 | — |
case-19 | pass→pass | 39,377 | 29,002 | -26% | 1 | 1 | 0% | 5,373 | 6,447 | +20% | 0 | 0 | — |
case-20 | pass→pass | 26,164 | 19,213 | -27% | 1 | 1 | 0% | 4,305 | 5,620 | +31% | 0 | 0 | — |
case-22 | pass→pass | 15,380 | 15,230 | -1% | 1 | 1 | 0% | 1,896 | 3,624 | +91% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/26/2026 | +59% |
| gemini-3.6-flash | verified | 8/12/2026 | +41% |
Other measured skills in the registry, with their headline benchmark lift.