Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Chunking, embeddings, and RAG pipeline integration
.claude/skills/xberg-io-chunking-embeddings/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 4% | 0% |
Text splitting, ONNX/static embedding generation, RAG pipeline integration
Locations: crates/xberg/src/chunking/ and crates/xberg/src/embeddings/ (both directories, not single files).
ExtractionConfig.chunking: Option<ChunkingConfig> drives it. The standalone entry points are chunking::chunk_text(text, &ChunkingConfig, page_boundaries) -> Result<ChunkingResult> (chunking/core.rs) and chunking::rag::chunk_for_rag(text, &ChunkingConfig) (chunking/rag.rs), which upgrades ChunkerType::Text to Markdown and fills each chunk's heading_path.
ChunkingResult { chunks: Vec<Chunk>, chunk_count: usize }. Chunk carries content, chunk_type, metadata, and the optional vectors embedding, sparse_embedding, late_interaction (types/extraction.rs).
ChunkerType — there is no strategy enum beyond thisText (default), Markdown, Yaml, Semantic (core/config/processing.rs). Semantic splits at embedding-based topic shifts when an EmbeddingConfig is present, and falls back to a structural-boundary heuristic otherwise — topic_threshold has no effect on the fallback path.
ChunkingConfig fields and their serde wire names| Field | Wire name (config file) | Default | | --- | --- | --- | | max_characters | max_chars (alias max_characters) | 1000 | | overlap | max_overlap (alias overlap) | 200 | | trim | trim | true | | chunker_type | chunker_type | Text | | preset | preset | none |
The renames are load-bearing: a config file that writes max_characters works only via the alias, and a typo'd key is silently ignored (see config-loading-precedence).
ChunkingConfig.preset resolves through resolve_preset(), which is #[cfg(feature = "embeddings")]-gated — without that feature it is a no-op and the preset name does nothing. A preset overrides max_characters and overlap and, if no embedding config was given, selects the model.
| Preset | chunk_size | overlap | dims | backend | | --- | --- | --- | --- | --- | | fast | 512 | 50 | 384 | ONNX | | balanced | 1024 | 100 | 768 | ONNX | | quality | 2000 | 200 | 1024 | ONNX | | multilingual | 1024 | 100 | 768 | ONNX | | gte-modernbert-base | 1024 | 100 | 768 | ONNX | | lightweight | 512 | 50 | 256 | static (model2vec) | | arctic-embed-m-v2.0 | 1024 | 100 | 768 | ONNX | | qwen3-embedding-0.6b | 2000 | 200 | 1024 | ONNX |
Source of truth: EMBEDDING_PRESETS in crates/xberg/src/embeddings/mod.rs.
There is no TextEmbeddingManager, no embed_chunks(), no ChunkWithEmbedding, no RagDocument, and no fastembed dependency — do not write code against any of those.
Model selection is EmbeddingModelType, a tagged enum (core/config/processing.rs): Preset { name } (recommended), Custom { … } (HuggingFace ONNX), Llm { … }, Plugin { … }.
Two defaults disagree and both are live: EmbeddingModelType::default() is the gte-modernbert-base preset (what bindings and #[serde(default)] get), while EmbeddingConfig::default() names balanced via default_balanced_embedding_model(). Read the constructor you are actually going through before assuming which model runs.
EmbeddingConfig defaults: normalize = true, batch_size = 32, max_embed_duration_secs = Some(60), max_sequence_length = None (falls back to 512, capped at the model's own model_max_length).
tomlembeddings = ["onnx-runtime", "dep:ndarray", "chunking", "tokio-runtime", "embedding-presets"]
ort-bundled (the default ORT linkage) downloads ONNX Runtime at build time — no system install, no ORT_DYLIB_PATH. That variable matters only under ort-dynamic.
static-embeddings is the pure-Rust model2vec path and the only dense embedder available on no-ort-target (WASM/Android). embedding-presets carries preset metadata alone and is WASM-safe.
embeddings feature is inert — resolve_preset() is compiled out.max_chars/max_overlap, not the Rust field names.EmbeddingConfig.normalize defaults to true; leave it on.ChunkingConfig is resolved and why typos are silentembeddings vs static-embeddings vs embedding-presets| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,467 | 9,127 | -37% | 1 | 1 | 0% | 2,320 | 3,071 | +32% | 0 | 0 | — |
case-02 | fail→pass | 15,554 | 10,749 | -31% | 1 | 1 | 0% | 2,481 | 3,303 | +33% | 0 | 0 | — |
case-03 | fail→pass | 20,054 | 10,682 | -47% | 1 | 1 | 0% | 3,152 | 3,346 | +6% | 0 | 0 | — |
case-04 | fail→pass | 10,829 | 3,188 | -71% | 1 | 1 | 0% | 1,660 | 1,962 | +18% | 0 | 0 | — |
case-05 | pass→pass | 16,245 | 3,370 | -79% | 1 | 1 | 0% | 2,701 | 2,030 | -25% | 0 | 0 | — |
case-06 | fail→pass | 10,410 | 3,878 | -63% | 1 | 1 | 0% | 1,567 | 2,103 | +34% | 0 | 0 | — |
case-07 | fail→pass | 12,008 | 3,833 | -68% | 1 | 1 | 0% | 1,910 | 1,983 | +4% | 0 | 0 | — |
case-08 | fail→pass | 12,070 | 4,951 | -59% | 1 | 1 | 0% | 1,821 | 2,279 | +25% | 0 | 0 | — |
case-09 | fail→pass | 12,024 | 5,183 | -57% | 1 | 1 | 0% | 1,964 | 2,223 | +13% | 0 | 0 | — |
case-10 | fail→pass | 12,510 | 1,661 | -87% | 1 | 1 | 0% | 2,009 | 1,710 | -15% | 0 | 0 | — |
case-11 | fail→pass | 11,740 | 2,544 | -78% | 1 | 1 | 0% | 1,744 | 1,817 | +4% | 0 | 0 | — |
case-12 | fail→pass | 22,360 | 2,384 | -89% | 1 | 1 | 0% | 2,580 | 1,837 | -29% | 0 | 0 | — |
case-13 | fail→pass | 11,471 | 1,876 | -84% | 1 | 1 | 0% | 1,863 | 1,693 | -9% | 0 | 0 | — |
case-14 | fail→pass | 19,423 | 4,604 | -76% | 1 | 1 | 0% | 3,276 | 2,290 | -30% | 0 | 0 | — |
case-15 | fail→pass | 13,449 | 4,115 | -69% | 1 | 1 | 0% | 1,984 | 2,069 | +4% | 0 | 0 | — |
case-16 | fail→pass | 9,708 | 1,997 | -79% | 1 | 1 | 0% | 1,538 | 1,746 | +14% | 0 | 0 | — |
case-17 | pass→pass | 11,286 | 1,414 | -87% | 1 | 1 | 0% | 1,893 | 1,597 | -16% | 0 | 0 | — |
case-18 | fail→pass | 16,178 | 15,380 | -5% | 1 | 1 | 0% | 2,402 | 2,297 | -4% | 0 | 0 | — |
case-19 | fail→fail | 18,672 | 10,788 | -42% | 1 | 1 | 0% | 3,395 | 3,394 | -0% | 0 | 0 | — |
case-20 | pass→pass | 11,928 | 5,550 | -53% | 1 | 1 | 0% | 1,880 | 2,279 | +21% | 0 | 0 | — |
case-21 | pass→pass | 11,873 | 8,175 | -31% | 1 | 1 | 0% | 2,155 | 2,940 | +36% | 0 | 0 | — |
case-22 | fail→pass | 11,671 | 2,584 | -78% | 1 | 1 | 0% | 1,827 | 1,855 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +73 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/11/2026 | +50% |
Other measured skills in the registry, with their headline benchmark lift.