Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 101% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 159% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -11% | 0% |
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:0b5cd4bec9d2a8f3452e07139ef56d6b87f9f7990714d15483a03435e751af59 Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->
Use this when feeding documents into an LLM context window or a vector store. Xberg chunks two ways: inline during extraction (chunks land on each document's chunks field), or standalone via the chunk command for text you already have. Sizing is character-based by default, or token-based when a tokenizer model is supplied.
Turn on chunking with --chunk and the chunks appear on the structured result under chunks:
bash# 1000-char chunks, 200-char overlap (defaults when --chunk is on) xberg extract report.pdf --chunk --format json | jq '.chunks | length' # Explicit size + overlap xberg extract report.pdf --chunk --chunk-size 1500 --chunk-overlap 300 --format json
Overlap must be smaller than chunk size — the CLI rejects --chunk-overlap >= --chunk-size. When you set only --chunk-overlap against an existing config, an overlap that exceeds the size is clamped to chunk_size / 4.
chunk commandChunk text you already have, from --text or stdin. Output defaults to JSON:
bash# From a flag xberg chunk --text "long document text ..." --chunk-size 800 --chunk-overlap 100 # From stdin (pipe extracted content straight in) xberg extract notes.md | xberg chunk --chunk-size 500 --format json
JSON output carries chunks (array of strings), chunk_count, the resolved config (max_characters, overlap, chunker_type), and input_size_bytes. Use --format text for a human-readable dump with --- chunk N --- separators.
> Note: in the JSON output, chunker_type is rendered capitalized ("Text", > "Markdown", "Yaml", "Semantic") because it is emitted via Rust's Debug > formatting, whereas the --chunker-type input flag is lowercase > (text, markdown, yaml, semantic). Lowercase the value before > comparing if you parse it back.
--chunker-type selects the splitting strategy (standalone chunk command):
| Type | Behavior | | ---------- | ------------------------------------------------------------------- | | text | Default. Plain character-window splitting with overlap. | | markdown | Markdown-aware — splits on structure (headings, blocks) where possible. | | yaml | YAML-aware splitting for structured config/data documents. | | semantic | Topic-boundary splitting driven by --topic-threshold (0.0–1.0, default 0.75). |
bash# Markdown-aware chunking keeps headings and blocks intact xberg chunk --text "$(cat README.md)" --chunker-type markdown # Semantic chunking — lower threshold = more, smaller topic chunks xberg chunk --text "$(cat transcript.txt)" --chunker-type semantic --topic-threshold 0.6
By default --chunk-size counts characters. To size chunks by tokens for a specific model, pass --chunking-tokenizer with a HuggingFace tokenizer id. On the extract command this implicitly enables chunking. Requires the chunking-tokenizers feature (present in the default CLI build).
bash# Size chunks by GPT-4o tokens during extraction xberg extract report.pdf --chunking-tokenizer Xenova/gpt-4o --format json # Or on the standalone command xberg chunk --text "$(cat doc.txt)" --chunking-tokenizer Xenova/gpt-4o --chunk-size 512
With a tokenizer set, --chunk-size is interpreted in tokens, not characters.
Field names in config files are snake_case under [chunking]:
toml[chunking] max_characters = 1000 overlap = 200 chunker_type = "markdown"
bashxberg extract report.pdf --config xberg.toml --format json
> CLI flags map to config fields as --chunk-size → max_characters and > --chunk-overlap → overlap. In config files use the snake_case names.
From Python, enable chunking on the config and read the chunks off the document in the result envelope (result.results[0].chunks):
pythonfrom xberg import ExtractInput, extract, ExtractionConfig, ChunkingConfig config = ExtractionConfig( chunking=ChunkingConfig(max_characters=1000, overlap=200), ) result = await extract(ExtractInput(uri="report.pdf"), config) for chunk in result.results[0].chunks or []: print(len(chunk.content))
> The public Python ChunkingConfig (a dataclass) uses constructor kwargs > max_characters / overlap; the Rust core struct fields are also > max_characters / overlap. TOML/JSON config keys are max_chars / > max_overlap (with max_characters / overlap accepted as serde aliases), > and dict-form config passed to ExtractionConfig likewise accepts the > max_chars / max_overlap aliases; Node's ChunkingConfig interface uses > maxCharacters / overlap. See references/python-api.md and > references/rust-api.md in the sibling xberg skill.
10–20% overlap. Use markdown chunking for docs to keep sections whole.
overlap; size by tokens to stay under the model window.
semantic chunker; tune --topic-thresholddown for finer splits, up for coarser ones.
extract; clamped to size / 4 whenonly overlap is changed against an existing config.
--chunking-tokenizer errors if theCLI was built without chunking-tokenizers. The default build includes it.
chunk command bails on empty text;provide --text or pipe non-empty stdin.
See references/configuration.md for the full [chunking] schema and references/cli-reference.md for every chunk flag.
Other measured skills in the registry, with their headline benchmark lift.