Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.
.claude/skills/openlair-sentencepiece/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 108% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 139% | 0% |
Unsupervised tokenizer that works on raw text without language-specific preprocessing.
Use SentencePiece when:
Performance:
Use alternatives instead:
bash# Python pip install sentencepiece # C++ (requires CMake) git clone https://github.com/google/sentencepiece.git cd sentencepiece mkdir build && cd build cmake .. && make -j $(nproc) sudo make install
bash# Command-line (BPE with 8000 vocab) spm_train --input=data.txt --model_prefix=m --vocab_size=8000 --model_type=bpe # Python API import sentencepiece as spm spm.SentencePieceTrainer.train( input='data.txt', model_prefix='m', vocab_size=8000, model_type='bpe' )
Training time: ~1-2 minutes for 100MB corpus
pythonimport sentencepiece as spm # Load model sp = spm.SentencePieceProcessor(model_file='m.model') # Encode to pieces pieces = sp.encode('This is a test', out_type=str) print(pieces) # ['▁This', '▁is', '▁a', '▁test'] # Encode to IDs ids = sp.encode('This is a test', out_type=int) print(ids) # [284, 47, 11, 1243] # Decode text = sp.decode(ids) print(text) # "This is a test"
pythontext = "Hello world" pieces = sp.encode(text, out_type=str) print(pieces) # ['▁Hello', '▁world'] # Decode preserves spaces decoded = sp.decode_pieces(pieces) print(decoded) # "Hello world"
Key principle: Treat text as raw Unicode, whitespace = ▁ (meta symbol)
pythonspm.SentencePieceTrainer.train( input='data.txt', model_prefix='bpe_model', vocab_size=16000, model_type='bpe' )
Used by: mBART
pythonspm.SentencePieceTrainer.train( input='data.txt', model_prefix='unigram_model', vocab_size=8000, model_type='unigram' )
Used by: T5, ALBERT, XLNet
pythonspm.SentencePieceTrainer.train( input='corpus.txt', model_prefix='m', vocab_size=32000, model_type='unigram', character_coverage=0.9995, # 1.0 for CJK user_defined_symbols=['[SEP]', '[CLS]'], unk_piece='<unk>', num_threads=16 )
| Language Type | Coverage | Rationale | |---------------|----------|-----------| | English | 0.9995 | Most common chars | | CJK (Chinese) | 1.0 | All characters needed | | Multilingual | 0.9995 | Balance |
python# Sample different tokenizations for _ in range(3): pieces = sp.encode('tokenization', out_type=str, enable_sampling=True, alpha=0.1) print(pieces) # Output (different each time): # ['▁token', 'ization'] # ['▁tok', 'en', 'ization']
Use case: Data augmentation for robustness.
pythonspm.SentencePieceTrainer.train( input='c4_corpus.txt', model_prefix='t5', vocab_size=32000, model_type='unigram', user_defined_symbols=[f'<extra_id_{i}>' for i in range(100)], unk_id=2, eos_id=1, pad_id=0 )
pythonfrom transformers import T5Tokenizer # T5 uses SentencePiece internally tokenizer = T5Tokenizer.from_pretrained('t5-base') inputs = tokenizer('translate English to French: Hello', return_tensors='pt')
| Corpus | BPE (16k) | Unigram (8k) | |--------|-----------|--------------| | 100 MB | 1-2 min | 3-4 min | | 1 GB | 10-15 min | 30-40 min |
T5 family: t5-base, t5-large (32k vocab, Unigram) ALBERT: albert-base-v2 (30k vocab, Unigram) XLNet: xlnet-base-cased (32k vocab, Unigram) mBART: facebook/mbart-large-50 (250k vocab, BPE)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 5,839 | 3,326 | -43% | 1 | 1 | 0% | 1,252 | 2,307 | +84% | 0 | 0 | — |
case-02 | fail→pass | 6,212 | 2,943 | -53% | 1 | 1 | 0% | 1,270 | 2,215 | +74% | 0 | 0 | — |
case-03 | pass→pass | 5,015 | 4,233 | -16% | 1 | 1 | 0% | 1,039 | 2,486 | +139% | 0 | 0 | — |
case-04 | pass→pass | 15,758 | 8,733 | -45% | 1 | 1 | 0% | 3,093 | 3,316 | +7% | 0 | 0 | — |
case-05 | pass→pass | 10,009 | 9,118 | -9% | 1 | 1 | 0% | 2,077 | 3,545 | +71% | 0 | 0 | — |
case-06 | pass→pass | 7,779 | 3,991 | -49% | 1 | 1 | 0% | 1,693 | 2,555 | +51% | 0 | 0 | — |
case-07 | pass→pass | 6,848 | 2,830 | -59% | 1 | 1 | 0% | 1,336 | 2,117 | +58% | 0 | 0 | — |
case-08 | pass→pass | 3,074 | 2,015 | -34% | 1 | 1 | 0% | 661 | 1,978 | +199% | 0 | 0 | — |
case-09 | pass→pass | 3,632 | 3,277 | -10% | 1 | 1 | 0% | 657 | 2,234 | +240% | 0 | 0 | — |
case-10 | pass→pass | 5,638 | 3,129 | -45% | 1 | 1 | 0% | 1,105 | 2,169 | +96% | 0 | 0 | — |
case-11 | pass→pass | 3,038 | 2,472 | -19% | 1 | 1 | 0% | 647 | 2,082 | +222% | 0 | 0 | — |
case-12 | pass→pass | 4,034 | 3,254 | -19% | 1 | 1 | 0% | 794 | 2,271 | +186% | 0 | 0 | — |
case-13 | pass→pass | 3,302 | 2,465 | -25% | 1 | 1 | 0% | 579 | 2,089 | +261% | 0 | 0 | — |
case-14 | pass→pass | 3,288 | 2,573 | -22% | 1 | 1 | 0% | 622 | 2,055 | +230% | 0 | 0 | — |
case-15 | pass→pass | 2,630 | 1,795 | -32% | 1 | 1 | 0% | 395 | 1,975 | +400% | 0 | 0 | — |
case-16 | pass→pass | 5,184 | 3,014 | -42% | 1 | 1 | 0% | 1,058 | 2,247 | +112% | 0 | 0 | — |
case-17 | pass→pass | 4,897 | 2,562 | -48% | 1 | 1 | 0% | 917 | 2,054 | +124% | 0 | 0 | — |
case-18 | pass→pass | 5,539 | 3,919 | -29% | 1 | 1 | 0% | 1,059 | 2,395 | +126% | 0 | 0 | — |
case-19 | fail→pass | 5,168 | 2,897 | -44% | 1 | 1 | 0% | 1,059 | 2,208 | +108% | 0 | 0 | — |
case-20 | fail→pass | 6,878 | 4,753 | -31% | 1 | 1 | 0% | 1,301 | 2,541 | +95% | 0 | 0 | — |
case-21 | pass→pass | 3,490 | 2,657 | -24% | 1 | 1 | 0% | 566 | 2,096 | +270% | 0 | 0 | — |
case-22 | pass→pass | 5,794 | 3,673 | -37% | 1 | 1 | 0% | 1,184 | 2,478 | +109% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.