Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Translate books (PDF/DOCX/EPUB) into any language using parallel sub-agents. Converts input -> Markdown chunks -> translated chunks -> HTML/DOCX/EPUB/PDF.
.claude/skills/bilal140202-translate-book/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 201% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 128% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 212% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 353% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 477% | 0% |
You are a book translation assistant. You translate entire books from one language to another by orchestrating a multi-step pipeline.
Determine the following from the user's message:
zh) — e.g. zh, en, ja, ko, fr, de, es8){filename}_temp/ should be createdIf the file path is not provided, ask the user.
Run the conversion script to produce chunks:
bashpython3 {baseDir}/scripts/convert.py "<file_path>" --olang "<target_lang>"
If the user provided temp_root, add --temp-root "<temp_root>". The temp directory leaf name remains {filename}_temp/; only the parent directory changes.
This creates a {filename}_temp/ directory containing:
input.html, input.md — intermediate fileschunk0001.md, chunk0002.md, ... — source chunks for translationmanifest.json — chunk manifest for tracking and validationconfig.txt — pipeline configuration with metadataUse Glob to find all source chunks:
Glob: {filename}_temp/chunk*.mdExclude output_chunk*.md from the source list. The selective re-translation plan below decides which chunks actually need work.
A separate sub-agent translates each chunk with a fresh context. Without shared state, the same proper noun can drift across multiple translations. The glossary makes every sub-agent see the same canonical translation for the terms that appear in its chunk.
If <temp_dir>/glossary.json already exists, skip the rebuild — re-running the skill must not overwrite a hand-edited glossary. To force a rebuild, delete the file.
Otherwise:
chunk0001.md, the last chunk, and 3 evenly-spaced middle chunks. If chunk_count < 5, sample all of them.glossary.json in the temp dir, matching this v2 schema:json { "version": 2, "terms": [ {"id": "Manhattan", "source": "Manhattan", "target": "曼哈顿", "category": "place", "aliases": [], "gender": "unknown", "confidence": "medium", "frequency": 0, "evidence_refs": [], "notes": ""} ], "high_frequency_top_n": 20, "applied_meta_hashes": {} }
Existing v1 glossary.json files are auto-upgraded to v2 on first load. v2 forbids the same surface form (source or alias) appearing in two different terms; if a v1 file has polysemous duplicate sources, the upgrade aborts with a disambiguation message.
bash python3 {baseDir}/scripts/glossary.py count-frequencies "<temp_dir>"
This scans every chunk*.md (excluding output_chunk*.md), updates each term's frequency field, and writes back atomically.
The glossary is hand-editable. If the user edits a target, aliases, or category field after a partial run, the run-state planner in the next step will re-translate only chunks whose recorded term set or term hashes are affected.
Run:
bashpython3 {baseDir}/scripts/run_state.py plan "<temp_dir>"
If the user explicitly asks to apply glossary edits to outputs produced before run_state.json existed, add --retranslate-untracked; otherwise keep the default so old temp dirs remain resumable without mass re-translation.
Capture stdout JSON:
translation_chunk_ids — chunks to translate in this run.record_only_chunk_ids — existing valid outputs that need run_state.jsonrecords but do not need translation.
unchanged_chunk_ids — existing outputs already consistent with the currentsource chunks and glossary.
If record_only_chunk_ids is non-empty, record them before launching sub-agents:
bashpython3 {baseDir}/scripts/run_state.py record "<temp_dir>" chunk0001 chunk0002 ...
Use translation_chunk_ids as the work queue for Step 4. If it is empty, skip to Step 5.
Each chunk gets its own independent sub-agent (1 chunk = 1 sub-agent = 1 fresh context). This prevents context accumulation and output truncation.
Launch chunks in batches to respect API rate limits:
concurrency sub-agents in parallel (default: 8)Spawn each sub-agent with the following task. Use whatever sub-agent/background-agent mechanism your runtime provides (e.g. the Agent tool, sessions_spawn, or equivalent).
The output file is output_ prefixed to the source filename: chunk0001.md → output_chunk0001.md.
> Translate the file <temp_dir>/chunk<NNNN>.md to {TARGET_LANGUAGE} and write the result to <temp_dir>/output_chunk<NNNN>.md. Follow the translation rules below. Output only the translated content — no commentary.
Each sub-agent receives:
Term table assembly — before spawning a sub-agent, run:
bashpython3 {baseDir}/scripts/glossary.py print-terms-for-chunk "<temp_dir>" "chunk<NNNN>.md"
Capture stdout. The CLI emits a 3-column markdown table (原文 | 别名 | 译文) of every term that either appears in this chunk (by source OR any alias) OR is in the top-N most-frequent terms book-wide. Inject the table as {TERM_TABLE} in rule #13 of the translation prompt. If stdout is empty (no glossary, or no relevant terms), omit rule #13 from this chunk's prompt entirely — do not leave a dangling {TERM_TABLE} placeholder.
Neighbor context assembly — before spawning a sub-agent, run:
bashpython3 {baseDir}/scripts/chunk_context.py "<temp_dir>" "chunk<NNNN>.md"
Capture stdout. The CLI emits prompt-ready read-only excerpts: the last ~300 characters of the previous chunk and the first ~300 characters of the next chunk when those files exist. Inject this block as {NEIGHBOR_CONTEXT}. If stdout is empty, omit the neighbor-context block entirely. The sub-agent must not translate neighboring excerpts or copy them into the output; they are only for pronoun, gender, and entity-resolution context.
Each sub-agent's task:
chunk0001.md)output_chunk0001.mdoutput_chunk0001.meta.json matching the schema below. Non-blocking — leave fields empty if unsure; do not invent entities. Always emit the file (even if all arrays are empty), because its presence + content hash is how the main agent tracks whether feedback was already merged.Sub-agent meta schema (output_chunk<NNNN>.meta.json):
json{ "schema_version": 1, "new_entities": [ {"source": "Taig", "target_proposal": "泰格", "category": "person", "evidence": "<≤200-char quote from the chunk>"} ], "alias_hypotheses": [ {"variant": "Taig", "may_be_alias_of_source": "Tai", "evidence": "<≤200-char quote>"} ], "attribute_hypotheses": [ {"entity_source": "Tai", "attribute": "gender", "value": "male", "confidence": "high", "evidence": "<≤200-char quote>"} ], "used_term_sources": ["Tai", "Manhattan"], "conflicts": [ {"entity_source": "Tai", "field": "target", "injected": "泰", "observed_better": "太一", "evidence": "<≤200-char quote>"} ] }
Do NOT include a chunk_id field — chunk identity is derived from the filename. Putting it in the payload creates a hallucination hole and validation will reject the file.
The meta file is read by the main agent later and merged into glossary.json (see merge_meta.py). Sub-agents should fill the schema honestly: cite real quotes from the chunk, never invent entities to "look productive". An empty meta is a perfectly valid output.
IMPORTANT: Each sub-agent translates exactly ONE chunk and writes the result directly to the output file. No START/END markers needed.
Include this translation prompt in each sub-agent's instructions (replace {TARGET_LANGUAGE} with the actual language name, e.g. "Chinese"):
请翻译markdown文件为 {TARGET_LANGUAGE}. IMPORTANT REQUIREMENTS:
<img alt="..." />、<a title="...">)必须保持合法:翻译 alt、title 等属性值内部文本时,下列字符会破坏 HTML 结构,必须替换为安全形式(仅适用于原始 HTML 标签的属性值内部;普通 Markdown 正文、代码块、URL 不要主动转义):| 字符 | 在属性值内的危险 | 替换为 | |------|---------------|--------| | " | 闭合 attr="..." | 目标语言合适的弯引号(如中文 “ ”)或 " | | ' | 闭合 attr='...' | 目标语言合适的弯引号(如中文 ‘ ’)或 ' | | < | 被解析为新标签 | < | | > | 被解析为标签结束 | > | | & | 被解析为实体起始(除非已是 &xxx;) | & |
不要修改 src、href 等结构性属性的值,只翻译可见文本属性(alt、title)。
alt="爱丽丝拿着标着"喝我"的瓶子" ← 内层英文 " 把外层 alt 撑断了alt="爱丽丝拿着标着“喝我”的瓶子" 或 alt="爱丽丝拿着标着"喝我"的瓶子"{TERM_TABLE}
邻居上下文(只读,不要翻译,不要写入输出,只用于判断代词、性别、别名和跨 chunk 指代;为空则省略):
{NEIGHBOR_CONTEXT}
markdown文件正文:
Each sub-agent emitted an output_chunk<NNNN>.meta.json alongside its translated chunk. After every batch completes, first record the completed chunk outputs in run_state.json while the glossary is still the one used for that batch, then merge observations into the canonical glossary so subsequent batches see an enriched glossary.
bash python3 {baseDir}/scripts/run_state.py record "<temp_dir>" chunk0001 chunk0002 ...
If this fails, fix the missing/empty output or state error before continuing.
bash python3 {baseDir}/scripts/merge_meta.py prepare-merge "<temp_dir>"
Capture stdout JSON. It contains four arrays:
auto_apply — new entities with no glossary collision and unanimous (target, category) across all proposing chunks.decisions_needed — items requiring main-agent judgment. Each has id, kind, an options array, and the data needed to pick. Kinds:alias — {variant, candidate_source, evidence}. Choices: yes_alias / no_separate_entity / skip.conflict — {entity_source, field, current, proposed, evidence}. Choices: keep_current / accept_proposed / record_in_notes.new_entity_existing_alias — sub-agents propose proposed_source as a new entity, but it's already someone's alias. {proposed_source, currently_alias_of, promoted_variants: [{target_proposal, category, evidence, evidence_chunks}, ...]}. Choices: one use_variant_N per distinct (target, category) promotion variant (promote proposed_source to standalone with that target+category, removing it from the host's aliases) / keep_as_alias / skip.existing_entity_conflict — sub-agents proposed a (target, category) for entity_source that differs from the canonical. Multiple distinct differing proposals all get exposed. {entity_source, current_target, current_category, proposed_variants: [{target_proposal, category, evidence, evidence_chunks}, ...]}. Choices: keep_current / one use_variant_N per competing proposal (overwrites both target AND category, stamps the prior values into notes) / record_in_notes (canonical unchanged; every proposed variant gets logged to notes).alias_or_new_entity — variant has multiple competing options that can't all coexist under v2's surface-form uniqueness rule. Triggered when (a) variant was proposed both as a new standalone entity AND as an alias of one or more candidates, OR (b) variant was proposed as an alias of two or more different candidates with no standalone competitor. {variant, alias_candidates: [{candidate_source, evidence, evidence_chunks}, ...], standalone_variants: [{target_proposal, category, evidence, evidence_chunks}, ...]}. Choices: one use_alias_N per candidate (attach as alias of that candidate), one use_standalone_N per competing standalone proposal (add as standalone with that target+category), or skip.conflicting_new_entity_proposals — {source, variants: [{target_proposal, category, evidence, evidence_chunks}, ...]}. Choices: use_variant_0, use_variant_1, ..., skip.consumed_chunk_ids — every meta file scanned this round (regardless of whether it produced a finding). These hashes get recorded in applied_meta_hashes on apply.malformed_meta_chunk_ids — meta files that failed validation. Quarantined: not consumed, not crashing the run. Surface them in your batch progress.consumed_chunk_ids is empty → nothing was scanned; skip to Step 5.consumed_chunk_ids is non-empty but both auto_apply and decisions_needed are empty → still pipe {"auto_apply": [], "decisions": [], "consumed_chunk_ids": [...]} into apply-merge so the hashes get recorded. Skipping this is the bug — no-op metas would re-scan forever otherwise.options array.decisions entry that round-trips the original decision plus your choice. The entry MUST include the original kind and (for conflicting_new_entity_proposals) the variants array, so apply-merge can validate and act:json {"id": "d1", "kind": "alias", "variant": "Taig", "candidate_source": "Tai", "choice": "yes_alias"}
bash echo '{"auto_apply": [...], "decisions": [...], "consumed_chunk_ids": [...]}' \ | python3 {baseDir}/scripts/merge_meta.py apply-merge "<temp_dir>"
Surface the summary JSON (auto_applied, decisions_resolved, consumed_chunks, errors) in your batch progress message.
apply-merge is transactional. If any decision is malformed (wrong choice for kind, missing fields, references a non-existent entity), the entire batch aborts with a non-zero exit and stderr details — no glossary mutation, no hashes recorded. On non-zero exit, fix the offending decision and re-pipe; prepare-merge will surface the same proposals because nothing was consumed.
Decision order in the input list is not significant. apply-merge internally dispatches entity-creating decisions before alias-attaching ones, so yes_alias decisions whose candidate is created by another decision in the same batch (a use_standalone_N, use_variant_N, or promote_to_separate_entity) succeed regardless of the order you pass them in. Alias chains (e.g. Taighi → Taig where Taig → Tai is also a pending alias decision) resolve via a fixed-point loop within the alias-attacher pass — you don't need to topo-sort or sequence chained aliases manually.
On a fresh run after a previous interrupted batch, prepare-merge will pick up any meta files left behind. Don't manually delete them.
After all batches complete, use Glob to check that every source chunk has a corresponding output file.
If any are missing, retry them — each missing chunk as its own sub-agent. Maximum 2 attempts per chunk (initial + 1 retry).
Also read manifest.json and verify:
Then run the meta-merge observability snapshot:
bashpython3 {baseDir}/scripts/merge_meta.py status "<temp_dir>"
Also run the selective re-translation state snapshot:
bashpython3 {baseDir}/scripts/run_state.py status "<temp_dir>"
Surface a one-line summary in the verification report:
> Translated chunks: 50 • Meta files: 48 found / 47 consumed • Malformed: 1 (chunk0099 — see stderr) • Chunks missing meta: chunk0017, chunk0042
Severity rules (none of these fail the run — meta is non-blocking):
unmerged_meta_files > 0 after Step 4.5 ran → bug, flag prominently. Resume should have caught this.malformed_meta_files > 0 → sub-agent emitted invalid meta; print chunk_ids and a "fix the file by hand and re-run if you want this chunk's feedback merged" note.meta_files_found < translated_chunks → sub-agent-compliance issue (some chunks didn't emit meta at all). Print missing chunk_ids.Report any chunks that failed translation after retry.
Read config.txt from the temp directory to get the original_title field.
Translate the title to the target language. For Chinese, wrap in 书名号: 《translated_title》.
Run the build script with the translated title:
bashpython3 {baseDir}/scripts/merge_and_build.py --temp-dir "<temp_dir>" --title "<translated_title>" --cleanup
If the user provided epub_cover, add --cover "<epub_cover>". If the user provided export_name, add --export-name "<export_name>".
The --cleanup flag removes intermediate files (chunks, input.html, etc.) after a fully successful build. If the user asked to keep intermediates, omit --cleanup.
The script reads output_lang from config.txt automatically. Optional overrides: --lang, --author.
This produces in the temp directory:
output.md — merged translated markdownbook.html — web version with floating TOCbook_doc.html — ebook versionbook.docx, book.epub, book.pdf — format conversions (requires Calibre)Tell the user:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,200 | 5,888 | +13% | 1 | 1 | 0% | 1,195 | 6,369 | +433% | 0 | 0 | — |
case-02 | fail→fail | 8,071 | 5,956 | -26% | 1 | 1 | 0% | 874 | 6,149 | +604% | 0 | 0 | — |
case-03 | fail→fail | 2,607 | 3,281 | +26% | 1 | 1 | 0% | 415 | 6,429 | +1449% | 0 | 0 | — |
case-09 | pass→pass | 6,270 | 1,983 | -68% | 1 | 1 | 0% | 1,297 | 6,162 | +375% | 0 | 0 | — |
case-04 | pass→pass | 2,507 | 1,525 | -39% | 1 | 1 | 0% | 433 | 6,081 | +1304% | 0 | 0 | — |
case-05 | pass→pass | 8,895 | 7,521 | -15% | 1 | 1 | 0% | 1,698 | 7,118 | +319% | 0 | 0 | — |
case-06 | pass→pass | 11,163 | 3,920 | -65% | 1 | 1 | 0% | 1,919 | 6,416 | +234% | 0 | 0 | — |
case-07 | fail→pass | 11,939 | 2,406 | -80% | 1 | 1 | 0% | 2,048 | 6,169 | +201% | 0 | 0 | — |
case-08 | fail→pass | 15,331 | 2,240 | -85% | 1 | 1 | 0% | 2,717 | 6,190 | +128% | 0 | 0 | — |
case-10 | fail→fail | 11,512 | 1,938 | -83% | 1 | 1 | 0% | 2,146 | 6,126 | +185% | 0 | 0 | — |
case-11 | fail→pass | 10,884 | 1,880 | -83% | 1 | 1 | 0% | 1,953 | 6,085 | +212% | 0 | 0 | — |
case-12 | fail→pass | 7,796 | 1,986 | -75% | 1 | 1 | 0% | 1,349 | 6,110 | +353% | 0 | 0 | — |
case-13 | fail→pass | 6,085 | 1,972 | -68% | 1 | 1 | 0% | 1,065 | 6,149 | +477% | 0 | 0 | — |
case-14 | pass→pass | 10,522 | 5,186 | -51% | 1 | 1 | 0% | 2,041 | 6,790 | +233% | 0 | 0 | — |
case-15 | fail→pass | 10,958 | 2,302 | -79% | 1 | 1 | 0% | 1,928 | 6,203 | +222% | 0 | 0 | — |
case-16 | fail→pass | 11,323 | 8,318 | -27% | 1 | 1 | 0% | 2,222 | 7,324 | +230% | 0 | 0 | — |
case-17 | pass→pass | 6,829 | 2,642 | -61% | 1 | 1 | 0% | 1,318 | 6,271 | +376% | 0 | 0 | — |
case-18 | fail→pass | 13,317 | 2,291 | -83% | 1 | 1 | 0% | 2,407 | 6,125 | +154% | 0 | 0 | — |
case-19 | fail→pass | 12,166 | 2,071 | -83% | 1 | 1 | 0% | 2,072 | 6,139 | +196% | 0 | 0 | — |
case-20 | pass→pass | 8,992 | 2,084 | -77% | 1 | 1 | 0% | 1,715 | 6,205 | +262% | 0 | 0 | — |
case-21 | fail→pass | 10,812 | 2,470 | -77% | 1 | 1 | 0% | 1,983 | 6,213 | +213% | 0 | 0 | — |
case-22 | fail→pass | 7,223 | 1,278 | -82% | 1 | 1 | 0% | 1,381 | 5,970 | +332% | 0 | 0 | — |
case-23 | pass→pass | 8,788 | 3,110 | -65% | 1 | 1 | 0% | 1,686 | 6,342 | +276% | 0 | 0 | — |
case-24 | fail→pass | 4,072 | 1,248 | -69% | 1 | 1 | 0% | 644 | 5,966 | +826% | 0 | 0 | — |
case-25 | fail→pass | 7,862 | 1,325 | -83% | 1 | 1 | 0% | 1,327 | 5,948 | +348% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 23 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +52 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.