Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。**遇到以下任一情况就主动使用本 skill**:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 / 幻灯片分析 / 发票提取 / 合同分析 / 文档统计 / 错别字 / 语病 / 字号检查 / 简历分析 / 多文档对比;③任务涉及从文档中提取表格、数值、图表、格式(颜色/高亮/字号)、组织架构、时间线等结构化信息。仅不用于:Excel/CSV 数据分析(使用 sn-da-excel-workflow)、纯图片分析(使用 sn-da-image-caption)。
.claude/skills/opensensenova-sn-da-non-spreadsheet-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-10 | ✓→✗ | ▼ Worse | 236% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 105% | 0% |
End-to-end workflow for Word, PDF, and PPT document parsing. Each format has specific parsing pitfalls — follow the format-specific sub-skill exactly.
pythonimport os input_path = "/mnt/data/..." # from user # Detect single file vs directory (multi-file scenario) if os.path.isdir(input_path): all_files = [ os.path.join(input_path, f) for f in os.listdir(input_path) if f.lower().endswith(('.docx', '.doc', '.pdf', '.pptx', '.ppt')) ] print(f"Found {len(all_files)} documents: {all_files}") else: all_files = [input_path] # Route by extension ext = os.path.splitext(all_files[0])[-1].lower() print(f"File type: {ext}")
> Critical rule: When input_path is a directory OR the user says "这些文件" / "所有文档", > process every file and aggregate. Never stop at the first file.
| Extension | Sub-skill to load | |-----------|------------------| | .docx / .doc | capability/word-analysis/SKILL.md | | .pdf | capability/pdf-analysis/SKILL.md | | .pptx / .ppt | capability/ppt-analysis/SKILL.md |
read_file(path="<skills_root>/sn-da-non-spreadsheet-analysis/capability/<format>-analysis/SKILL.md")Load only the sub-skill you need — do not load all three at once.
Follow the sub-skill's extraction pattern. For all formats:
caption.pyAfter extracting data, verify before answering:
python# For count/statistics questions: spot-check 3-5 items sample = result_list[:3] print(f"Sample check: {sample}") print(f"Total count: {len(result_list)}") # For numeric calculations: print intermediate values print(f"Max={max_val}, Min={min_val}, Range={max_val - min_val}") # For unit-sensitive answers: always include the unit print(f"Answer: {value} {unit}") # e.g., "475 千港元" not just "475"
for page in doc, for slide in prs.slides, for para in doc.paragraphscaption.py for OCRcaption.pypytesseract or easyocr as primary OCR — they are not installed; use caption.pyWhen a page, slide, or embedded image needs vision understanding, load the sn-da-image-caption skill first, then use its scripts/caption.py:
read_file(path="<skills_root>/sn-da-image-caption/SKILL.md")pythonimport subprocess, json CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py" def caption_image(image_path, prompt=None): cmd = ["python3", CAPTION, image_path, "--json"] if prompt: cmd += ["--prompt", prompt] result = subprocess.run(cmd, capture_output=True, text=True, timeout=60) if result.returncode != 0: raise RuntimeError(f"caption failed: {result.stderr[:200]}") return json.loads(result.stdout)["description"] # Example prompts by content type: # Table: "提取表格所有内容,Markdown 表格格式,保持行列结构,数值不四舍五入。" # Chart: "提取图表标题、坐标轴标签、每个数据点的数值。Markdown 表格输出。" # Diagram: "描述所有节点和连接关系。"
sn-da-non-spreadsheet-analysis/capability/word-analysis/SKILL.md — .docx/.doc
sn-da-non-spreadsheet-analysis/capability/pdf-analysis/SKILL.md — .pdf
sn-da-non-spreadsheet-analysis/capability/ppt-analysis/SKILL.md — .pptx/.ppt| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | fail→fail | 12,142 | 8,549 | -30% | 1 | 1 | 0% | 632 | 1,660 | +163% | 0 | 0 | — |
case-01 | fail→fail | 6,614 | 9,833 | +49% | 1 | 1 | 0% | 336 | 1,818 | +441% | 0 | 0 | — |
case-02 | fail→fail | 8,685 | 8,219 | -5% | 1 | 1 | 0% | 389 | 1,761 | +353% | 0 | 0 | — |
case-03 | fail→fail | 9,889 | 8,056 | -19% | 1 | 1 | 0% | 675 | 1,693 | +151% | 0 | 0 | — |
case-04 | fail→fail | 6,844 | 8,592 | +26% | 1 | 1 | 0% | 1,066 | 2,520 | +136% | 0 | 0 | — |
case-05 | fail→pass | 13,037 | 8,239 | -37% | 1 | 1 | 0% | 2,024 | 2,856 | +41% | 0 | 0 | — |
case-06 | fail→pass | 12,093 | 3,610 | -70% | 1 | 1 | 0% | 1,810 | 1,967 | +9% | 0 | 0 | — |
case-07 | fail→fail | 9,846 | 10,327 | +5% | 1 | 1 | 0% | 564 | 1,836 | +226% | 0 | 0 | — |
case-09 | pass→pass | 4,829 | 9,841 | +104% | 1 | 1 | 0% | 847 | 2,482 | +193% | 0 | 0 | — |
case-10 | pass→fail | 3,101 | 13,092 | +322% | 1 | 1 | 0% | 567 | 1,906 | +236% | 0 | 0 | — |
case-11 | fail→fail | 4,861 | 8,869 | +82% | 1 | 1 | 0% | 771 | 1,649 | +114% | 0 | 0 | — |
case-12 | fail→fail | 13,491 | 7,512 | -44% | 1 | 1 | 0% | 2,624 | 1,621 | -38% | 0 | 0 | — |
case-13 | pass→pass | 6,249 | 2,158 | -65% | 1 | 1 | 0% | 984 | 1,634 | +66% | 0 | 0 | — |
case-14 | pass→pass | 9,227 | 4,445 | -52% | 1 | 1 | 0% | 1,515 | 2,124 | +40% | 0 | 0 | — |
case-15 | pass→pass | 8,596 | 2,333 | -73% | 1 | 1 | 0% | 1,357 | 1,594 | +17% | 0 | 0 | — |
case-16 | fail→fail | 5,617 | 5,262 | -6% | 1 | 1 | 0% | 796 | 2,151 | +170% | 0 | 0 | — |
case-17 | pass→pass | 10,370 | 10,188 | -2% | 1 | 1 | 0% | 1,633 | 3,224 | +97% | 0 | 0 | — |
case-18 | fail→pass | 9,807 | 2,195 | -78% | 1 | 1 | 0% | 1,424 | 1,673 | +17% | 0 | 0 | — |
case-19 | pass→pass | 14,093 | 9,934 | -30% | 1 | 1 | 0% | 2,167 | 2,168 | +0% | 0 | 0 | — |
case-20 | fail→fail | 9,877 | 10,863 | +10% | 1 | 1 | 0% | 1,000 | 2,129 | +113% | 0 | 0 | — |
case-21 | pass→fail | 5,223 | 9,459 | +81% | 1 | 1 | 0% | 861 | 1,761 | +105% | 0 | 0 | — |
case-22 | fail→fail | 17,045 | 7,036 | -59% | 1 | 1 | 0% | 3,282 | 1,610 | -51% | 0 | 0 | — |
case-23 | fail→fail | 28,603 | 6,602 | -77% | 1 | 1 | 0% | 6,183 | 1,619 | -74% | 0 | 0 | — |
case-24 | fail→fail | 7,310 | 9,724 | +33% | 1 | 1 | 0% | 1,319 | 3,046 | +131% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 12 counted toward the lift figure. The other 12 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +4 percentage points is the difference between those two pass rates over the 12 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.