Install any skill in seconds. Free to start, no credit card required.
Get Started Free →对Excel文件进行文本标准化清洗(如去除异常前缀、提取纯中文字符等),并,最终输出清洗后的Excel文件并提供下载链接。
.claude/skills/opensensenova-text-normalization-and-large-file-processing/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 94% | 0% |
> This sub-skill covers one capability of the Excel workflow. For reading/counting/Parquet optimization, see the parent workflow SKILL.md.
Step1 识别并清洗包含前缀符号的异常数值字段,统一转换为整数类型;同时使用正则表达式清洗文本字段,仅保留 Unicode 范围内的中文字符。
pythonimport re import numpy as np target_numeric_col = '需要转数字的文本列' # 示例:'获赞' target_text_col = '需要提取中文的列' # 示例:'收货人' # 1. 清洗包含前缀符号的数值字段 prefix_patterns = ['.', 'I ', '■ ', '一 ', '_', '. '] def clean_numeric_with_prefix(value): val_str = str(value).strip() if val_str in ['None', 'nan', '', 'nan']: return np.nan for prefix in prefix_patterns: if val_str.startswith(prefix): val_str = val_str[len(prefix):].strip() break if val_str == '': return np.nan try: return int(val_str) except ValueError: return np.nan # 2. 清洗文本字段,仅保留 Unicode 范围内的中文字符(\u4e00-\u9fff) def clean_chinese_name(name): if pd.isna(name): return name s = str(name) chinese_chars = re.findall(r'[\u4e00-\u9fff]', s) cleaned = ''.join(chinese_chars) return cleaned if cleaned else '' if target_numeric_col in df.columns: df[f'{target_numeric_col}_清洗后'] = df[target_numeric_col].apply(clean_numeric_with_prefix) if target_text_col in df.columns: df[f'{target_text_col}_清洗后'] = df[target_text_col].apply(clean_chinese_name)
Step2 将清洗后的结果保存为 Excel 文件,在报告中提供下载链接,并执行内存清理以应对大文件处理时的内存压力。
pythonoutput_path = '/mnt/data/标准化清洗结果.xlsx' # 保存清洗结果 df.to_excel(output_path, index=False, engine='openpyxl') print(f'清洗结果已保存到: {output_path}') # 生成可下载链接 print(f'[下载清洗结果表](sandbox:{output_path})') # 内存清理 if 'df' in locals(): del df gc.collect()
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,170 | 5,602 | -50% | 1 | 1 | 0% | 2,362 | 1,894 | -20% | 0 | 0 | — |
case-02 | fail→pass | 12,935 | 11,224 | -13% | 1 | 1 | 0% | 2,606 | 2,577 | -1% | 0 | 0 | — |
case-03 | fail→pass | 12,527 | 8,862 | -29% | 1 | 1 | 0% | 2,657 | 2,604 | -2% | 0 | 0 | — |
case-04 | pass→pass | 12,734 | 11,874 | -7% | 1 | 1 | 0% | 2,592 | 2,989 | +15% | 0 | 0 | — |
case-05 | pass→pass | 14,481 | 12,311 | -15% | 1 | 1 | 0% | 2,753 | 3,150 | +14% | 0 | 0 | — |
case-06 | pass→pass | 10,286 | 8,371 | -19% | 1 | 1 | 0% | 2,138 | 2,306 | +8% | 0 | 0 | — |
case-07 | fail→fail | 11,335 | 5,522 | -51% | 1 | 1 | 0% | 1,939 | 1,538 | -21% | 0 | 0 | — |
case-08 | pass→pass | 10,106 | 4,533 | -55% | 1 | 1 | 0% | 1,908 | 1,565 | -18% | 0 | 0 | — |
case-09 | fail→pass | 14,570 | 2,199 | -85% | 1 | 1 | 0% | 2,139 | 1,035 | -52% | 0 | 0 | — |
case-10 | fail→pass | 3,470 | 2,787 | -20% | 1 | 1 | 0% | 564 | 1,093 | +94% | 0 | 0 | — |
case-19 | fail→pass | 10,318 | 4,068 | -61% | 1 | 1 | 0% | 1,600 | 1,136 | -29% | 0 | 0 | — |
case-11 | pass→pass | 10,025 | 3,339 | -67% | 1 | 1 | 0% | 1,903 | 1,083 | -43% | 0 | 0 | — |
case-12 | fail→pass | 7,777 | 3,025 | -61% | 1 | 1 | 0% | 1,400 | 1,040 | -26% | 0 | 0 | — |
case-13 | fail→pass | 13,621 | 12,479 | -8% | 1 | 1 | 0% | 2,584 | 3,210 | +24% | 0 | 0 | — |
case-14 | fail→pass | 17,431 | 6,029 | -65% | 1 | 1 | 0% | 2,504 | 1,833 | -27% | 0 | 0 | — |
case-15 | fail→fail | 14,379 | 19,063 | +33% | 1 | 1 | 0% | 2,090 | 3,288 | +57% | 0 | 0 | — |
case-16 | pass→pass | 10,574 | 2,938 | -72% | 1 | 1 | 0% | 1,509 | 1,168 | -23% | 0 | 0 | — |
case-17 | pass→pass | 7,033 | 4,375 | -38% | 1 | 1 | 0% | 1,353 | 1,279 | -5% | 0 | 0 | — |
case-18 | pass→pass | 9,280 | 2,455 | -74% | 1 | 1 | 0% | 1,683 | 1,093 | -35% | 0 | 0 | — |
case-20 | fail→pass | 14,837 | 10,225 | -31% | 1 | 1 | 0% | 2,434 | 2,348 | -4% | 0 | 0 | — |
case-21 | fail→pass | 4,501 | 1,538 | -66% | 1 | 1 | 0% | 798 | 879 | +10% | 0 | 0 | — |
case-22 | fail→pass | 4,982 | 2,280 | -54% | 1 | 1 | 0% | 900 | 1,005 | +12% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.