Install any skill in seconds. Free to start, no credit card required.
Get Started Free →構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。
.claude/skills/affaan-m-regex-vs-llm-structured-text/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 44% | 0% |
一个用于解析结构化文本(测验、表单、发票、文档)的实用决策框架。核心见解是:正则表达式能以低成本、确定性的方式处理 95-98% 的情况。将昂贵的 LLM 调用留给剩余的边缘情况。
文本格式是否一致且重复?
├── 是 (>90% 遵循某种模式) → 从正则表达式开始
│ ├── 正则表达式处理 95%+ → 完成,无需 LLM
│ └── 正则表达式处理 <95% → 仅为边缘情况添加 LLM
└── 否 (自由格式,高度可变) → 直接使用 LLM[正则表达式解析器] ─── 提取结构(95-98% 准确率)
│
▼
[文本清理器] ─── 去除噪声(标记、页码、伪影)
│
▼
[置信度评分器] ─── 标记低置信度提取项
│
├── 高置信度(≥0.95)→ 直接输出
│
└── 低置信度(<0.95)→ [LLM 验证器] → 输出pythonimport re from dataclasses import dataclass @dataclass(frozen=True) class ParsedItem: id: str text: str choices: tuple[str, ...] answer: str confidence: float = 1.0 def parse_structured_text(content: str) -> list[ParsedItem]: """Parse structured text using regex patterns.""" pattern = re.compile( r"(?P<id>\d+)\.\s*(?P<text>.+?)\n" r"(?P<choices>(?:[A-D]\..+?\n)+)" r"Answer:\s*(?P<answer>[A-D])", re.MULTILINE | re.DOTALL, ) items = [] for match in pattern.finditer(content): choices = tuple( c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices")) ) items.append(ParsedItem( id=match.group("id"), text=match.group("text").strip(), choices=choices, answer=match.group("answer"), )) return items
标记可能需要 LLM 审核的项:
python@dataclass(frozen=True) class ConfidenceFlag: item_id: str score: float reasons: tuple[str, ...] def score_confidence(item: ParsedItem) -> ConfidenceFlag: """Score extraction confidence and flag issues.""" reasons = [] score = 1.0 if len(item.choices) < 3: reasons.append("few_choices") score -= 0.3 if not item.answer: reasons.append("missing_answer") score -= 0.5 if len(item.text) < 10: reasons.append("short_text") score -= 0.2 return ConfidenceFlag( item_id=item.id, score=max(0.0, score), reasons=tuple(reasons), ) def identify_low_confidence( items: list[ParsedItem], threshold: float = 0.95, ) -> list[ConfidenceFlag]: """Return items below confidence threshold.""" flags = [score_confidence(item) for item in items] return [f for f in flags if f.score < threshold]
pythondef validate_with_llm( item: ParsedItem, original_text: str, client, ) -> ParsedItem: """Use LLM to fix low-confidence extractions.""" response = client.messages.create( model="claude-haiku-4-5-20251001", # Cheapest model for validation max_tokens=500, messages=[{ "role": "user", "content": ( f"Extract the question, choices, and answer from this text.\n\n" f"Text: {original_text}\n\n" f"Current extraction: {item}\n\n" f"Return corrected JSON if needed, or 'CORRECT' if accurate." ), }], ) # Parse LLM response and return corrected item... return corrected_item
pythondef process_document( content: str, *, llm_client=None, confidence_threshold: float = 0.95, ) -> list[ParsedItem]: """Full pipeline: regex -> confidence check -> LLM for edge cases.""" # Step 1: Regex extraction (handles 95-98%) items = parse_structured_text(content) # Step 2: Confidence scoring low_confidence = identify_low_confidence(items, confidence_threshold) if not low_confidence or llm_client is None: return items # Step 3: LLM validation (only for flagged items) low_conf_ids = {f.item_id for f in low_confidence} result = [] for item in items: if item.id in low_conf_ids: result.append(validate_with_llm(item, content, llm_client)) else: result.append(item) return result
来自一个生产中的测验解析管道(410 个项目):
| 指标 | 值 | |--------|-------| | 正则表达式成功率 | 98.0% | | 低置信度项目 | 8 (2.0%) | | 所需 LLM 调用次数 | ~5 | | 相比全 LLM 的成本节省 | ~95% | | 测试覆盖率 | 93% |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 43,728 | 49,071 | +12% | 1 | 1 | 0% | 3,732 | 5,878 | +58% | 0 | 0 | — |
case-02 | fail→fail | 24,360 | 26,888 | +10% | 1 | 1 | 0% | 3,891 | 5,938 | +53% | 0 | 0 | — |
case-03 | fail→pass | 17,532 | 9,045 | -48% | 1 | 1 | 0% | 2,595 | 3,481 | +34% | 0 | 0 | — |
case-04 | pass→pass | 14,957 | 13,235 | -12% | 1 | 1 | 0% | 2,591 | 4,058 | +57% | 0 | 0 | — |
case-05 | pass→pass | 16,667 | 16,296 | -2% | 1 | 1 | 0% | 2,529 | 4,397 | +74% | 0 | 0 | — |
case-06 | pass→pass | 15,391 | 13,798 | -10% | 1 | 1 | 0% | 2,226 | 3,963 | +78% | 0 | 0 | — |
case-07 | fail→pass | 17,110 | 13,116 | -23% | 1 | 1 | 0% | 2,640 | 3,803 | +44% | 0 | 0 | — |
case-08 | fail→pass | 14,732 | 5,360 | -64% | 1 | 1 | 0% | 2,278 | 2,611 | +15% | 0 | 0 | — |
case-09 | fail→pass | 19,853 | 15,727 | -21% | 1 | 1 | 0% | 3,029 | 4,461 | +47% | 0 | 0 | — |
case-10 | pass→pass | 14,189 | 15,221 | +7% | 1 | 1 | 0% | 2,224 | 4,301 | +93% | 0 | 0 | — |
case-11 | fail→pass | 10,332 | 3,623 | -65% | 1 | 1 | 0% | 1,733 | 2,494 | +44% | 0 | 0 | — |
case-12 | pass→pass | 8,470 | 5,497 | -35% | 1 | 1 | 0% | 1,398 | 2,796 | +100% | 0 | 0 | — |
case-13 | pass→pass | 12,821 | 7,687 | -40% | 1 | 1 | 0% | 2,052 | 3,283 | +60% | 0 | 0 | — |
case-14 | fail→pass | 9,375 | 2,639 | -72% | 1 | 1 | 0% | 1,498 | 2,241 | +50% | 0 | 0 | — |
case-15 | fail→pass | 12,109 | 12,523 | +3% | 1 | 1 | 0% | 2,120 | 3,721 | +76% | 0 | 0 | — |
case-16 | pass→pass | 9,851 | 7,407 | -25% | 1 | 1 | 0% | 1,499 | 2,993 | +100% | 0 | 0 | — |
case-17 | pass→pass | 11,877 | 10,987 | -7% | 1 | 1 | 0% | 1,975 | 3,541 | +79% | 0 | 0 | — |
case-18 | pass→pass | 9,377 | 5,652 | -40% | 1 | 1 | 0% | 1,380 | 2,790 | +102% | 0 | 0 | — |
case-19 | pass→pass | 13,684 | 10,538 | -23% | 1 | 1 | 0% | 2,271 | 3,606 | +59% | 0 | 0 | — |
case-20 | pass→pass | 16,915 | 18,044 | +7% | 1 | 1 | 0% | 2,675 | 5,148 | +92% | 0 | 0 | — |
case-21 | pass→pass | 17,154 | 17,873 | +4% | 1 | 1 | 0% | 2,961 | 4,958 | +67% | 0 | 0 | — |
case-22 | fail→pass | 17,836 | 20,007 | +12% | 1 | 1 | 0% | 3,372 | 5,589 | +66% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.