Install any skill in seconds. Free to start, no credit card required.
Get Started Free →验证信源可信度的研究过程知识,指导 agent 在涉及事实声明时交叉核验信源
.claude/skills/hezaohezao-source-verification/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-24 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-09 | ✓→✗ | ▼ Worse | -80% | 0% |
| case-12 | ✓→✗ | ▼ Worse | -74% | 0% |
| case-17 | ✓→✗ | ▼ Worse | -67% | 0% |
研究涉及事实声明、统计数据、引用或任何需要可信背书的信息时。不适用于纯推理或用户主观偏好类问题。
任何事实声明在写入 observation 前,至少需一个独立信源交叉核验。 单一信源的可信度不足以支撑结论。
web_search 找 2 个以上独立来源(不同机构/作者/时间)。优先一手来源(官方报告、论文原文)而非二手转述。browse_page 读全文(非 snippet),比对关键数字/日期/结论是否一致。| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 19,116 | 9,290 | -51% | 1 | 1 | 0% | 3,164 | 896 | -72% | 0 | 0 | — |
case-02 | fail→fail | 18,407 | 9,088 | -51% | 1 | 1 | 0% | 2,635 | 938 | -64% | 0 | 0 | — |
case-03 | pass→pass | 10,346 | 15,072 | +46% | 1 | 1 | 0% | 1,915 | 2,906 | +52% | 0 | 0 | — |
case-04 | pass→pass | 27,423 | 34,314 | +25% | 1 | 1 | 0% | 3,347 | 4,922 | +47% | 0 | 0 | — |
case-05 | pass→pass | 14,271 | 10,326 | -28% | 1 | 1 | 0% | 2,722 | 2,466 | -9% | 0 | 0 | — |
case-06 | pass→pass | 11,315 | 12,960 | +15% | 1 | 1 | 0% | 1,594 | 1,940 | +22% | 0 | 0 | — |
case-07 | pass→pass | 12,698 | 7,907 | -38% | 1 | 1 | 0% | 1,651 | 1,327 | -20% | 0 | 0 | — |
case-08 | fail→fail | 14,358 | 11,345 | -21% | 1 | 1 | 0% | 2,340 | 880 | -62% | 0 | 0 | — |
case-09 | pass→fail | 22,099 | 8,912 | -60% | 1 | 1 | 0% | 3,465 | 704 | -80% | 0 | 0 | — |
case-10 | pass→pass | 15,357 | 18,577 | +21% | 1 | 1 | 0% | 2,162 | 2,868 | +33% | 0 | 0 | — |
case-11 | pass→pass | 15,130 | 12,969 | -14% | 1 | 1 | 0% | 2,052 | 1,838 | -10% | 0 | 0 | — |
case-12 | pass→fail | 15,498 | 8,848 | -43% | 1 | 1 | 0% | 2,514 | 649 | -74% | 0 | 0 | — |
case-13 | fail→fail | 13,893 | 10,230 | -26% | 1 | 1 | 0% | 1,859 | 1,696 | -9% | 0 | 0 | — |
case-14 | pass→pass | 19,832 | 20,310 | +2% | 1 | 1 | 0% | 3,015 | 3,223 | +7% | 0 | 0 | — |
case-15 | pass→pass | 14,397 | 13,440 | -7% | 1 | 1 | 0% | 2,106 | 2,030 | -4% | 0 | 0 | — |
case-16 | fail→pass | 19,530 | 15,477 | -21% | 1 | 1 | 0% | 2,767 | 2,823 | +2% | 0 | 0 | — |
case-17 | pass→fail | 20,024 | 25,537 | +28% | 1 | 1 | 0% | 2,903 | 968 | -67% | 0 | 0 | — |
case-18 | pass→pass | 17,804 | 16,056 | -10% | 1 | 1 | 0% | 2,637 | 2,115 | -20% | 0 | 0 | — |
case-19 | fail→fail | 22,239 | 10,996 | -51% | 1 | 1 | 0% | 3,024 | 872 | -71% | 0 | 0 | — |
case-20 | fail→fail | 14,125 | 13,428 | -5% | 1 | 1 | 0% | 1,921 | 2,154 | +12% | 0 | 0 | — |
case-21 | pass→fail | 18,409 | 11,760 | -36% | 1 | 1 | 0% | 2,578 | 999 | -61% | 0 | 0 | — |
case-22 | pass→pass | 21,214 | 18,100 | -15% | 1 | 1 | 0% | 2,979 | 2,944 | -1% | 0 | 0 | — |
case-23 | pass→fail | 14,851 | 8,932 | -40% | 1 | 1 | 0% | 2,506 | 768 | -69% | 0 | 0 | — |
case-24 | fail→pass | 14,397 | 14,970 | +4% | 1 | 1 | 0% | 2,230 | 2,502 | +12% | 0 | 0 | — |
case-25 | pass→pass | 14,855 | 10,264 | -31% | 1 | 1 | 0% | 2,070 | 1,641 | -21% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 16 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -12 percentage points is the difference between those two pass rates over the 16 comparable cases. 8 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.