Install any skill in seconds. Free to start, no credit card required.
Get Started Free →收集系统全链路操作日志,生成可追溯的执行证据链与向量索引,是多智能体系统可观测性与安全审计的底座。
.claude/skills/anbeime-antinet-provenance/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-01 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -3% | 0% |
trace_id / 时间窗 / 官署维度的检索接口,支撑审计面板与回滚定位。event:操作事件,含 actor(官署名)、action、target、timestamp、trace_idpayload:(可选)与事件关联的结构化产物引用evidence_chain:按 trace_id 串联的可追溯证据链vector_index:用于语义检索的向量索引条目query_api:按条件检索历史证据的接口描述stale。scripts/run_provenance.pycore.runtime.AgentSession.run_stage("provenance"),调用 memory.taishige.TaiShiGeAgent.writeback(把全链路事件与四色卡片回流证据链,纯 Python 可离线)。python skills/provenance/scripts/run_provenance.pyexamples/snse_survey/provenance/(trace.jsonl + trace_summary.json,含每一条军机处派发与官署执行事件)。| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | fail→pass | 17,558 | 16,555 | -6% | 1 | 1 | 0% | 2,090 | 2,912 | +39% | 0 | 0 | — |
case-12 | fail→pass | 20,777 | 17,648 | -15% | 1 | 1 | 0% | 2,704 | 2,838 | +5% | 0 | 0 | — |
case-01 | fail→pass | 36,823 | 34,685 | -6% | 1 | 1 | 0% | 7,466 | 4,509 | -40% | 0 | 0 | — |
case-02 | fail→pass | 18,254 | 20,539 | +13% | 1 | 1 | 0% | 2,348 | 3,587 | +53% | 0 | 0 | — |
case-03 | pass→pass | 25,614 | 26,981 | +5% | 1 | 1 | 0% | 3,923 | 4,885 | +25% | 0 | 0 | — |
case-04 | pass→pass | 23,942 | 22,687 | -5% | 1 | 1 | 0% | 3,422 | 3,537 | +3% | 0 | 0 | — |
case-05 | pass→pass | 27,294 | 37,287 | +37% | 1 | 1 | 0% | 4,160 | 6,378 | +53% | 0 | 0 | — |
case-06 | fail→pass | 20,392 | 15,105 | -26% | 1 | 1 | 0% | 2,289 | 2,211 | -3% | 0 | 0 | — |
case-07 | pass→pass | 17,484 | 9,204 | -47% | 1 | 1 | 0% | 1,991 | 1,215 | -39% | 0 | 0 | — |
case-08 | fail→pass | 17,868 | 7,966 | -55% | 1 | 1 | 0% | 2,084 | 979 | -53% | 0 | 0 | — |
case-09 | fail→pass | 27,161 | 22,590 | -17% | 1 | 1 | 0% | 4,378 | 4,362 | -0% | 0 | 0 | — |
case-10 | fail→pass | 21,751 | 21,866 | +1% | 1 | 1 | 0% | 2,732 | 3,402 | +25% | 0 | 0 | — |
case-13 | fail→pass | 20,781 | 7,366 | -65% | 1 | 1 | 0% | 2,469 | 935 | -62% | 0 | 0 | — |
case-14 | pass→pass | 14,580 | 8,306 | -43% | 1 | 1 | 0% | 1,546 | 1,041 | -33% | 0 | 0 | — |
case-15 | fail→pass | 25,503 | 19,862 | -22% | 1 | 1 | 0% | 4,574 | 4,851 | +6% | 0 | 0 | — |
case-16 | fail→pass | 26,041 | 7,307 | -72% | 1 | 1 | 0% | 3,505 | 875 | -75% | 0 | 0 | — |
case-17 | fail→fail | 8,483 | 1,846 | -78% | 1 | 1 | 0% | 1,331 | 860 | -35% | 0 | 0 | — |
case-18 | fail→pass | 11,233 | 1,845 | -84% | 1 | 1 | 0% | 1,775 | 839 | -53% | 0 | 0 | — |
case-19 | fail→pass | 11,720 | 6,027 | -49% | 1 | 1 | 0% | 1,813 | 1,454 | -20% | 0 | 0 | — |
case-20 | pass→fail | 17,062 | 20,148 | +18% | 1 | 1 | 0% | 2,712 | 4,360 | +61% | 0 | 0 | — |
case-21 | pass→pass | 8,872 | 2,431 | -73% | 1 | 1 | 0% | 1,348 | 878 | -35% | 0 | 0 | — |
case-22 | fail→pass | 17,745 | 13,401 | -24% | 1 | 1 | 0% | 2,785 | 2,790 | +0% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.