Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns. Use when building data pipelines, ETL workflows, stream processing, or data quality checks.
.claude/skills/telagod-engineering-data-pipelines/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-01 | ✓→✓ | = Same ✓ | -21% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 14% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 10% | 0% |
> 判断先于执行:决定「是否做 / 选什么 / 如何取舍」(栈、方案、架构、权衡)前,先读领域判断内核 skills/_kernel/backend/SKILL.md——它管 judgment,本秘典管 execution;冲突时以内核判断为准。
编排:Airflow(调度) | Dagster(资产) | Prefect(现代流)
流处理:Kafka Streams(嵌入式) | Flink(集群) | Spark Streaming
质量:Great Expectations | dbt tests | Soda Core幂等(UPSERT/分区覆盖) | 增量(WHERE updated_at > last_run) | 事件驱动触发 | 跨 DAG 依赖 | 数据血缘(ref()/Asset deps)
时间语义选择 | Watermark 乱序容忍 | 状态 TTL 防膨胀 | Checkpoint 间隔 | 端到端 Exactly-Once | 背压监控
分层验证(源→转换→目标) | 完整性+准确性+一致性 | 及时性阈值 | 加权评分 | 告警(Slack/PagerDuty)
工具对比、API 用法、质量维度详见 references/details.md
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 12,925 | 17,532 | +36% | 1 | 1 | 0% | 1,951 | 1,549 | -21% | 0 | 0 | — |
case-02 | pass→pass | 10,579 | 10,158 | -4% | 1 | 1 | 0% | 1,671 | 1,902 | +14% | 0 | 0 | — |
case-03 | pass→pass | 15,445 | 14,598 | -5% | 1 | 1 | 0% | 2,352 | 2,578 | +10% | 0 | 0 | — |
case-04 | pass→pass | 14,854 | 13,146 | -11% | 1 | 1 | 0% | 2,395 | 2,243 | -6% | 0 | 0 | — |
case-05 | pass→pass | 9,326 | 8,084 | -13% | 1 | 1 | 0% | 1,412 | 1,473 | +4% | 0 | 0 | — |
case-06 | fail→pass | 6,802 | 8,064 | +19% | 1 | 1 | 0% | 1,005 | 1,633 | +62% | 0 | 0 | — |
case-07 | pass→pass | 9,902 | 11,818 | +19% | 1 | 1 | 0% | 1,467 | 2,252 | +54% | 0 | 0 | — |
case-08 | pass→pass | 7,131 | 5,592 | -22% | 1 | 1 | 0% | 1,147 | 1,147 | 0% | 0 | 0 | — |
case-09 | pass→pass | 5,611 | 6,426 | +15% | 1 | 1 | 0% | 843 | 1,233 | +46% | 0 | 0 | — |
case-10 | pass→pass | 15,061 | 12,684 | -16% | 1 | 1 | 0% | 2,166 | 2,184 | +1% | 0 | 0 | — |
case-11 | pass→pass | 17,319 | 15,436 | -11% | 1 | 1 | 0% | 2,505 | 2,623 | +5% | 0 | 0 | — |
case-12 | pass→pass | 15,602 | 16,385 | +5% | 1 | 1 | 0% | 2,342 | 2,717 | +16% | 0 | 0 | — |
case-13 | pass→pass | 19,196 | 18,468 | -4% | 1 | 1 | 0% | 2,955 | 3,278 | +11% | 0 | 0 | — |
case-14 | pass→pass | 17,843 | 18,835 | +6% | 1 | 1 | 0% | 2,840 | 3,363 | +18% | 0 | 0 | — |
case-15 | pass→pass | 16,997 | 15,615 | -8% | 1 | 1 | 0% | 2,452 | 2,772 | +13% | 0 | 0 | — |
case-16 | pass→pass | 4,390 | 5,009 | +14% | 1 | 1 | 0% | 624 | 1,043 | +67% | 0 | 0 | — |
case-17 | pass→pass | 13,122 | 12,277 | -6% | 1 | 1 | 0% | 2,256 | 2,501 | +11% | 0 | 0 | — |
case-18 | pass→pass | 6,890 | 6,585 | -4% | 1 | 1 | 0% | 1,051 | 1,300 | +24% | 0 | 0 | — |
case-19 | pass→pass | 13,642 | 6,605 | -52% | 1 | 1 | 0% | 1,846 | 1,293 | -30% | 0 | 0 | — |
case-20 | fail→pass | 16,671 | 15,299 | -8% | 1 | 1 | 0% | 2,394 | 2,408 | +1% | 0 | 0 | — |
case-21 | pass→pass | 17,689 | 23,551 | +33% | 1 | 1 | 0% | 2,587 | 3,856 | +49% | 0 | 0 | — |
case-22 | pass→pass | 18,033 | 12,720 | -29% | 1 | 1 | 0% | 2,661 | 2,335 | -12% | 0 | 0 | — |
case-23 | pass→pass | 20,122 | 22,589 | +12% | 1 | 1 | 0% | 3,348 | 4,622 | +38% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.