Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Operate as an agentic engineer using eval-first execution, decomposition, and cost-aware model routing. Use when AI agents perform most implementation work and humans enforce quality and risk controls.
.claude/skills/affaan-m-agentic-engineering/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -61% | 0% |
在 AI 智能体执行大部分实施工作、而人类负责质量与风险控制的工程工作流中使用此技能。
应用 15 分钟单元规则:
优先审查:
当自动化格式化/代码检查工具已强制执行代码风格时,不要在仅涉及风格分歧的审查上浪费周期。
按任务跟踪:
仅当较低层级的模型失败且存在清晰的推理差距时,才升级模型层级。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | pass→pass | 13,932 | 8,302 | -40% | 1 | 1 | 0% | 2,753 | 1,878 | -32% | 0 | 0 | — |
case-04 | pass→pass | 15,919 | 14,299 | -10% | 1 | 1 | 0% | 2,753 | 2,983 | +8% | 0 | 0 | — |
case-05 | pass→pass | 14,937 | 14,438 | -3% | 1 | 1 | 0% | 2,656 | 2,803 | +6% | 0 | 0 | — |
case-06 | pass→fail | 12,424 | 12,439 | +0% | 1 | 1 | 0% | 2,135 | 2,417 | +13% | 0 | 0 | — |
case-01 | fail→pass | 25,053 | 17,380 | -31% | 1 | 1 | 0% | 4,416 | 3,348 | -24% | 0 | 0 | — |
case-02 | fail→pass | 17,987 | 16,542 | -8% | 1 | 1 | 0% | 3,215 | 3,336 | +4% | 0 | 0 | — |
case-03 | fail→fail | 24,595 | 19,107 | -22% | 1 | 1 | 0% | 4,067 | 3,792 | -7% | 0 | 0 | — |
case-07 | fail→pass | 15,488 | 13,917 | -10% | 1 | 1 | 0% | 2,751 | 2,897 | +5% | 0 | 0 | — |
case-08 | pass→pass | 13,424 | 10,617 | -21% | 1 | 1 | 0% | 2,321 | 2,372 | +2% | 0 | 0 | — |
case-09 | pass→pass | 7,017 | 3,855 | -45% | 1 | 1 | 0% | 1,292 | 1,054 | -18% | 0 | 0 | — |
case-10 | pass→pass | 7,330 | 3,096 | -58% | 1 | 1 | 0% | 1,302 | 848 | -35% | 0 | 0 | — |
case-11 | fail→pass | 13,631 | 7,228 | -47% | 1 | 1 | 0% | 2,323 | 1,615 | -30% | 0 | 0 | — |
case-12 | fail→pass | 13,647 | 2,763 | -80% | 1 | 1 | 0% | 2,262 | 880 | -61% | 0 | 0 | — |
case-13 | pass→pass | 9,734 | 4,554 | -53% | 1 | 1 | 0% | 1,590 | 1,165 | -27% | 0 | 0 | — |
case-14 | pass→pass | 8,208 | 5,985 | -27% | 1 | 1 | 0% | 1,383 | 1,382 | -0% | 0 | 0 | — |
case-15 | pass→pass | 5,370 | 3,012 | -44% | 1 | 1 | 0% | 871 | 853 | -2% | 0 | 0 | — |
case-16 | pass→pass | 9,933 | 4,417 | -56% | 1 | 1 | 0% | 1,660 | 1,182 | -29% | 0 | 0 | — |
case-17 | pass→pass | 13,808 | 12,624 | -9% | 1 | 1 | 0% | 2,353 | 2,526 | +7% | 0 | 0 | — |
case-19 | fail→pass | 9,150 | 6,177 | -32% | 1 | 1 | 0% | 1,502 | 1,519 | +1% | 0 | 0 | — |
case-20 | fail→pass | 15,510 | 10,693 | -31% | 1 | 1 | 0% | 2,825 | 2,233 | -21% | 0 | 0 | — |
case-21 | fail→pass | 6,913 | 3,043 | -56% | 1 | 1 | 0% | 1,118 | 996 | -11% | 0 | 0 | — |
case-22 | pass→pass | 17,659 | 12,911 | -27% | 1 | 1 | 0% | 2,511 | 2,471 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.