Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Technical leadership guidance for engineering teams, architecture decisions, and technology strategy. Includes tech debt analyzer, team scaling calculator, engineering metrics frameworks, technology evaluation tools, and ADR templates. Use when assessing technical debt, scaling engineering teams, evaluating technologies, making architecture decisions, establishing engineering metrics, or when user mentions CTO, tech debt, technical debt, team scaling, architecture decisions, technology evaluatio
.claude/skills/ibrahim-3d-cto-advisor/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 73% | 0% |
Strategic frameworks and tools for technology leadership, team scaling, and engineering excellence.
CTO, chief technology officer, technical leadership, tech debt, technical debt, engineering team, team scaling, architecture decisions, technology evaluation, engineering metrics, DORA metrics, ADR, architecture decision records, technology strategy, engineering leadership, engineering organization, team structure, hiring plan, technical strategy, vendor evaluation, technology selection
bashpython scripts/tech_debt_analyzer.py
Analyzes system architecture and provides prioritized debt reduction plan.
bashpython scripts/team_scaling_calculator.py
Calculates optimal hiring plan and team structure for growth.
Review references/architecture_decision_records.md for ADR templates and examples.
Use framework in references/technology_evaluation_framework.md for vendor selection.
Implement KPIs from references/engineering_metrics.md for team performance tracking.
bash# Assess current debt python scripts/tech_debt_analyzer.py # Allocate capacity - Critical debt: 40% capacity - High debt: 25% capacity - Medium debt: 15% capacity - Low debt: Ongoing maintenance
bash# Calculate scaling needs python scripts/team_scaling_calculator.py # Key ratios to maintain: - Manager:Engineer = 1:8 - Senior:Mid:Junior = 3:4:2 - Product:Engineering = 1:10 - QA:Engineering = 1.5:10
Use ADR template from references/architecture_decision_records.md:
Follow framework in references/technology_evaluation_framework.md:
From references/engineering_metrics.md:
DORA Metrics (Deploy to production targets):
Quality Metrics:
Team Health:
Monthly:
Quarterly:
1. Executive Summary (1 slide)
2. Current State Assessment (2 slides)
3. Vision & Strategy (2 slides)
4. Roadmap & Milestones (3 slides)
5. Investment Required (1 slide)
6. Risks & Mitigation (1 slide)
7. Success Metrics (1 slide)1. Wins & Recognition (5 min)
2. Metrics Review (5 min)
3. Strategic Updates (10 min)
4. Demo/Deep Dive (15 min)
5. Q&A (10 min)Subject: Engineering Update - [Month]
Highlights:
* [Major achievement]
* [Key metric improvement]
* [Strategic progress]
Challenges:
* [Issue and mitigation]
Next Month:
* [Priority 1]
* [Priority 2]
Detailed metrics attached.Technical Excellence
Team Success
Business Impact
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 19,674 | 35,207 | +79% | 1 | 1 | 0% | 3,703 | 4,790 | +29% | 0 | 0 | — |
case-02 | fail→fail | 21,563 | 20,829 | -3% | 1 | 1 | 0% | 3,655 | 5,955 | +63% | 0 | 0 | — |
case-03 | fail→fail | 15,521 | 16,910 | +9% | 1 | 1 | 0% | 2,642 | 4,798 | +82% | 0 | 0 | — |
case-04 | pass→pass | 14,127 | 13,138 | -7% | 1 | 1 | 0% | 2,349 | 4,310 | +83% | 0 | 0 | — |
case-05 | pass→pass | 13,918 | 12,048 | -13% | 1 | 1 | 0% | 2,678 | 4,346 | +62% | 0 | 0 | — |
case-06 | pass→pass | 5,745 | 4,682 | -19% | 1 | 1 | 0% | 1,096 | 2,935 | +168% | 0 | 0 | — |
case-07 | fail→pass | 11,690 | 5,149 | -56% | 1 | 1 | 0% | 1,840 | 2,824 | +53% | 0 | 0 | — |
case-08 | fail→pass | 12,820 | 3,500 | -73% | 1 | 1 | 0% | 2,129 | 2,593 | +22% | 0 | 0 | — |
case-09 | pass→pass | 13,599 | 11,058 | -19% | 1 | 1 | 0% | 2,222 | 3,793 | +71% | 0 | 0 | — |
case-10 | pass→pass | 10,082 | 4,401 | -56% | 1 | 1 | 0% | 1,631 | 2,748 | +68% | 0 | 0 | — |
case-11 | pass→pass | 13,177 | 13,675 | +4% | 1 | 1 | 0% | 2,109 | 4,226 | +100% | 0 | 0 | — |
case-12 | pass→pass | 12,716 | 6,724 | -47% | 1 | 1 | 0% | 2,173 | 3,107 | +43% | 0 | 0 | — |
case-13 | fail→pass | 11,282 | 2,658 | -76% | 1 | 1 | 0% | 1,746 | 2,499 | +43% | 0 | 0 | — |
case-14 | pass→pass | 14,719 | 9,392 | -36% | 1 | 1 | 0% | 2,279 | 3,502 | +54% | 0 | 0 | — |
case-15 | fail→pass | 10,894 | 5,778 | -47% | 1 | 1 | 0% | 1,694 | 2,931 | +73% | 0 | 0 | — |
case-16 | pass→pass | 15,154 | 12,347 | -19% | 1 | 1 | 0% | 2,423 | 3,940 | +63% | 0 | 0 | — |
case-17 | fail→pass | 15,273 | 7,273 | -52% | 1 | 1 | 0% | 2,508 | 3,225 | +29% | 0 | 0 | — |
case-18 | fail→pass | 14,640 | 16,331 | +12% | 1 | 1 | 0% | 2,421 | 4,670 | +93% | 0 | 0 | — |
case-19 | fail→pass | 9,318 | 1,737 | -81% | 1 | 1 | 0% | 1,375 | 2,336 | +70% | 0 | 0 | — |
case-20 | fail→pass | 12,068 | 5,271 | -56% | 1 | 1 | 0% | 1,827 | 2,941 | +61% | 0 | 0 | — |
case-21 | pass→pass | 7,487 | 3,598 | -52% | 1 | 1 | 0% | 1,180 | 2,563 | +117% | 0 | 0 | — |
case-22 | pass→pass | 7,333 | 3,312 | -55% | 1 | 1 | 0% | 1,164 | 2,563 | +120% | 0 | 0 | — |
case-23 | pass→pass | 11,860 | 3,912 | -67% | 1 | 1 | 0% | 1,955 | 2,701 | +38% | 0 | 0 | — |
case-24 | fail→pass | 13,570 | 10,403 | -23% | 1 | 1 | 0% | 2,316 | 3,828 | +65% | 0 | 0 | — |
case-25 | pass→pass | 14,236 | 15,088 | +6% | 1 | 1 | 0% | 2,390 | 4,608 | +93% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +40 percentage points is the difference between those two pass rates over the 25 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.