Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Discovers and invokes agent skills. Use when starting a session or when you need to discover which skill applies to the current task. This is the meta-skill that governs how all other skills are discovered and invoked.
.claude/skills/addyosmani-using-agent-skills/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 124% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 234% | 0% |
Agent Skills is a collection of engineering workflow skills organized by development phase. Each skill encodes a specific process that senior engineers follow. This meta-skill helps you discover and apply the right skill for your current task.
When a task arrives, identify the development phase and apply the corresponding skill:
Task arrives
│
├── Don't know what you want yet? ──────→ interview-me
├── Have a rough concept, need variants? → idea-refine
├── New project/feature/change? ──→ spec-driven-development
├── No quality bar written down? ──→ constraint-driven-development
├── Have a spec, need tasks? ──────→ planning-and-task-breakdown
├── Implementing code? ────────────→ incremental-implementation
│ ├── UI work? ─────────────────→ frontend-ui-engineering
│ ├── API work? ────────────────→ api-and-interface-design
│ ├── Need better context? ─────→ context-engineering
│ ├── Need doc-verified code? ───→ source-driven-development
│ └── Stakes high / unfamiliar code? ──→ doubt-driven-development
├── Writing/running tests? ────────→ test-driven-development
│ └── Browser-based? ───────────→ browser-testing-with-devtools
├── Something broke? ──────────────→ debugging-and-error-recovery
├── Reviewing code? ───────────────→ code-review-and-quality
│ ├── Too complex? ─────────────→ code-simplification
│ ├── Security concerns? ───────→ security-and-hardening
│ └── Performance concerns? ────→ performance-optimization
├── Committing/branching? ─────────→ git-workflow-and-versioning
├── CI/CD pipeline work? ──────────→ ci-cd-and-automation
├── Deprecating/migrating? ────────→ deprecation-and-migration
├── Writing docs/ADRs? ───────────→ documentation-and-adrs
├── Adding logs/metrics/alerts? ───→ observability-and-instrumentation
└── Deploying/launching? ─────────→ shipping-and-launchThese behaviors apply at all times, across all skills. They are non-negotiable.
Before implementing anything non-trivial, explicitly state your assumptions:
ASSUMPTIONS I'M MAKING:
1. [assumption about requirements]
2. [assumption about architecture]
3. [assumption about scope]
→ Correct me now or I'll proceed with these.Don't silently fill in ambiguous requirements. The most common failure mode is making wrong assumptions and running with them unchecked. Surface uncertainty early — it's cheaper than rework.
When you encounter inconsistencies, conflicting requirements, or unclear specifications:
Bad: Silently picking one interpretation and hoping it's right. Good: "I see X in the spec but Y in the existing code. Which takes precedence?"
You are not a yes-machine. When an approach has clear problems:
Sycophancy is a failure mode. "Of course!" followed by implementing a bad idea helps no one. Honest technical disagreement is more valuable than false agreement.
Your natural tendency is to overcomplicate. Actively resist it.
Before finishing any implementation, ask:
If you build 1000 lines and 100 would suffice, you have failed. Prefer the boring, obvious solution. Cleverness is expensive.
Touch only what you're asked to touch.
Do NOT:
Your job is surgical precision, not unsolicited renovation.
Every skill includes a verification step. A task is not complete until verification passes. "Seems right" is never sufficient — there must be evidence (passing tests, build output, runtime data).
Per-skill verification is the local check. The project-wide bar that applies to every change, regardless of which skill is active, is the Definition of Done: tests pass, no regressions, behavior verified at runtime, docs updated. See ../../references/definition-of-done.md. It complements each task's acceptance criteria rather than replacing them.
These are the subtle errors that look like productivity but create problems:
idea-refine → spec-driven-development → planning-and-task-breakdown → incremental-implementation → test-driven-development → code-review-and-quality → code-simplification → shipping-and-launch in sequence.spec-driven-development.For a complete feature, the typical skill sequence is:
1. interview-me → Extract what the user actually wants
2. idea-refine → Refine vague ideas
3. spec-driven-development → Define what we're building
4. planning-and-task-breakdown → Break into verifiable chunks
5. context-engineering → Load the right context
6. source-driven-development → Verify against official docs
7. incremental-implementation → Build slice by slice
8. observability-and-instrumentation → Instrument as you build (runs parallel with 7-9, not after)
9. doubt-driven-development → Cross-examine non-trivial decisions in-flight
10. test-driven-development → Prove each slice works
11. code-review-and-quality → Review before merge
12. code-simplification → Reduce unnecessary complexity while preserving behavior
13. git-workflow-and-versioning → Clean commit history
14. documentation-and-adrs → Document decisions
15. deprecation-and-migration → Retire old systems and move users safely when needed
16. shipping-and-launch → Deploy safelyNot every task needs every skill. A bug fix might only need: debugging-and-error-recovery → test-driven-development → code-review-and-quality.
| Phase | Skill | One-Line Summary | |-------|-------|-----------------| | Define | interview-me | Surface what the user actually wants before any plan, spec, or code exists | | Define | idea-refine | Refine ideas through structured divergent and convergent thinking | | Define | spec-driven-development | Requirements and acceptance criteria before code | | Plan | planning-and-task-breakdown | Decompose into small, verifiable tasks | | Build | incremental-implementation | Thin vertical slices, test each before expanding | | Build | source-driven-development | Verify against official docs before implementing | | Build | doubt-driven-development | Adversarial fresh-context review of every non-trivial decision | | Build | context-engineering | Right context at the right time | | Build | frontend-ui-engineering | Production-quality UI with accessibility | | Build | api-and-interface-design | Stable interfaces with clear contracts | | Verify | test-driven-development | Failing test first, then make it pass | | Verify | browser-testing-with-devtools | Chrome DevTools MCP for runtime verification | | Verify | debugging-and-error-recovery | Reproduce → localize → fix → guard | | Review | code-review-and-quality | Five-axis review with quality gates | | Review | code-simplification | Preserve behavior while reducing unnecessary complexity | | Review | security-and-hardening | OWASP prevention, input validation, least privilege | | Review | performance-optimization | Measure first, optimize only what matters | | Ship | git-workflow-and-versioning | Atomic commits, clean history | | Ship | ci-cd-and-automation | Automated quality gates on every change | | Ship | deprecation-and-migration | Remove old systems and migrate users safely | | Ship | documentation-and-adrs | Document the why, not just the what | | Ship | observability-and-instrumentation | Structured logs, RED metrics, traces, symptom-based alerts | | Ship | shipping-and-launch | Pre-launch checklist, monitoring, rollback plan |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→pass | 21,820 | 20,048 | -8% | 1 | 1 | 0% | 2,701 | 4,544 | +68% | 0 | 0 | — |
case-01 | fail→fail | 28,300 | 28,718 | +1% | 1 | 1 | 0% | 3,292 | 5,722 | +74% | 0 | 0 | — |
case-02 | fail→fail | 23,432 | 21,878 | -7% | 1 | 1 | 0% | 2,887 | 4,568 | +58% | 0 | 0 | — |
case-03 | fail→fail | 48,686 | 31,282 | -36% | 1 | 1 | 0% | 7,873 | 6,028 | -23% | 0 | 0 | — |
case-04 | fail→fail | 18,247 | 14,959 | -18% | 1 | 1 | 0% | 2,066 | 4,045 | +96% | 0 | 0 | — |
case-06 | pass→pass | 22,248 | 16,927 | -24% | 1 | 1 | 0% | 2,353 | 3,840 | +63% | 0 | 0 | — |
case-07 | fail→fail | 16,897 | 10,386 | -39% | 1 | 1 | 0% | 1,447 | 3,013 | +108% | 0 | 0 | — |
case-08 | fail→pass | 17,929 | 10,350 | -42% | 1 | 1 | 0% | 1,951 | 3,195 | +64% | 0 | 0 | — |
case-09 | fail→pass | 14,154 | 10,296 | -27% | 1 | 1 | 0% | 1,327 | 2,973 | +124% | 0 | 0 | — |
case-10 | pass→pass | 12,689 | 8,879 | -30% | 1 | 1 | 0% | 1,120 | 2,916 | +160% | 0 | 0 | — |
case-11 | fail→pass | 16,306 | 8,919 | -45% | 1 | 1 | 0% | 1,593 | 2,924 | +84% | 0 | 0 | — |
case-12 | pass→pass | 13,474 | 8,144 | -40% | 1 | 1 | 0% | 1,376 | 2,908 | +111% | 0 | 0 | — |
case-13 | pass→pass | 19,889 | 3,441 | -83% | 1 | 1 | 0% | 2,293 | 2,708 | +18% | 0 | 0 | — |
case-14 | pass→pass | 13,265 | 10,037 | -24% | 1 | 1 | 0% | 1,295 | 2,890 | +123% | 0 | 0 | — |
case-15 | pass→pass | 5,841 | 3,662 | -37% | 1 | 1 | 0% | 714 | 2,821 | +295% | 0 | 0 | — |
case-16 | pass→pass | 13,732 | 11,782 | -14% | 1 | 1 | 0% | 1,849 | 3,222 | +74% | 0 | 0 | — |
case-17 | pass→pass | 13,608 | 9,536 | -30% | 1 | 1 | 0% | 1,161 | 2,808 | +142% | 0 | 0 | — |
case-18 | fail→pass | 11,171 | 8,803 | -21% | 1 | 1 | 0% | 877 | 2,930 | +234% | 0 | 0 | — |
case-19 | fail→pass | 14,778 | 4,784 | -68% | 1 | 1 | 0% | 1,272 | 2,845 | +124% | 0 | 0 | — |
case-20 | pass→pass | 14,743 | 25,331 | +72% | 1 | 1 | 0% | 2,508 | 4,459 | +78% | 0 | 0 | — |
case-21 | pass→fail | 14,787 | 10,050 | -32% | 1 | 1 | 0% | 1,367 | 2,955 | +116% | 0 | 0 | — |
case-22 | pass→pass | 15,834 | 11,742 | -26% | 1 | 1 | 0% | 1,515 | 4,000 | +164% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/11/2026 | +36% |
| gemini-3.6-flash | verified | 8/8/2026 | +27% |
Other measured skills in the registry, with their headline benchmark lift.