Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Map the end-to-end service delivery system including frontstage actions, backstage processes, and supporting infrastructure.
.claude/skills/owl-listener-service-blueprint/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-03 | ✓→✗ | ▼ Worse | -1% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 110% | 0% |
You are an expert in service design and systems-level experience mapping.
You create service blueprints that reveal how a service is delivered across all channels and actors — giving teams a shared view of the full system, not just the user-facing touchpoints.
A blueprint maps five horizontal swim lanes:
Line of interaction: separates user actions from frontstage Line of visibility: separates frontstage (visible to user) from backstage (invisible) Line of internal interaction: separates backstage from support processes
| | Journey Map | Service Blueprint | |---|---|---| | Focus | User experience | Entire delivery system | | Actors | User | User + employees + systems | | Purpose | Understand emotional journey | Reveal operational gaps and dependencies | | When | Research and ideation | System design and coordination | Use journey maps to understand the experience; use blueprints to design and fix the system delivering it.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | pass→pass | 8,315 | 8,030 | -3% | 1 | 1 | 0% | 1,390 | 2,179 | +57% | 0 | 0 | — |
case-01 | fail→pass | 23,696 | 20,907 | -12% | 1 | 1 | 0% | 4,111 | 4,642 | +13% | 0 | 0 | — |
case-02 | pass→pass | 16,801 | 18,058 | +7% | 1 | 1 | 0% | 3,029 | 4,108 | +36% | 0 | 0 | — |
case-03 | pass→fail | 21,207 | 17,615 | -17% | 1 | 1 | 0% | 3,797 | 3,773 | -1% | 0 | 0 | — |
case-04 | pass→pass | 16,514 | 15,597 | -6% | 1 | 1 | 0% | 2,906 | 3,508 | +21% | 0 | 0 | — |
case-05 | pass→pass | 11,071 | 9,482 | -14% | 1 | 1 | 0% | 1,905 | 2,473 | +30% | 0 | 0 | — |
case-06 | fail→pass | 6,237 | 8,224 | +32% | 1 | 1 | 0% | 1,147 | 2,272 | +98% | 0 | 0 | — |
case-07 | pass→fail | 4,153 | 4,309 | +4% | 1 | 1 | 0% | 786 | 1,649 | +110% | 0 | 0 | — |
case-08 | pass→pass | 12,457 | 15,055 | +21% | 1 | 1 | 0% | 2,386 | 3,386 | +42% | 0 | 0 | — |
case-09 | fail→pass | 13,789 | 11,127 | -19% | 1 | 1 | 0% | 2,593 | 2,571 | -1% | 0 | 0 | — |
case-10 | fail→fail | 9,737 | 9,405 | -3% | 1 | 1 | 0% | 1,637 | 2,225 | +36% | 0 | 0 | — |
case-11 | pass→pass | 12,291 | 10,464 | -15% | 1 | 1 | 0% | 2,175 | 2,651 | +22% | 0 | 0 | — |
case-13 | pass→pass | 9,918 | 10,579 | +7% | 1 | 1 | 0% | 1,720 | 2,606 | +52% | 0 | 0 | — |
case-14 | pass→pass | 9,407 | 6,877 | -27% | 1 | 1 | 0% | 1,388 | 1,903 | +37% | 0 | 0 | — |
case-15 | pass→pass | 11,379 | 10,957 | -4% | 1 | 1 | 0% | 1,829 | 2,552 | +40% | 0 | 0 | — |
case-16 | pass→pass | 7,514 | 4,216 | -44% | 1 | 1 | 0% | 1,210 | 1,524 | +26% | 0 | 0 | — |
case-17 | pass→pass | 9,141 | 5,681 | -38% | 1 | 1 | 0% | 1,309 | 1,659 | +27% | 0 | 0 | — |
case-18 | pass→pass | 10,585 | 8,481 | -20% | 1 | 1 | 0% | 1,673 | 1,957 | +17% | 0 | 0 | — |
case-19 | pass→pass | 8,714 | 7,331 | -16% | 1 | 1 | 0% | 1,500 | 1,989 | +33% | 0 | 0 | — |
case-20 | pass→pass | 14,114 | 12,281 | -13% | 1 | 1 | 0% | 2,027 | 2,631 | +30% | 0 | 0 | — |
case-21 | pass→pass | 7,746 | 4,482 | -42% | 1 | 1 | 0% | 1,494 | 1,633 | +9% | 0 | 0 | — |
case-22 | pass→pass | 6,743 | 6,517 | -3% | 1 | 1 | 0% | 1,063 | 1,892 | +78% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.