Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute BambooHR production deployment checklist and rollback procedures. Use when deploying BambooHR integrations to production, preparing for launch, or implementing go-live procedures with BambooHR API. Trigger with phrases like "bamboohr production", "deploy bamboohr", "bamboohr go-live", "bamboohr launch checklist", "bamboohr prod ready".
.claude/skills/jeremylongshore-bamboohr-prod-checklist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 10% | 0% |
Issue a reproducible go/no-go verdict. A checklist item passes only with current evidence tied to the release artifact and environment; "configured" or "tested before" is not a receipt.
The reviewed BambooHR contract includes tenant-local hosts, OAuth/API-key auth, caller-owned refreshed-token persistence, typed request IDs/errors, finite SDK retries, dataset v2, and one-time webhook keys. Legacy dataset/report deprecations and Python package-registry uncertainty must be explicit.
Require auth mode, identity owner, tenant binding, field/operation permissions, credential age, rotation/revocation runbook, and OAuth state/token-persistence tests. Verify references, not secret values.
dependency lock, vulnerability/secret scans, and required CI are green.
token/key storage, destination namespace, queues, and checkpoints.
retention, encryption, deletion, backup, log, and support-evidence controls.
reconciliation, retries, ambiguous writes, schema drift, webhook HMAC/replay, and redaction.
circuit breaker, dead letter, alerts, dashboards, on-call ownership, and body-free health checks.
public registry release until reverified. For PHP, record the locked Packagist version.
scheduler/queue ownership and checkpoint compatibility for both versions.
observe through the agreed window.
reason. Any unresolved critical gate yields NO-GO.
Use Read, Glob, and Grep to collect exact release evidence. Use Write/Edit only for the approved readiness receipt and missing tests/configuration. This skill does not deploy, rotate secrets, call tenants, or alter production state.
Launch approval must identify engineering, HR/data owner, security/privacy, and operations owners as applicable. A waiver needs owner, rationale, compensating control, and expiry.
Return release and artifact identity, gate table with evidence links, canary and reconciliation result, waivers, rollback readiness, approvers, and final GO/NO-GO.
Read official evidence during the release review.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 24,279 | 13,809 | -43% | 1 | 1 | 0% | 2,711 | 4,102 | +51% | 0 | 0 | — |
case-02 | fail→fail | 18,367 | 17,322 | -6% | 1 | 1 | 0% | 3,972 | 5,331 | +34% | 0 | 0 | — |
case-03 | fail→fail | 25,707 | 27,695 | +8% | 1 | 1 | 0% | 4,255 | 6,995 | +64% | 0 | 0 | — |
case-04 | fail→pass | 11,312 | 7,974 | -30% | 1 | 1 | 0% | 2,097 | 3,307 | +58% | 0 | 0 | — |
case-05 | fail→pass | 13,644 | 12,330 | -10% | 1 | 1 | 0% | 2,320 | 3,994 | +72% | 0 | 0 | — |
case-19 | fail→pass | 13,899 | 10,495 | -24% | 1 | 1 | 0% | 2,483 | 3,804 | +53% | 0 | 0 | — |
case-06 | pass→pass | 10,522 | 5,388 | -49% | 1 | 1 | 0% | 1,706 | 2,462 | +44% | 0 | 0 | — |
case-07 | pass→pass | 9,070 | 7,257 | -20% | 1 | 1 | 0% | 1,573 | 2,960 | +88% | 0 | 0 | — |
case-08 | pass→pass | 11,313 | 8,477 | -25% | 1 | 1 | 0% | 2,016 | 3,243 | +61% | 0 | 0 | — |
case-09 | fail→pass | 14,183 | 4,671 | -67% | 1 | 1 | 0% | 2,250 | 2,481 | +10% | 0 | 0 | — |
case-10 | fail→fail | 17,006 | 10,444 | -39% | 1 | 1 | 0% | 2,687 | 3,691 | +37% | 0 | 0 | — |
case-11 | fail→pass | 14,533 | 3,865 | -73% | 1 | 1 | 0% | 2,246 | 2,304 | +3% | 0 | 0 | — |
case-12 | fail→pass | 8,018 | 6,753 | -16% | 1 | 1 | 0% | 1,320 | 2,838 | +115% | 0 | 0 | — |
case-13 | fail→pass | 11,781 | 5,963 | -49% | 1 | 1 | 0% | 2,256 | 2,781 | +23% | 0 | 0 | — |
case-14 | fail→pass | 13,391 | 10,526 | -21% | 1 | 1 | 0% | 2,069 | 3,373 | +63% | 0 | 0 | — |
case-15 | fail→pass | 16,549 | 15,578 | -6% | 1 | 1 | 0% | 2,523 | 4,218 | +67% | 0 | 0 | — |
case-16 | pass→pass | 11,634 | 19,923 | +71% | 1 | 1 | 0% | 1,762 | 3,414 | +94% | 0 | 0 | — |
case-17 | pass→pass | 17,613 | 14,101 | -20% | 1 | 1 | 0% | 2,914 | 4,002 | +37% | 0 | 0 | — |
case-18 | fail→pass | 29,357 | 3,039 | -90% | 1 | 1 | 0% | 1,825 | 2,191 | +20% | 0 | 0 | — |
case-20 | fail→fail | 18,626 | 21,184 | +14% | 1 | 1 | 0% | 3,508 | 5,671 | +62% | 0 | 0 | — |
case-21 | pass→pass | 20,194 | 17,067 | -15% | 1 | 1 | 0% | 4,388 | 5,629 | +28% | 0 | 0 | — |
case-22 | pass→pass | 17,773 | 14,441 | -19% | 1 | 1 | 0% | 3,991 | 4,912 | +23% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.