Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute Flexport production deployment checklist for logistics integrations. Use when deploying shipment tracking, booking automation, or supply chain integrations to production with proper monitoring and rollback. Trigger: "flexport production", "deploy flexport", "flexport go-live checklist".
.claude/skills/jeremylongshore-flexport-prod-checklist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 1% | 0% |
A checklist item passes only with a receipt tied to the immutable release and configuration. Read access, bookings, trade records, and webhook processing have different risk boundaries and must be approved separately.
Verify account, credential alias/type, endpoint resources, token caching, secret injection, revocation, and MCP role permissions where used.
Confirm v3/account version behavior, tolerant additive-field parsing, documented endpoints/tools, error classification, and pagination.
Approve field minimization, destinations, telemetry redaction, retention/deletion, and tenant boundaries.
Require durable operation keys, business approval for bookings/writes, ambiguous-outcome reconciliation, and one authoritative writer.
Prove raw-body X-Hub-Signature-256 validation, durable enqueue before 200, deduplication, event allowlist, and gap reconciliation.
Run a read-only canary, expand by cohort, test rollback, and attach all receipts to the release decision.
REST calls authenticate with a cached OAuth 2.0 client-credentials Bearer token using audience https://api.flexport.com, or an explicitly accepted broad API key. Use distinct credentials per workload and never log credentials or tokens. MCP calls use the authenticated connection to https://mcp.flexport.com/mcp and remain subject to each tool's documented account permissions.
Use Read and Grep for discovery and evidence. Use Write or Edit only for the approved artifact, code, configuration, test, or receipt described by this workflow; do not make an unapproved Flexport-side change.
Return a machine-reviewable receipt in this shape; adapt the operation values, but never place credentials or provider payloads in it:
yamlsurface: rest-v3 operation: shipment-read decision: approved outcome: verified evidence: release_sha: recorded-out-of-band provider_reference: redacted rollback_owner: logistics-platform
The launch record enables shipment reads and signed webhook ingestion but leaves booking disabled because its human approval and ambiguous-outcome reconciliation evidence is incomplete.
| Failure | Response | | --- | --- | | Evidence belongs to another SHA | Re-run the gate on the release candidate. | | Owner or rollback absent | Do not launch. | | Broad API key lacks exception | Replace it or obtain explicit risk acceptance. | | Mutation reconciliation unproven | Keep that capability disabled. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 24,044 | 32,179 | +34% | 1 | 1 | 0% | 3,347 | 5,787 | +73% | 0 | 0 | — |
case-02 | fail→pass | 23,291 | 9,397 | -60% | 1 | 1 | 0% | 3,239 | 2,511 | -22% | 0 | 0 | — |
case-03 | fail→pass | 18,298 | 12,636 | -31% | 1 | 1 | 0% | 2,531 | 2,324 | -8% | 0 | 0 | — |
case-04 | fail→fail | 21,245 | 25,439 | +20% | 1 | 1 | 0% | 3,926 | 4,655 | +19% | 0 | 0 | — |
case-05 | pass→pass | 17,649 | 24,902 | +41% | 1 | 1 | 0% | 2,827 | 4,311 | +52% | 0 | 0 | — |
case-06 | pass→pass | 22,319 | 23,752 | +6% | 1 | 1 | 0% | 3,305 | 4,426 | +34% | 0 | 0 | — |
case-07 | fail→pass | 13,055 | 8,090 | -38% | 1 | 1 | 0% | 1,582 | 1,329 | -16% | 0 | 0 | — |
case-08 | fail→pass | 14,497 | 5,105 | -65% | 1 | 1 | 0% | 1,668 | 1,687 | +1% | 0 | 0 | — |
case-09 | fail→pass | 6,409 | 11,058 | +73% | 1 | 1 | 0% | 1,110 | 1,808 | +63% | 0 | 0 | — |
case-10 | pass→pass | 14,600 | 6,259 | -57% | 1 | 1 | 0% | 1,617 | 1,776 | +10% | 0 | 0 | — |
case-11 | fail→pass | 14,622 | 12,050 | -18% | 1 | 1 | 0% | 2,174 | 2,016 | -7% | 0 | 0 | — |
case-12 | pass→pass | 18,989 | 17,535 | -8% | 1 | 1 | 0% | 2,261 | 2,757 | +22% | 0 | 0 | — |
case-13 | pass→pass | 16,713 | 6,965 | -58% | 1 | 1 | 0% | 1,883 | 1,843 | -2% | 0 | 0 | — |
case-14 | pass→pass | 16,253 | 10,862 | -33% | 1 | 1 | 0% | 1,873 | 1,788 | -5% | 0 | 0 | — |
case-15 | pass→pass | 6,113 | 1,802 | -71% | 1 | 1 | 0% | 977 | 974 | -0% | 0 | 0 | — |
case-16 | fail→pass | 13,570 | 4,957 | -63% | 1 | 1 | 0% | 2,234 | 1,488 | -33% | 0 | 0 | — |
case-17 | fail→pass | 13,246 | 7,247 | -45% | 1 | 1 | 0% | 2,323 | 1,061 | -54% | 0 | 0 | — |
case-18 | fail→pass | 11,604 | 1,792 | -85% | 1 | 1 | 0% | 2,005 | 995 | -50% | 0 | 0 | — |
case-19 | fail→pass | 12,417 | 1,365 | -89% | 1 | 1 | 0% | 1,485 | 954 | -36% | 0 | 0 | — |
case-20 | pass→pass | 7,341 | 1,749 | -76% | 1 | 1 | 0% | 452 | 1,080 | +139% | 0 | 0 | — |
case-21 | pass→pass | 16,800 | 8,082 | -52% | 1 | 1 | 0% | 2,300 | 2,204 | -4% | 0 | 0 | — |
case-22 | pass→pass | 12,503 | 17,076 | +37% | 1 | 1 | 0% | 2,494 | 3,327 | +33% | 0 | 0 | — |
case-23 | pass→pass | 12,501 | 13,474 | +8% | 1 | 1 | 0% | 2,696 | 2,636 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +48 percentage points is the difference between those two pass rates over the 23 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.