Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute Flexport primary workflow: shipment booking and purchase order management. Use when creating bookings, managing purchase orders, tracking freight, or building the core shipment lifecycle integration. Trigger: "flexport booking", "flexport purchase order", "create flexport shipment".
.claude/skills/jeremylongshore-flexport-core-workflow-a/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 100% | 0% |
Separate discovery from commitment. MCP rate tools can search, evaluate, request, and book; a booking must never be inferred from a search result or performed before a named approver accepts the final normalized terms.
https://mcp.flexport.com/mcpResolve addresses, company entities, ports, addresses, and commodity inputs with read-only network tools. Preserve opaque Flexport identifiers.
Call rates_search_instant_price with approved inputs and store the search/result identifiers plus a redacted input digest.
Use rates_evaluate_total_price_from_instant_price_search for the selected result. Present currency, charge breakdown, timing, assumptions, and expiry exactly as returned.
If instant booking is unavailable, use rates_request_rate and reconcile the quote request. Do not substitute rates_book_without_rate unless that supplier flow was explicitly approved.
Bind the approver, exact evaluated result, cargo digest, and approval expiry. Re-evaluate when any material input or provider term changes.
Persist an operation key before rates_instant_book; store the returned booking reference and reconcile ambiguous outcomes before any retry.
REST calls authenticate with a cached OAuth 2.0 client-credentials Bearer token using audience https://api.flexport.com, or an explicitly accepted broad API key. Use distinct credentials per workload and never log credentials or tokens. MCP calls use the authenticated connection to https://mcp.flexport.com/mcp and remain subject to each tool's documented account permissions.
Use Read and Grep for discovery and evidence. Use Write or Edit only for the approved artifact, code, configuration, test, or receipt described by this workflow; do not make an unapproved Flexport-side change.
Return a machine-reviewable receipt in this shape; adapt the operation values, but never place credentials or provider payloads in it:
yamlsurface: rest-v3 operation: shipment-read decision: approved outcome: verified evidence: release_sha: recorded-out-of-band provider_reference: redacted rollback_owner: logistics-platform
An operator searches an ocean lane, evaluates one returned rate, and receives a review card. Only after the logistics owner approves that exact result does the service make one booking call and record the returned reference.
| Failure | Response | | --- | --- | | No instant rate | Create a quote request or stop; never fabricate a price. | | Permission denied | Route to an authorized role instead of escalating credentials. | | Terms changed after approval | Invalidate approval and present a fresh evaluation. | | Booking outcome ambiguous | Reconcile by known references before considering another mutation. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 12,663 | 13,470 | +6% | 1 | 1 | 0% | 2,819 | 3,471 | +23% | 0 | 0 | — |
case-02 | fail→pass | 20,709 | 15,873 | -23% | 1 | 1 | 0% | 3,326 | 4,039 | +21% | 0 | 0 | — |
case-03 | fail→pass | 21,810 | 8,493 | -61% | 1 | 1 | 0% | 3,393 | 3,254 | -4% | 0 | 0 | — |
case-04 | fail→fail | 13,577 | 4,118 | -70% | 1 | 1 | 0% | 1,530 | 2,193 | +43% | 0 | 0 | — |
case-05 | fail→pass | 17,020 | 8,491 | -50% | 1 | 1 | 0% | 2,988 | 3,375 | +13% | 0 | 0 | — |
case-06 | fail→fail | 14,659 | 6,176 | -58% | 1 | 1 | 0% | 1,989 | 2,716 | +37% | 0 | 0 | — |
case-07 | fail→pass | 12,221 | 13,124 | +7% | 1 | 1 | 0% | 1,542 | 3,089 | +100% | 0 | 0 | — |
case-08 | pass→pass | 22,444 | 12,240 | -45% | 1 | 1 | 0% | 3,276 | 3,021 | -8% | 0 | 0 | — |
case-09 | fail→pass | 12,289 | 11,875 | -3% | 1 | 1 | 0% | 1,325 | 2,435 | +84% | 0 | 0 | — |
case-10 | pass→pass | 14,328 | 11,287 | -21% | 1 | 1 | 0% | 1,731 | 2,563 | +48% | 0 | 0 | — |
case-11 | fail→pass | 17,173 | 14,636 | -15% | 1 | 1 | 0% | 2,259 | 3,234 | +43% | 0 | 0 | — |
case-12 | fail→pass | 12,364 | 10,284 | -17% | 1 | 1 | 0% | 2,259 | 2,458 | +9% | 0 | 0 | — |
case-13 | pass→pass | 16,146 | 10,285 | -36% | 1 | 1 | 0% | 2,182 | 2,614 | +20% | 0 | 0 | — |
case-14 | fail→pass | 12,210 | 9,701 | -21% | 1 | 1 | 0% | 2,394 | 2,433 | +2% | 0 | 0 | — |
case-15 | pass→pass | 8,595 | 4,267 | -50% | 1 | 1 | 0% | 1,646 | 2,270 | +38% | 0 | 0 | — |
case-16 | pass→pass | 8,666 | 10,430 | +20% | 1 | 1 | 0% | 1,748 | 2,626 | +50% | 0 | 0 | — |
case-17 | pass→pass | 14,580 | 9,352 | -36% | 1 | 1 | 0% | 1,927 | 2,354 | +22% | 0 | 0 | — |
case-18 | pass→pass | 7,408 | 8,125 | +10% | 1 | 1 | 0% | 1,430 | 2,065 | +44% | 0 | 0 | — |
case-19 | pass→pass | 11,736 | 8,369 | -29% | 1 | 1 | 0% | 2,653 | 3,343 | +26% | 0 | 0 | — |
case-20 | fail→fail | 20,959 | 16,041 | -23% | 1 | 1 | 0% | 2,794 | 4,226 | +51% | 0 | 0 | — |
case-21 | fail→fail | 10,204 | 12,170 | +19% | 1 | 1 | 0% | 2,120 | 2,872 | +35% | 0 | 0 | — |
case-22 | fail→fail | 17,454 | 13,636 | -22% | 1 | 1 | 0% | 2,750 | 3,350 | +22% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.