Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Collect Flexport API debug evidence for support tickets and troubleshooting. Use when encountering persistent API issues, preparing support tickets, or collecting diagnostic information for Flexport logistics problems. Trigger: "flexport debug", "flexport support bundle", "flexport diagnostic".
.claude/skills/jeremylongshore-flexport-debug-bundle/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 22 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 48% | 0% |
Collect the smallest evidence that can distinguish contract, permission, version, transport, and provider failures. Never archive an entire environment, log directory, request body, webhook payload, or credential file.
Declare exact metadata fields before collection: timestamps, release SHA, surface, operation, version header, status/code/message, correlation ID, and payload digest.
Record variable names, endpoint hosts, feature flags, and credential aliases—not values, tokens, secret suffixes, or .env content.
List bounded attempt timestamps and outcomes. Exclude raw request/response bodies, document data, route details, names, emails, and addresses.
Include a minimal sanitized fixture or field-name diff only when necessary to reproduce the failure.
Run automated secret/sensitive-data scanning, then require a human owner to approve the exact manifest and recipient.
Use the approved support channel, record a checksum and expiry, then delete according to incident evidence policy.
REST calls authenticate with a cached OAuth 2.0 client-credentials Bearer token using audience https://api.flexport.com, or an explicitly accepted broad API key. Use distinct credentials per workload and never log credentials or tokens. MCP calls use the authenticated connection to https://mcp.flexport.com/mcp and remain subject to each tool's documented account permissions.
Use Read and Grep for discovery and evidence. Use Write or Edit only for the approved artifact, code, configuration, test, or receipt described by this workflow; do not make an unapproved Flexport-side change.
Return a machine-reviewable receipt in this shape; adapt the operation values, but never place credentials or provider payloads in it:
yamlsurface: rest-v3 operation: shipment-read decision: approved outcome: verified evidence: release_sha: recorded-out-of-band provider_reference: redacted rollback_owner: logistics-platform
A webhook escalation contains the receiver release SHA, event-type string, signature-header presence, body digest, status outcome, and timestamps. It contains neither the webhook body nor the secret.
| Failure | Response | | --- | --- | | Secret scanner fires | Stop transfer, remove the material, rotate if exposure occurred, and rescan. | | Payload needed for reproduction | Create a synthetic minimal fixture rather than shipping production content. | | Recipient or purpose unclear | Do not build or send the manifest. | | Archive tool proposes broad paths | Reject it and use the explicit allowlist only. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 23,697 | 18,639 | -21% | 1 | 1 | 0% | 3,644 | 3,653 | +0% | 0 | 0 | — |
case-02 | fail→fail | 45,971 | 19,364 | -58% | 1 | 1 | 0% | 3,659 | 3,997 | +9% | 0 | 0 | — |
case-03 | fail→fail | 25,721 | 22,245 | -14% | 1 | 1 | 0% | 4,152 | 4,492 | +8% | 0 | 0 | — |
case-04 | pass→pass | 18,682 | 17,349 | -7% | 1 | 1 | 0% | 3,321 | 3,403 | +2% | 0 | 0 | — |
case-05 | pass→pass | 22,074 | 18,204 | -18% | 1 | 1 | 0% | 3,414 | 3,609 | +6% | 0 | 0 | — |
case-06 | pass→pass | 14,299 | 13,951 | -2% | 1 | 1 | 0% | 1,604 | 2,583 | +61% | 0 | 0 | — |
case-07 | fail→pass | 18,589 | 11,663 | -37% | 1 | 1 | 0% | 1,190 | 2,305 | +94% | 0 | 0 | — |
case-08 | fail→fail | 16,482 | 10,963 | -33% | 1 | 1 | 0% | 1,919 | 2,078 | +8% | 0 | 0 | — |
case-09 | pass→pass | 30,157 | 8,933 | -70% | 1 | 1 | 0% | 2,438 | 2,582 | +6% | 0 | 0 | — |
case-10 | pass→pass | 11,011 | 10,223 | -7% | 1 | 1 | 0% | 1,185 | 2,028 | +71% | 0 | 0 | — |
case-11 | fail→pass | 10,380 | 11,211 | +8% | 1 | 1 | 0% | 1,035 | 2,180 | +111% | 0 | 0 | — |
case-12 | fail→pass | 12,872 | 11,321 | -12% | 1 | 1 | 0% | 1,384 | 2,097 | +52% | 0 | 0 | — |
case-13 | fail→pass | 32,057 | 11,232 | -65% | 1 | 1 | 0% | 1,530 | 2,258 | +48% | 0 | 0 | — |
case-14 | pass→pass | 11,308 | 10,851 | -4% | 1 | 1 | 0% | 1,092 | 2,065 | +89% | 0 | 0 | — |
case-15 | fail→fail | 15,065 | 7,458 | -50% | 1 | 1 | 0% | 1,887 | 2,406 | +28% | 0 | 0 | — |
case-16 | fail→pass | 19,853 | 11,847 | -40% | 1 | 1 | 0% | 2,477 | 2,190 | -12% | 0 | 0 | — |
case-17 | pass→pass | 5,436 | 8,037 | +48% | 1 | 1 | 0% | 1,019 | 1,583 | +55% | 0 | 0 | — |
case-18 | fail→fail | 11,528 | 12,090 | +5% | 1 | 1 | 0% | 2,137 | 2,333 | +9% | 0 | 0 | — |
case-19 | fail→pass | 18,592 | 15,272 | -18% | 1 | 1 | 0% | 2,822 | 2,978 | +6% | 0 | 0 | — |
case-20 | fail→pass | 15,146 | 17,542 | +16% | 1 | 1 | 0% | 1,992 | 4,818 | +142% | 0 | 0 | — |
case-21 | fail→pass | 7,316 | 4,129 | -44% | 1 | 1 | 0% | 1,524 | 1,887 | +24% | 0 | 0 | — |
case-22 | fail→pass | 11,044 | 5,642 | -49% | 1 | 1 | 0% | 1,408 | 2,220 | +58% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 21 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.