Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Implement data handling for Flexport supply chain data including PII redaction, shipment data retention, GDPR compliance, and secure document management. Trigger: "flexport data handling", "flexport PII", "flexport GDPR", "flexport data retention".
.claude/skills/jeremylongshore-flexport-data-handling/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 46% | 0% |
Flexport payloads can contain commercial, personal, customs, financial, and document content. This skill does not invent universal GDPR, CCPA, ITAR, or retention rules; it makes the accountable owner and policy explicit.
Map only fields needed for the approved outcome, including identifiers embedded in routes, parties, documents, invoices, customs, tags, or MCP tracking output.
Apply the organization's authoritative classification and legal guidance. Do not infer a regime from an HS code, endpoint, country, or shipment mode.
Prefer opaque resource IDs and derived operational state over raw documents, addresses, names, emails, entry numbers, or charge details.
Use approved encrypted channels and stores, workload-scoped credentials, tenant boundaries, and access logging.
Redact logs, metrics, prompts, tickets, debug bundles, and analytics. Never include document bodies, tokens, or full route/party payloads by default.
Attach retention/deletion rules to the approved data class, test deletion and backup behavior, and record exceptions with owner and expiry.
REST calls authenticate with a cached OAuth 2.0 client-credentials Bearer token using audience https://api.flexport.com, or an explicitly accepted broad API key. Use distinct credentials per workload and never log credentials or tokens. MCP calls use the authenticated connection to https://mcp.flexport.com/mcp and remain subject to each tool's documented account permissions.
Use Read and Grep for discovery and evidence. Use Write or Edit only for the approved artifact, code, configuration, test, or receipt described by this workflow; do not make an unapproved Flexport-side change.
Return a machine-reviewable receipt in this shape; adapt the operation values, but never place credentials or provider payloads in it:
yamlsurface: rest-v3 operation: shipment-read decision: approved outcome: verified evidence: release_sha: recorded-out-of-band provider_reference: redacted rollback_owner: logistics-platform
A delay alert stores a shipment's opaque Flexport ID, milestone category, and alert state. It does not copy route addresses, customs entries, tags, or document content into observability systems.
| Failure | Response | | --- | --- | | Purpose cannot justify a field | Remove it before ingestion. | | Destination lacks approved controls | Block export and route to the data owner. | | Legal rule uncertain | Ask counsel or policy owner; do not encode a guessed retention period. | | Sensitive data reaches logs | Contain access, purge where possible, and repair redaction. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 26,010 | 19,749 | -24% | 1 | 1 | 0% | 5,607 | 4,743 | -15% | 0 | 0 | — |
case-02 | fail→pass | 17,009 | 12,311 | -28% | 1 | 1 | 0% | 3,512 | 4,110 | +17% | 0 | 0 | — |
case-03 | pass→pass | 12,007 | 3,509 | -71% | 1 | 1 | 0% | 2,133 | 2,015 | -6% | 0 | 0 | — |
case-04 | pass→pass | 20,850 | 4,867 | -77% | 1 | 1 | 0% | 2,511 | 2,297 | -9% | 0 | 0 | — |
case-05 | pass→pass | 14,019 | 9,535 | -32% | 1 | 1 | 0% | 2,324 | 2,170 | -7% | 0 | 0 | — |
case-06 | fail→fail | 17,497 | 14,160 | -19% | 1 | 1 | 0% | 2,624 | 3,407 | +30% | 0 | 0 | — |
case-07 | fail→fail | 21,098 | 15,471 | -27% | 1 | 1 | 0% | 2,992 | 4,362 | +46% | 0 | 0 | — |
case-08 | fail→pass | 16,392 | 9,631 | -41% | 1 | 1 | 0% | 2,673 | 2,209 | -17% | 0 | 0 | — |
case-09 | pass→fail | 20,812 | 15,987 | -23% | 1 | 1 | 0% | 2,830 | 3,297 | +17% | 0 | 0 | — |
case-10 | pass→pass | 20,713 | 9,834 | -53% | 1 | 1 | 0% | 2,617 | 3,369 | +29% | 0 | 0 | — |
case-11 | fail→pass | 23,755 | 21,843 | -8% | 1 | 1 | 0% | 2,980 | 4,621 | +55% | 0 | 0 | — |
case-12 | pass→pass | 17,105 | 7,474 | -56% | 1 | 1 | 0% | 2,222 | 2,869 | +29% | 0 | 0 | — |
case-13 | pass→pass | 13,030 | 7,885 | -39% | 1 | 1 | 0% | 1,967 | 2,710 | +38% | 0 | 0 | — |
case-14 | fail→pass | 15,265 | 16,599 | +9% | 1 | 1 | 0% | 2,423 | 3,528 | +46% | 0 | 0 | — |
case-15 | pass→pass | 22,769 | 20,006 | -12% | 1 | 1 | 0% | 2,676 | 4,173 | +56% | 0 | 0 | — |
case-16 | pass→pass | 14,536 | 3,224 | -78% | 1 | 1 | 0% | 1,653 | 1,750 | +6% | 0 | 0 | — |
case-17 | pass→pass | 17,063 | 3,869 | -77% | 1 | 1 | 0% | 2,739 | 2,099 | -23% | 0 | 0 | — |
case-18 | fail→pass | 9,643 | 1,932 | -80% | 1 | 1 | 0% | 870 | 1,752 | +101% | 0 | 0 | — |
case-19 | fail→pass | 17,475 | 2,286 | -87% | 1 | 1 | 0% | 2,190 | 1,827 | -17% | 0 | 0 | — |
case-20 | fail→pass | 10,150 | 1,827 | -82% | 1 | 1 | 0% | 1,636 | 1,640 | +0% | 0 | 0 | — |
case-21 | pass→pass | 16,596 | 19,483 | +17% | 1 | 1 | 0% | 2,866 | 4,974 | +74% | 0 | 0 | — |
case-22 | pass→pass | 19,617 | 12,656 | -35% | 1 | 1 | 0% | 2,765 | 3,643 | +32% | 0 | 0 | — |
case-23 | pass→pass | 28,888 | 22,388 | -23% | 1 | 1 | 0% | 3,957 | 6,542 | +65% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +30 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.