Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Set up local development for Clari API integrations with mock data. Use when building forecast dashboards, testing export pipelines, or iterating on Clari data transformations locally. Trigger with phrases like "clari dev setup", "clari local testing", "develop with clari", "clari mock data".
.claude/skills/jeremylongshore-clari-local-dev-loop/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -36% | 0% |
Make offline development the default and live provider calls an explicit integration stage. Model the provider’s authentication headers, job transitions, pagination, error objects, and schema drift without storing secrets or production records.
List endpoints, headers, request fields, response fields, job states, and pagination semantics used by the integration.
Create minimal payloads that preserve shapes and identifiers while containing no real people, accounts, calls, forecasts, or deal values.
Make the test server move deterministically through queued, running, completed, aborted, rate-limited, and timed-out paths.
Run parser, schema, retry, pagination, idempotency, and redaction tests against the simulator with network access disabled.
Use a dedicated non-production identity and the smallest read-only request; skip it unless credentials and explicit integration-test approval are present.
When the provider contract changes, review the diff, update fixtures and assertions together, and retain the old failing fixture as migration evidence.
Offline tests must use obvious dummy values. A gated live test may read credentials from the approved secret manager at runtime, but must never serialize headers or provider payloads into test output.
Use Read and Grep to inspect configuration, provider contracts, fixtures, logs, schemas, and existing tests before proposing a change. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not issue, rotate, revoke, create, update, cancel, delete, export, ingest, or publish provider data without explicit operator approval.
Return the exact surface, environment, resource or job identifiers, contract fingerprint, evidence, unresolved risks, and final decision without exposing credentials or sensitive customer data.
A local export client receives synthetic SCHEDULED, STARTED, and DONE responses, then a 429 and an ABORTED job. CI proves bounded retry and redaction without consuming a Clari export.
| Failure | Response | | --- | --- | | Fixture contains real customer data | Quarantine and remove it, rotate any exposed credential, and replace it with generated values. | | Mock diverges from the provider contract | Pin the current contract fingerprint and add a regression fixture for the observed delta. | | Live smoke runs unexpectedly | Fail closed unless an explicit environment gate and non-production credential are both present. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | pass→pass | 13,841 | 12,507 | -10% | 1 | 1 | 0% | 2,403 | 2,843 | +18% | 0 | 0 | — |
case-01 | fail→pass | 15,694 | 11,986 | -24% | 1 | 1 | 0% | 3,211 | 3,562 | +11% | 0 | 0 | — |
case-02 | fail→pass | 13,259 | 6,774 | -49% | 1 | 1 | 0% | 2,997 | 2,273 | -24% | 0 | 0 | — |
case-03 | fail→fail | 17,069 | 11,550 | -32% | 1 | 1 | 0% | 3,026 | 3,368 | +11% | 0 | 0 | — |
case-04 | pass→pass | 18,783 | 17,380 | -7% | 1 | 1 | 0% | 3,432 | 4,448 | +30% | 0 | 0 | — |
case-05 | pass→pass | 23,011 | 17,430 | -24% | 1 | 1 | 0% | 4,977 | 4,515 | -9% | 0 | 0 | — |
case-06 | pass→pass | 16,835 | 16,666 | -1% | 1 | 1 | 0% | 3,103 | 4,308 | +39% | 0 | 0 | — |
case-07 | fail→pass | 11,949 | 2,305 | -81% | 1 | 1 | 0% | 1,536 | 1,280 | -17% | 0 | 0 | — |
case-08 | fail→pass | 14,582 | 4,244 | -71% | 1 | 1 | 0% | 2,648 | 1,847 | -30% | 0 | 0 | — |
case-09 | fail→pass | 15,428 | 5,303 | -66% | 1 | 1 | 0% | 2,809 | 1,789 | -36% | 0 | 0 | — |
case-10 | fail→pass | 13,076 | 4,440 | -66% | 1 | 1 | 0% | 2,332 | 1,789 | -23% | 0 | 0 | — |
case-11 | pass→pass | 7,810 | 2,445 | -69% | 1 | 1 | 0% | 1,334 | 1,384 | +4% | 0 | 0 | — |
case-12 | pass→pass | 4,667 | 2,822 | -40% | 1 | 1 | 0% | 818 | 1,396 | +71% | 0 | 0 | — |
case-13 | pass→pass | 9,986 | 7,300 | -27% | 1 | 1 | 0% | 1,824 | 2,412 | +32% | 0 | 0 | — |
case-15 | fail→pass | 11,303 | 10,049 | -11% | 1 | 1 | 0% | 1,877 | 2,938 | +57% | 0 | 0 | — |
case-16 | pass→pass | 10,954 | 2,168 | -80% | 1 | 1 | 0% | 1,850 | 1,318 | -29% | 0 | 0 | — |
case-17 | fail→pass | 12,743 | 2,116 | -83% | 1 | 1 | 0% | 2,314 | 1,309 | -43% | 0 | 0 | — |
case-18 | fail→pass | 9,365 | 2,900 | -69% | 1 | 1 | 0% | 1,608 | 1,345 | -16% | 0 | 0 | — |
case-19 | fail→fail | 12,095 | 2,292 | -81% | 1 | 1 | 0% | 887 | 1,260 | +42% | 0 | 0 | — |
case-20 | fail→pass | 5,941 | 1,526 | -74% | 1 | 1 | 0% | 943 | 1,188 | +26% | 0 | 0 | — |
case-21 | fail→pass | 28,360 | 3,247 | -89% | 1 | 1 | 0% | 2,890 | 1,716 | -41% | 0 | 0 | — |
case-22 | fail→pass | 7,937 | 2,202 | -72% | 1 | 1 | 0% | 1,609 | 1,435 | -11% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 21 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.