Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute Cohere production deployment checklist and rollback procedures. Use when deploying Cohere integrations to production, preparing for launch, or implementing go-live procedures for Cohere-powered apps. Trigger with phrases like "cohere production", "deploy cohere", "cohere go-live", "cohere launch checklist".
.claude/skills/jeremylongshore-cohere-prod-checklist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -14% | 0% |
Convert a staged Cohere integration into a reversible production release with explicit owners, thresholds, and evidence.
Use Read, Glob, and Grep to inspect code, configuration, and evidence. Use WebFetch only for current Cohere primary documentation. Use Write or Edit only when the user requested implementation and the exact target files are known; never write credentials or customer content.
Use an environment-specific key injected from an approved secret manager. Never print, persist, commit, or place CO_API_KEY in an example. Confirm access with the least costly bounded operation appropriate to the task, and treat key creation, rotation, revocation, role changes, and production-capacity requests as owner-approved actions.
Do not expose or rotate keys, change Cohere Team roles, accept commercial terms, enable sensitive production data, increase spend or capacity, switch production models, send a support bundle, or execute model-proposed side effects without the accountable owner's approval. Keep diagnosis read-only unless implementation was requested.
Return the resolved API and model contract, files or settings inspected, evidence collected, validation result, remaining risk, owner, and rollback or next action. Redact keys, authorization headers, prompts, retrieved documents, embeddings, customer identifiers, and unrestricted environment output.
| Condition | Response | |---|---| | Trial key | Do not launch a public production workload on evaluation capacity. | | No quality gate | Block release until representative evaluation thresholds exist. | | No rollback | Block release until traffic can be restored safely. | | Model near retirement | Migrate or obtain an explicit time-bound exception. |
Use this compact handoff shape to keep the selected scope, validation evidence, and operational result reviewable.
Input:
textservice=support-rag; canary=5%; quality-threshold=approved; rollback=ready
Expected handoff:
textdecision=GO; model=resolved; capacity=verified; evidence=linked
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,463 | 13,221 | -20% | 1 | 1 | 0% | 2,931 | 4,375 | +49% | 0 | 0 | — |
case-02 | fail→pass | 16,539 | 16,223 | -2% | 1 | 1 | 0% | 3,535 | 5,348 | +51% | 0 | 0 | — |
case-03 | fail→fail | 19,225 | 16,731 | -13% | 1 | 1 | 0% | 3,765 | 5,097 | +35% | 0 | 0 | — |
case-04 | pass→pass | 14,730 | 16,502 | +12% | 1 | 1 | 0% | 2,845 | 5,013 | +76% | 0 | 0 | — |
case-05 | pass→pass | 17,552 | 11,783 | -33% | 1 | 1 | 0% | 3,733 | 4,334 | +16% | 0 | 0 | — |
case-06 | pass→pass | 15,206 | 14,226 | -6% | 1 | 1 | 0% | 2,809 | 4,396 | +56% | 0 | 0 | — |
case-07 | fail→pass | 5,923 | 2,838 | -52% | 1 | 1 | 0% | 1,070 | 2,210 | +107% | 0 | 0 | — |
case-08 | pass→pass | 7,821 | 2,275 | -71% | 1 | 1 | 0% | 1,497 | 2,009 | +34% | 0 | 0 | — |
case-09 | fail→pass | 9,694 | 2,382 | -75% | 1 | 1 | 0% | 1,859 | 1,998 | +7% | 0 | 0 | — |
case-10 | pass→pass | 5,213 | 1,704 | -67% | 1 | 1 | 0% | 956 | 1,928 | +102% | 0 | 0 | — |
case-11 | fail→pass | 13,148 | 2,634 | -80% | 1 | 1 | 0% | 2,428 | 2,087 | -14% | 0 | 0 | — |
case-12 | pass→pass | 4,251 | 2,600 | -39% | 1 | 1 | 0% | 711 | 1,979 | +178% | 0 | 0 | — |
case-13 | pass→pass | 4,456 | 2,535 | -43% | 1 | 1 | 0% | 797 | 2,040 | +156% | 0 | 0 | — |
case-14 | fail→pass | 9,708 | 10,732 | +11% | 1 | 1 | 0% | 2,041 | 3,675 | +80% | 0 | 0 | — |
case-15 | fail→pass | 10,273 | 4,520 | -56% | 1 | 1 | 0% | 1,803 | 2,516 | +40% | 0 | 0 | — |
case-16 | pass→pass | 9,663 | 3,017 | -69% | 1 | 1 | 0% | 1,705 | 2,091 | +23% | 0 | 0 | — |
case-17 | fail→pass | 10,903 | 6,275 | -42% | 1 | 1 | 0% | 1,853 | 2,678 | +45% | 0 | 0 | — |
case-18 | pass→pass | 9,253 | 1,645 | -82% | 1 | 1 | 0% | 1,549 | 1,921 | +24% | 0 | 0 | — |
case-19 | fail→pass | 11,723 | 3,459 | -70% | 1 | 1 | 0% | 2,035 | 2,208 | +9% | 0 | 0 | — |
case-20 | pass→pass | 12,567 | 2,546 | -80% | 1 | 1 | 0% | 2,180 | 2,076 | -5% | 0 | 0 | — |
case-21 | fail→pass | 4,410 | 2,384 | -46% | 1 | 1 | 0% | 852 | 2,026 | +138% | 0 | 0 | — |
case-22 | fail→pass | 7,489 | 2,993 | -60% | 1 | 1 | 0% | 1,273 | 2,111 | +66% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.