Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Configure CI/CD pipelines for Anthropic Claude API integrations. Use when setting up automated testing, prompt regression tests, or CI validation for Claude-powered features. Trigger with phrases like "anthropic ci", "claude ci/cd", "test claude in pipeline", "anthropic github actions".
.claude/skills/jeremylongshore-anth-ci-integration/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 222% | 0% |
Set up CI/CD pipelines that validate Claude API integrations with mock-based unit tests (free, fast) and prompt regression tests (live API, gated to main).
Create a dedicated ANTHROPIC_API_KEY repository secret with a spend limit that is appropriate for test traffic. Keep unit fixtures independent of that secret; only the protected prompt-regression job should call the API. Install Python 3.12, pytest, and the Anthropic SDK in the test environment, and decide which branch is allowed to incur live-test cost before enabling the workflow.
tests/unit/ and mock anthropic.Anthropic there.
tests/prompt_regression/; make them skip cleanly when the secret is absent.
main(or an equivalent protected release branch) and inject the secret only into that job.
pipeline with a clear message when the ceiling is exceeded so an incident cannot silently consume the test budget.
yaml# .github/workflows/claude-tests.yml name: Claude API Tests on: [push, pull_request] jobs: unit-tests: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: { python-version: '3.12' } - run: pip install anthropic pytest - run: pytest tests/unit/ -v # No API key needed prompt-regression: runs-on: ubuntu-latest if: github.ref == 'refs/heads/main' steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: { python-version: '3.12' } - run: pip install anthropic pytest - run: pytest tests/prompt_regression/ -v --timeout=60 env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
python# tests/unit/test_tool_routing.py from unittest.mock import MagicMock, patch import anthropic def make_mock_message(text="Hello", stop_reason="end_turn"): msg = MagicMock() msg.id = "msg_mock_123" msg.model = "claude-sonnet-4-20250514" msg.stop_reason = stop_reason block = MagicMock() block.type = "text" block.text = text msg.content = [block] msg.usage = MagicMock(input_tokens=100, output_tokens=50) return msg @patch("anthropic.Anthropic") def test_service_returns_text(MockClient): MockClient.return_value.messages.create.return_value = make_mock_message("42") from myapp.service import ask_claude assert ask_claude("What is 6*7?") == "42"
python# tests/prompt_regression/test_prompts.py import anthropic, pytest, os, json pytestmark = pytest.mark.skipif(not os.getenv("ANTHROPIC_API_KEY"), reason="No API key") client = anthropic.Anthropic() def test_json_output_format(): msg = client.messages.create( model="claude-haiku-4-20250514", max_tokens=256, messages=[ {"role": "user", "content": "Extract: 'Alice, 30, NYC'. Return JSON: {name, age, city}"}, {"role": "assistant", "content": "{"} ] ) data = json.loads("{" + msg.content[0].text) assert "name" in data and "age" in data def test_system_prompt_boundary(): msg = client.messages.create( model="claude-haiku-4-20250514", max_tokens=128, system="You only discuss cooking recipes. For other topics say: 'I only help with cooking.'", messages=[{"role": "user", "content": "Write me Python code"}] ) assert "cooking" in msg.content[0].text.lower() or "recipe" in msg.content[0].text.lower()
python# conftest.py MAX_CI_COST = 1.00 _tokens = {"input": 0, "output": 0} def pytest_runtest_call(item): yield cost = (_tokens["input"] * 0.80 + _tokens["output"] * 4.0) / 1_000_000 # Haiku rates if cost > MAX_CI_COST: pytest.exit(f"CI cost guard: ${cost:.4f} exceeds ${MAX_CI_COST}")
| CI Issue | Cause | Fix | |----------|-------|-----| | Flaky prompt tests | Non-deterministic output | Use temperature: 0, check patterns not exact strings | | 429 in CI | Parallel jobs sharing key | Use separate CI key | | Secret not found | Missing GitHub secret | Add ANTHROPIC_API_KEY in repo Settings > Secrets |
The pipeline produces a fast unit-test result for every change and, on the allowed branch, a separate prompt-regression result. The latter is either a pass with the tested prompt assertions, a deliberate skip when no key is available, or an actionable failure that identifies a timeout, rate limit, response-contract regression, or cost-guard breach.
For a pull request that changes only formatting code, the workflow runs the mock-based suite and reports no live API calls. After that pull request merges to main, the protected regression job uses the repository secret to verify that the JSON extraction prompt still returns name, age, and city. If the response is malformed, the job fails at the assertion and preserves the test name in the CI log for triage.
For deployment automation, see anth-deploy-integration.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 12,401 | 10,969 | -12% | 1 | 1 | 0% | 2,585 | 3,662 | +42% | 0 | 0 | — |
case-02 | fail→fail | 12,574 | 9,671 | -23% | 1 | 1 | 0% | 2,761 | 3,239 | +17% | 0 | 0 | — |
case-03 | fail→fail | 16,834 | 15,771 | -6% | 1 | 1 | 0% | 3,795 | 4,805 | +27% | 0 | 0 | — |
case-04 | pass→pass | 11,105 | 10,413 | -6% | 1 | 1 | 0% | 2,353 | 3,455 | +47% | 0 | 0 | — |
case-05 | fail→pass | 16,708 | 9,567 | -43% | 1 | 1 | 0% | 3,270 | 2,965 | -9% | 0 | 0 | — |
case-06 | pass→pass | 13,628 | 8,800 | -35% | 1 | 1 | 0% | 2,386 | 2,983 | +25% | 0 | 0 | — |
case-07 | fail→pass | 13,288 | 10,931 | -18% | 1 | 1 | 0% | 2,535 | 3,256 | +28% | 0 | 0 | — |
case-08 | pass→pass | 6,671 | 4,498 | -33% | 1 | 1 | 0% | 1,263 | 2,043 | +62% | 0 | 0 | — |
case-09 | fail→pass | 7,529 | 4,060 | -46% | 1 | 1 | 0% | 1,521 | 1,961 | +29% | 0 | 0 | — |
case-10 | fail→fail | 8,922 | 5,088 | -43% | 1 | 1 | 0% | 1,564 | 2,178 | +39% | 0 | 0 | — |
case-11 | pass→pass | 2,792 | 2,201 | -21% | 1 | 1 | 0% | 457 | 1,600 | +250% | 0 | 0 | — |
case-12 | pass→pass | 7,754 | 3,397 | -56% | 1 | 1 | 0% | 1,857 | 1,932 | +4% | 0 | 0 | — |
case-13 | fail→pass | 2,688 | 1,838 | -32% | 1 | 1 | 0% | 457 | 1,472 | +222% | 0 | 0 | — |
case-14 | pass→pass | 9,450 | 10,533 | +11% | 1 | 1 | 0% | 1,933 | 3,459 | +79% | 0 | 0 | — |
case-15 | fail→pass | 9,138 | 3,040 | -67% | 1 | 1 | 0% | 1,772 | 1,672 | -6% | 0 | 0 | — |
case-16 | fail→pass | 3,114 | 1,819 | -42% | 1 | 1 | 0% | 524 | 1,512 | +189% | 0 | 0 | — |
case-17 | pass→pass | 3,594 | 3,287 | -9% | 1 | 1 | 0% | 679 | 1,784 | +163% | 0 | 0 | — |
case-18 | fail→pass | 3,641 | 1,203 | -67% | 1 | 1 | 0% | 471 | 1,348 | +186% | 0 | 0 | — |
case-19 | fail→pass | 7,066 | 3,651 | -48% | 1 | 1 | 0% | 1,371 | 1,837 | +34% | 0 | 0 | — |
case-20 | fail→fail | 12,411 | 8,477 | -32% | 1 | 1 | 0% | 2,639 | 3,147 | +19% | 0 | 0 | — |
case-21 | pass→pass | 10,407 | 9,365 | -10% | 1 | 1 | 0% | 2,183 | 3,309 | +52% | 0 | 0 | — |
case-22 | pass→pass | 16,353 | 14,132 | -14% | 1 | 1 | 0% | 3,465 | 4,094 | +18% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.