Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Full CLI test pipeline — monitor CI for types/lint, then run local binary test, then run bundled binary test via CI.
.claude/skills/bilal140202-full-cli-test/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 99% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 76% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 20% | 0% |
End-to-end CLI validation pipeline that ensures types/lint pass in CI, then tests the CLI binary built from source, and finally tests the CI-bundled binary.
This skill runs three phases sequentially:
/cli-test)./cli-test-with-bundling).Use /loop 5m to poll the CI status for the current branch every 5 minutes. Check that all type-check and lint jobs pass for the CLI package before proceeding.
bash# Get the latest CI run for the current branch gh run list \ --repo ComposioHQ/composio \ --branch "$(git rev-parse --abbrev-ref HEAD)" \ --limit 5 \ --json databaseId,status,conclusion,name,headBranch
Look for workflow runs that cover type-checking and linting (e.g., the main CI workflow). If none are running or completed yet, wait for one to appear.
Use /loop 5m to poll: run the gh run list / gh run view command every 5 minutes until the relevant jobs succeed. Once all type/lint checks are green, move on.
/cli-test)Once CI lint/types pass, run the /cli-test skill:
pnpm installpnpm turbo buildpnpm --dir ts/packages/cli build:binarybash ./ts/packages/cli/dist/composio version ./ts/packages/cli/dist/composio whoami ./ts/packages/cli/dist/composio --help
bash ./ts/packages/cli/dist/composio run '<SLACK_TEST_SCRIPT>'
If any command fails, report to the user and stop — do NOT proceed to Phase 3.
/cli-test-with-bundling)Once the local binary test passes, run the /cli-test-with-bundling skill:
ts/packages/cli/package.json and append a -beta.<timestamp> suffix (e.g. 1.2.3-beta.20260331143022) — always trigger as a beta releasebuild-cli-binaries.yml workflow via gh workflow run with the beta version/loop 5m to poll)bash $BINARY version $BINARY whoami $BINARY --help $BINARY run 'console.log("hello from composio run")' $BINARY run 'const result = await experimental_subAgent({ goal: "What is 2+2?", toolNames: [] }); console.log(result)'
bash $BINARY run '<SLACK_TEST_SCRIPT>'
This test validates execute(), experimental_subAgent(), and end-to-end Slack connectivity. It must be run in both Phase 2 and Phase 3 to ensure it works with both the locally-built and CI-bundled binaries.
#buzz-skill-based-cli-testing — dedicated Slack channel for automated CLI test runs.
Use composio run (or $BINARY run in Phase 3) with the following script. Replace $BINARY with the appropriate binary path for the phase.
bash$BINARY run ' // Step 1: Find the #buzz-skill-based-cli-testing channel const channels = await execute("SLACK_LIST_CHANNELS", { types: "public_channel", limit: 200, }); const channel = channels.data?.channels?.find( (c) => c.name === "buzz-skill-based-cli-testing" ); if (!channel) throw new Error("Channel #buzz-skill-based-cli-testing not found"); // Step 2: Send an initial message tagging @cryogenicplanet const buildType = "local"; // use "bundled" in Phase 3 await execute("SLACK_SEND_A_MESSAGE_TO_A_SLACK_CHANNEL", { channel: channel.id, text: `<@cryogenicplanet> CLI test run started (${buildType} build) at ${new Date().toISOString()}`, }); // Step 3: Fetch recent channel history for the subAgent to summarize const history = await execute("SLACK_GET_CHANNEL_HISTORY", { channel: channel.id, limit: 20, }); // Step 4: Use experimental_subAgent to summarize what happened in the channel const summary = await experimental_subAgent( `Summarize the recent activity in this Slack channel in 2-3 sentences. Focus on what tests were run and their outcomes.\n\n${history.prompt()}`, { schema: z.object({ summary: z.string(), messageCount: z.number(), }), } ); // Step 5: Post the summary back to the channel await execute("SLACK_SEND_A_MESSAGE_TO_A_SLACK_CHANNEL", { channel: channel.id, text: `CLI Test Summary (${buildType} build):\n${summary.structuredOutput.summary}\n(${summary.structuredOutput.messageCount} messages analyzed)`, }); console.log("Slack integration test passed:", summary.structuredOutput); '
| Capability | How it's tested | |---|---| | execute() with Slack tools | List channels, send messages, fetch history | | experimental_subAgent() | Summarizes channel history with structured output via z schema | | z (Zod) global | Used in the subAgent schema definition | | result.prompt() | Feeds channel history into the subAgent | | End-to-end Slack connectivity | Reads from and writes to a real Slack channel |
buildType to "bundled" when running in Phase 3.<@cryogenicplanet> should resolve to the correct user in your workspace. If the mention doesn't resolve, find the Slack user ID first and use <@U_XXXXX> format.composio connect slack first.| File | Purpose | |---|---| | .github/workflows/build-cli-binaries.yml | CI workflow that builds binaries | | ts/packages/cli/package.json | Source of CLI version | | ts/packages/cli/scripts/build-binary.ts | Local binary build script | | ts/packages/cli/dist/composio | Built binary output |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,804 | 2,843 | -68% | 1 | 1 | 0% | 1,811 | 2,365 | +31% | 0 | 0 | — |
case-02 | fail→fail | 7,724 | 3,128 | -60% | 1 | 1 | 0% | 1,478 | 2,335 | +58% | 0 | 0 | — |
case-03 | fail→fail | 5,102 | 4,620 | -9% | 1 | 1 | 0% | 737 | 2,689 | +265% | 0 | 0 | — |
case-04 | pass→fail | 10,067 | 3,735 | -63% | 1 | 1 | 0% | 2,311 | 2,079 | -10% | 0 | 0 | — |
case-05 | pass→fail | 7,921 | 4,016 | -49% | 1 | 1 | 0% | 1,548 | 2,115 | +37% | 0 | 0 | — |
case-06 | pass→pass | 7,225 | 3,684 | -49% | 1 | 1 | 0% | 1,386 | 2,584 | +86% | 0 | 0 | — |
case-07 | fail→pass | 5,977 | 1,969 | -67% | 1 | 1 | 0% | 1,084 | 2,159 | +99% | 0 | 0 | — |
case-08 | fail→pass | 9,884 | 1,498 | -85% | 1 | 1 | 0% | 1,704 | 2,087 | +22% | 0 | 0 | — |
case-09 | fail→pass | 8,254 | 2,590 | -69% | 1 | 1 | 0% | 1,315 | 2,311 | +76% | 0 | 0 | — |
case-10 | fail→pass | 11,143 | 1,470 | -87% | 1 | 1 | 0% | 2,235 | 2,168 | -3% | 0 | 0 | — |
case-11 | fail→pass | 9,207 | 1,422 | -85% | 1 | 1 | 0% | 1,802 | 2,167 | +20% | 0 | 0 | — |
case-12 | fail→pass | 9,357 | 1,974 | -79% | 1 | 1 | 0% | 1,840 | 2,265 | +23% | 0 | 0 | — |
case-13 | fail→pass | 7,245 | 2,277 | -69% | 1 | 1 | 0% | 1,339 | 2,310 | +73% | 0 | 0 | — |
case-14 | fail→pass | 7,763 | 1,468 | -81% | 1 | 1 | 0% | 1,385 | 2,141 | +55% | 0 | 0 | — |
case-15 | fail→pass | 10,384 | 1,981 | -81% | 1 | 1 | 0% | 1,734 | 2,274 | +31% | 0 | 0 | — |
case-16 | fail→pass | 7,050 | 1,278 | -82% | 1 | 1 | 0% | 1,180 | 2,046 | +73% | 0 | 0 | — |
case-17 | fail→pass | 5,335 | 2,845 | -47% | 1 | 1 | 0% | 951 | 2,350 | +147% | 0 | 0 | — |
case-18 | pass→pass | 7,530 | 1,616 | -79% | 1 | 1 | 0% | 1,274 | 2,199 | +73% | 0 | 0 | — |
case-19 | pass→pass | 5,391 | 1,744 | -68% | 1 | 1 | 0% | 1,070 | 2,196 | +105% | 0 | 0 | — |
case-20 | fail→pass | 10,589 | 2,850 | -73% | 1 | 1 | 0% | 1,986 | 2,449 | +23% | 0 | 0 | — |
case-21 | pass→pass | 6,956 | 1,739 | -75% | 1 | 1 | 0% | 1,171 | 2,199 | +88% | 0 | 0 | — |
case-22 | fail→fail | 6,238 | 1,655 | -73% | 1 | 1 | 0% | 1,050 | 2,150 | +105% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.