Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Analyze a GitHub issue, verify claims against the codebase, and close invalid issues with a technical response.
.claude/skills/aden-hive-triage-issue-skill/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 91% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 11% | 0% |
Analyze a GitHub issue, verify claims against the codebase, and close invalid issues with a technical response.
User provides a GitHub issue URL or number, e.g.:
/triage-issue 1970/triage-issue https://github.com/adenhq/hive/issues/1970bashgh issue view <number> --repo adenhq/hive --json title,body,state,labels,author
Extract:
If issue is already closed, inform user and stop.
Read the issue body and identify:
For each technical claim:
Categorize the issue as one of:
| Category | Action | |----------|--------| | Valid Bug | Do NOT close. Inform user this is a real issue. | | Valid Feature Request | Do NOT close. Suggest labeling appropriately. | | Misunderstanding | Prepare technical explanation for why behavior is correct. | | Fundamentally Flawed | Prepare critique explaining the technical impossibility or design rationale. | | Duplicate | Find the original issue and prepare duplicate notice. | | Incomplete | Prepare request for more information. |
For issues to be closed, draft a response that:
Use this template:
markdown## Analysis [Brief summary of what was investigated] ## Technical Details [Explanation with code references] ## Why This Is Working As Designed [Rationale] ## Recommendation [What the user should do instead, if applicable] --- *This issue was reviewed and closed by the maintainers.*
Present the draft to the user with:
## Issue #<number>: <title>
**Claim:** <summary of claim>
**Finding:** <valid/invalid/misunderstanding/etc>
**Draft Response:**
<the markdown response>
---
Do you want me to post this comment and close the issue?Use AskUserQuestion with options:
If user approves:
bash# Post comment gh issue comment <number> --repo adenhq/hive --body "<response>" # Close issue gh issue close <number> --repo adenhq/hive --reason "not planned"
Report success with link to the issue.
> "The claim that secrets are exposed in plaintext misunderstands the encryption architecture. While SecretStr is used for logging protection, actual encryption is provided by Fernet (AES-128-CBC) at the storage layer. The code path is: serialize → encrypt → write. Only encrypted bytes touch disk."
> "The requested feature would require X] which violates fundamental constraint]. This is not a limitation of our implementation but a fundamental property of technology/protocol]."
> "This scenario is already handled by code reference]. The reporter may be using an older version or misconfigured environment."
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | fail→pass | 4,430 | 1,811 | -59% | 1 | 1 | 0% | 727 | 1,387 | +91% | 0 | 0 | — |
case-01 | fail→fail | 7,249 | 3,569 | -51% | 1 | 1 | 0% | 316 | 1,233 | +290% | 0 | 0 | — |
case-02 | fail→fail | 28,214 | 29,240 | +4% | 1 | 1 | 0% | 3,734 | 1,322 | -65% | 0 | 0 | — |
case-03 | fail→fail | 12,350 | 3,936 | -68% | 1 | 1 | 0% | 1,390 | 1,249 | -10% | 0 | 0 | — |
case-04 | pass→fail | 6,907 | 6,686 | -3% | 1 | 1 | 0% | 1,192 | 1,461 | +23% | 0 | 0 | — |
case-05 | pass→fail | 8,142 | 7,686 | -6% | 1 | 1 | 0% | 1,335 | 1,516 | +14% | 0 | 0 | — |
case-06 | fail→fail | 6,264 | 6,139 | -2% | 1 | 1 | 0% | 425 | 1,339 | +215% | 0 | 0 | — |
case-07 | pass→fail | 12,329 | 5,772 | -53% | 1 | 1 | 0% | 1,960 | 1,312 | -33% | 0 | 0 | — |
case-08 | fail→fail | 10,222 | 7,221 | -29% | 1 | 1 | 0% | 1,480 | 1,454 | -2% | 0 | 0 | — |
case-09 | pass→pass | 10,470 | 3,714 | -65% | 1 | 1 | 0% | 1,582 | 1,633 | +3% | 0 | 0 | — |
case-10 | pass→pass | 7,445 | 2,175 | -71% | 1 | 1 | 0% | 1,131 | 1,407 | +24% | 0 | 0 | — |
case-11 | pass→pass | 7,253 | 2,554 | -65% | 1 | 1 | 0% | 1,239 | 1,477 | +19% | 0 | 0 | — |
case-12 | pass→pass | 3,873 | 1,735 | -55% | 1 | 1 | 0% | 668 | 1,314 | +97% | 0 | 0 | — |
case-13 | fail→pass | 21,770 | 1,613 | -93% | 1 | 1 | 0% | 1,397 | 1,275 | -9% | 0 | 0 | — |
case-14 | fail→pass | 7,744 | 1,481 | -81% | 1 | 1 | 0% | 1,390 | 1,355 | -3% | 0 | 0 | — |
case-16 | fail→pass | 8,810 | 2,741 | -69% | 1 | 1 | 0% | 1,388 | 1,509 | +9% | 0 | 0 | — |
case-17 | pass→pass | 3,533 | 2,468 | -30% | 1 | 1 | 0% | 544 | 1,480 | +172% | 0 | 0 | — |
case-18 | pass→pass | 5,268 | 3,205 | -39% | 1 | 1 | 0% | 812 | 1,588 | +96% | 0 | 0 | — |
case-19 | fail→fail | 7,846 | 1,702 | -78% | 1 | 1 | 0% | 1,247 | 1,320 | +6% | 0 | 0 | — |
case-20 | pass→pass | 5,844 | 2,413 | -59% | 1 | 1 | 0% | 992 | 1,344 | +35% | 0 | 0 | — |
case-21 | fail→pass | 6,761 | 2,217 | -67% | 1 | 1 | 0% | 1,217 | 1,354 | +11% | 0 | 0 | — |
case-22 | pass→pass | 7,301 | 5,721 | -22% | 1 | 1 | 0% | 1,204 | 2,175 | +81% | 0 | 0 | — |
case-23 | pass→pass | 14,537 | 5,706 | -61% | 1 | 1 | 0% | 2,244 | 2,041 | -9% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 14 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 14 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.