Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Collect Abridge debug evidence for support tickets and troubleshooting. Use when filing Abridge support tickets, collecting diagnostic data, or preparing evidence for escalation to Abridge engineering. Trigger: "abridge debug bundle", "abridge support ticket", "abridge diagnostics", "collect abridge evidence".
.claude/skills/jeremylongshore-abridge-debug-bundle/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 127% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 47% | 0% |
Collect configuration presence, component versions, coarse timestamps, correlation references, and observed outcomes without copying recordings, transcripts, generated notes, patient identifiers, tokens, or unrestricted logs.
Use Read, Glob, and Grep to inspect repository configuration, adapters, tests, policies, and existing evidence. Use WebFetch only for current official Abridge, HHS, or named EHR documentation. Use Write or Edit only after confirming scope, environment, owners, patient-data boundary, and approval state. These tools do not confer access to Abridge, an EHR, or a clinical record; return exact operator steps or an approval-gated handoff for live actions.
Use only the health system's provisioned Abridge application access, SSO, administrative role, or tenant-specific partner authentication documented for the approved environment. Do not infer public API credentials, reuse production secrets in tests, or expose tokens and session material. Verify identity owner, least privilege, environment binding, storage, rotation, and revocation before any authenticated action.
Read, Glob, and Grep to locate approved logs and configuration names; do not recursively collect an application directory.Write or Edit to create the bounded bundle and a manifest; have a second reviewer attest to the redaction.WebFetch only to confirm current official Abridge security and support guidance before transfer.Do not attach a bundle to email, chat, or a public issue unless the privacy and security owners approve that channel and the bundle passes review.
Return manifest, source classes, allowlisted fields, redaction counts, reviewer, destination, expiry, and excluded evidence classes. Separate verified facts, tenant-specific evidence, assumptions, and actions still awaiting approval.
| Condition | Response | |---|---| | Sensitive value cannot be classified | Exclude it by default. | | Recipient requests raw logs | Move to the approved protected disclosure process. | | Bundle cannot reproduce the symptom | Document the limitation; do not broaden collection silently. |
The example is a redacted operational receipt, not patient data or proof of vendor certification.
textincident=INC-2041; window=15m-rounded; files=3; rejected-fields=17; phi=none-observed; reviewer=privacy-oncall; expiry=7d
Read the source map before changing a workflow. Recheck tenant-specific implementation evidence for every interface or capability that public documentation does not define.
Revalidate the evidence date and tenant-specific authority before repeating this workflow in another environment, cohort, care setting, or integration mode.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,780 | 18,411 | +17% | 1 | 1 | 0% | 3,612 | 6,098 | +69% | 0 | 0 | — |
case-02 | fail→pass | 14,604 | 10,178 | -30% | 1 | 1 | 0% | 3,500 | 3,939 | +13% | 0 | 0 | — |
case-03 | fail→pass | 10,917 | 15,660 | +43% | 1 | 1 | 0% | 2,350 | 5,331 | +127% | 0 | 0 | — |
case-04 | pass→pass | 18,951 | 13,774 | -27% | 1 | 1 | 0% | 4,244 | 4,813 | +13% | 0 | 0 | — |
case-05 | pass→pass | 18,188 | 14,176 | -22% | 1 | 1 | 0% | 3,967 | 4,756 | +20% | 0 | 0 | — |
case-06 | pass→pass | 11,187 | 9,552 | -15% | 1 | 1 | 0% | 2,474 | 3,722 | +50% | 0 | 0 | — |
case-07 | fail→pass | 23,504 | 2,415 | -90% | 1 | 1 | 0% | 2,323 | 1,958 | -16% | 0 | 0 | — |
case-08 | fail→pass | 16,232 | 16,482 | +2% | 1 | 1 | 0% | 3,675 | 5,404 | +47% | 0 | 0 | — |
case-09 | pass→pass | 17,662 | 9,337 | -47% | 1 | 1 | 0% | 3,695 | 3,900 | +6% | 0 | 0 | — |
case-10 | fail→fail | 14,673 | 10,189 | -31% | 1 | 1 | 0% | 3,163 | 3,969 | +25% | 0 | 0 | — |
case-11 | fail→pass | 11,712 | 8,042 | -31% | 1 | 1 | 0% | 2,452 | 3,264 | +33% | 0 | 0 | — |
case-12 | fail→pass | 15,128 | 9,684 | -36% | 1 | 1 | 0% | 2,964 | 3,662 | +24% | 0 | 0 | — |
case-17 | pass→pass | 5,608 | 3,893 | -31% | 1 | 1 | 0% | 1,121 | 2,402 | +114% | 0 | 0 | — |
case-13 | pass→pass | 11,966 | 8,870 | -26% | 1 | 1 | 0% | 2,412 | 3,282 | +36% | 0 | 0 | — |
case-14 | fail→pass | 17,840 | 15,036 | -16% | 1 | 1 | 0% | 3,817 | 4,810 | +26% | 0 | 0 | — |
case-15 | fail→pass | 4,530 | 2,878 | -36% | 1 | 1 | 0% | 949 | 2,162 | +128% | 0 | 0 | — |
case-16 | pass→pass | 4,349 | 3,484 | -20% | 1 | 1 | 0% | 855 | 2,352 | +175% | 0 | 0 | — |
case-18 | fail→pass | 11,578 | 5,826 | -50% | 1 | 1 | 0% | 2,534 | 2,817 | +11% | 0 | 0 | — |
case-19 | fail→pass | 12,656 | 4,573 | -64% | 1 | 1 | 0% | 2,455 | 2,593 | +6% | 0 | 0 | — |
case-20 | fail→pass | 9,158 | 8,850 | -3% | 1 | 1 | 0% | 1,742 | 3,403 | +95% | 0 | 0 | — |
case-21 | fail→pass | 15,462 | 5,767 | -63% | 1 | 1 | 0% | 2,851 | 2,778 | -3% | 0 | 0 | — |
case-22 | fail→fail | 14,970 | 8,542 | -43% | 1 | 1 | 0% | 2,729 | 3,320 | +22% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.