Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Implement Abridge reference architecture for clinical AI integration. Use when designing a new Abridge deployment, reviewing project structure, or planning multi-site health system rollouts with EHR integration. Trigger: "abridge architecture", "abridge project structure", "abridge system design", "abridge multi-site".
.claude/skills/jeremylongshore-abridge-reference-architecture/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 126% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 89% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 121% | 0% |
Produce a tenant-grounded component and trust-boundary model for clinician capture, Abridge processing, review, Linked Evidence, EHR handoff, identity, support, telemetry, and downtime. Show unknown private interfaces explicitly instead of filling them with generic REST or FHIR assumptions.
Use Read, Glob, and Grep to inspect repository configuration, adapters, tests, policies, and existing evidence. Use WebFetch only for current official Abridge, HHS, or named EHR documentation. Use Write or Edit only after confirming scope, environment, owners, patient-data boundary, and approval state. These tools do not confer access to Abridge, an EHR, or a clinical record; return exact operator steps or an approval-gated handoff for live actions.
Use only the health system's provisioned Abridge application access, SSO, administrative role, or tenant-specific partner authentication documented for the approved environment. Do not infer public API credentials, reuse production secrets in tests, or expose tokens and session material. Verify identity owner, least privilege, environment binding, storage, rotation, and revocation before any authenticated action.
Read, Glob, and Grep to locate interface control documents, data-flow diagrams, retention decisions, and threat models.WebFetch only for current official Abridge product context and label all tenant-specific facts by source revision.Write or Edit to update the architecture record and unresolved evidence register.Do not label an inferred interface, FHIR resource, webhook, hosting model, or retention period as implemented without authoritative evidence.
Return component map, flows, trust boundaries, data classes, systems of record, owners, failure paths, evidence revisions, and open assumptions. Separate verified facts, tenant-specific evidence, assumptions, and actions still awaiting approval.
| Condition | Response | |---|---| | Edge lacks an owner | Keep the design unapproved. | | PHI crosses an undocumented boundary | Stop and escalate to privacy and security review. | | Diagram conflicts with tenant evidence | Update the diagram and record the superseded assumption. |
The example is a redacted operational receipt, not patient data or proof of vendor certification.
textsettings=outpatient+ed; ehr=epic; boundaries=7; undocumented-edges=2; phi-flows=owner-reviewed; status=design-review
Read the source map before changing a workflow. Recheck tenant-specific implementation evidence for every interface or capability that public documentation does not define.
Revalidate the evidence date and tenant-specific authority before repeating this workflow in another environment, cohort, care setting, or integration mode.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 29,960 | 26,605 | -11% | 1 | 1 | 0% | 6,197 | 7,936 | +28% | 0 | 0 | — |
case-01 | fail→fail | 33,213 | 28,516 | -14% | 1 | 1 | 0% | 6,201 | 7,940 | +28% | 0 | 0 | — |
case-03 | fail→pass | 33,212 | 22,900 | -31% | 1 | 1 | 0% | 6,205 | 7,032 | +13% | 0 | 0 | — |
case-04 | fail→pass | 17,425 | 22,770 | +31% | 1 | 1 | 0% | 3,089 | 6,983 | +126% | 0 | 0 | — |
case-05 | pass→pass | 6,798 | 5,853 | -14% | 1 | 1 | 0% | 1,102 | 2,828 | +157% | 0 | 0 | — |
case-06 | fail→pass | 14,667 | 14,211 | -3% | 1 | 1 | 0% | 2,366 | 4,465 | +89% | 0 | 0 | — |
case-07 | fail→pass | 9,054 | 9,137 | +1% | 1 | 1 | 0% | 1,518 | 3,360 | +121% | 0 | 0 | — |
case-08 | pass→pass | 14,472 | 15,520 | +7% | 1 | 1 | 0% | 2,501 | 5,001 | +100% | 0 | 0 | — |
case-09 | fail→fail | 14,625 | 10,857 | -26% | 1 | 1 | 0% | 3,163 | 4,209 | +33% | 0 | 0 | — |
case-10 | pass→pass | 12,937 | 10,425 | -19% | 1 | 1 | 0% | 2,350 | 3,886 | +65% | 0 | 0 | — |
case-11 | fail→pass | 9,436 | 5,493 | -42% | 1 | 1 | 0% | 1,843 | 2,736 | +48% | 0 | 0 | — |
case-20 | pass→pass | 3,250 | 3,547 | +9% | 1 | 1 | 0% | 634 | 2,425 | +282% | 0 | 0 | — |
case-12 | fail→pass | 17,002 | 10,609 | -38% | 1 | 1 | 0% | 3,214 | 4,078 | +27% | 0 | 0 | — |
case-13 | fail→pass | 13,028 | 3,624 | -72% | 1 | 1 | 0% | 2,529 | 2,329 | -8% | 0 | 0 | — |
case-14 | fail→pass | 9,324 | 2,291 | -75% | 1 | 1 | 0% | 1,737 | 2,139 | +23% | 0 | 0 | — |
case-15 | fail→pass | 11,283 | 1,653 | -85% | 1 | 1 | 0% | 2,077 | 1,974 | -5% | 0 | 0 | — |
case-16 | fail→pass | 14,370 | 14,217 | -1% | 1 | 1 | 0% | 2,861 | 5,031 | +76% | 0 | 0 | — |
case-17 | fail→pass | 6,946 | 1,551 | -78% | 1 | 1 | 0% | 1,253 | 1,983 | +58% | 0 | 0 | — |
case-18 | pass→pass | 9,095 | 5,180 | -43% | 1 | 1 | 0% | 1,597 | 2,578 | +61% | 0 | 0 | — |
case-19 | fail→pass | 8,768 | 2,201 | -75% | 1 | 1 | 0% | 1,507 | 2,118 | +41% | 0 | 0 | — |
case-21 | pass→pass | 8,920 | 7,638 | -14% | 1 | 1 | 0% | 1,714 | 3,176 | +85% | 0 | 0 | — |
case-22 | pass→pass | 15,064 | 14,038 | -7% | 1 | 1 | 0% | 2,731 | 4,456 | +63% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.