Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Turn walkthrough notes, photos, or voice-memo transcripts into a proper construction punch list with location, trade, and spec reference per item. Use when asked to build a punch list, clean up walkthrough notes, organise a deficiency list, prep for substantial completion, or track punch items to closeout. Produces a numbered punch list grouped by location with severity tiers, responsible subcontractor, back-charge candidates, and closeout/retainage linkage.
.claude/skills/mohitagw15856-punch-list-builder/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 242% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 233% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 120% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 580% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 77% | 0% |
"Fix paint in hallway" closes nothing. A punch item that closes reads: Level 2, Corridor 2C, north wall — drywall finish fails Level 4 requirement per spec 09 29 00; responsible: [drywall sub]; verify: repaint entire wall section, re-inspect under raking light. This skill converts messy walkthrough notes into that — a numbered, trade-assigned, spec-referenced punch list an owner's rep can sign off against and a super can actually run subs from.
Ask for what's missing; from raw notes alone, build the list and tag inferred fields [verify]:
[assign]Tier every item:
| Tier | Definition | Consequence | |---|---|---| | A — Blocks completion | Life-safety, code, non-functional systems, missing inspections | Blocks substantial completion / certificate of occupancy | | B — Blocks acceptance | Doesn't meet contract documents — wrong product, failed finish tolerance, incomplete scope | Blocks final payment / retainage for that trade | | C — Cosmetic | Touch-up, adjustment, cleaning within spec tolerance | Track to zero, but don't hold the project on it |
Never let a Tier A hide inside a room's list of C's — pull Tier A items to the top summary.
Assign one responsible party per item. "GC to coordinate" is not an assignee. Where trades overlap (e.g. scratched frame — painter or the trade who scratched it?), assign the most likely party and flag the dispute.
Back-charge screening. Flag as back-charge candidates: damage to completed work by another trade, rework of previously-accepted work, and items a sub was already directed to fix once. Note the evidence needed (photo, prior notice date) — a back-charge without paper is a gift.
Closeout linkage. Mark which items gate substantial completion (Tier A), which gate retainage release per trade (Tier B), and which convert to warranty items if the owner accepts occupancy first.
1. Summary — item counts by tier and by trade; the Tier A list in full. 2. Punch items by location — table per area:
| # | Location | Description of deficiency | Spec/Dwg ref | Trade | Responsible sub | Tier | Back-charge? | Acceptance criterion | Status | |---|---|---|---|---|---|---|---|---|---|
3. Trade rollups — per-sub extract with item numbers and due date field. 4. Back-charge candidates — item #, basis, evidence held/needed. 5. Closeout linkage — items gating substantial completion; items gating retainage by trade; warranty conversions.
[verify] where inferred| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 66,249 | 14,852 | -78% | 1 | 1 | 0% | 8,308 | 2,434 | -71% | 0 | 0 | — |
case-02 | fail→fail | 53,344 | 22,354 | -58% | 1 | 1 | 0% | 6,253 | 2,545 | -59% | 0 | 0 | — |
case-03 | fail→fail | 46,974 | 37,477 | -20% | 1 | 1 | 0% | 6,930 | 8,054 | +16% | 0 | 0 | — |
case-04 | pass→pass | 21,622 | 19,933 | -8% | 1 | 1 | 0% | 1,869 | 3,330 | +78% | 0 | 0 | — |
case-05 | pass→pass | 32,337 | 22,609 | -30% | 1 | 1 | 0% | 3,984 | 4,682 | +18% | 0 | 0 | — |
case-06 | pass→pass | 26,246 | 35,647 | +36% | 1 | 1 | 0% | 3,225 | 5,137 | +59% | 0 | 0 | — |
case-07 | fail→pass | 14,041 | 25,788 | +84% | 1 | 1 | 0% | 1,146 | 3,923 | +242% | 0 | 0 | — |
case-08 | fail→pass | 12,644 | 21,347 | +69% | 1 | 1 | 0% | 1,063 | 3,540 | +233% | 0 | 0 | — |
case-09 | fail→pass | 8,355 | 19,693 | +136% | 1 | 1 | 0% | 1,352 | 2,970 | +120% | 0 | 0 | — |
case-10 | fail→fail | 6,034 | 16,289 | +170% | 1 | 1 | 0% | 953 | 2,425 | +154% | 0 | 0 | — |
case-11 | pass→pass | 13,442 | 18,168 | +35% | 1 | 1 | 0% | 1,741 | 3,076 | +77% | 0 | 0 | — |
case-12 | fail→pass | 10,467 | 25,336 | +142% | 1 | 1 | 0% | 598 | 4,064 | +580% | 0 | 0 | — |
case-13 | fail→pass | 12,912 | 25,373 | +97% | 1 | 1 | 0% | 1,706 | 3,027 | +77% | 0 | 0 | — |
case-14 | pass→pass | 6,899 | 8,949 | +30% | 1 | 1 | 0% | 1,217 | 2,589 | +113% | 0 | 0 | — |
case-15 | pass→pass | 9,233 | 8,512 | -8% | 1 | 1 | 0% | 1,512 | 2,466 | +63% | 0 | 0 | — |
case-16 | pass→fail | 19,273 | 15,374 | -20% | 1 | 1 | 0% | 2,592 | 3,281 | +27% | 0 | 0 | — |
case-17 | fail→pass | 5,601 | 13,994 | +150% | 1 | 1 | 0% | 846 | 2,685 | +217% | 0 | 0 | — |
case-18 | pass→pass | 10,413 | 12,643 | +21% | 1 | 1 | 0% | 1,616 | 3,139 | +94% | 0 | 0 | — |
case-19 | fail→fail | 14,964 | 7,227 | -52% | 1 | 1 | 0% | 1,824 | 2,352 | +29% | 0 | 0 | — |
case-20 | pass→pass | 10,488 | 16,148 | +54% | 1 | 1 | 0% | 1,651 | 3,221 | +95% | 0 | 0 | — |
case-21 | fail→pass | 4,980 | 15,320 | +208% | 1 | 1 | 0% | 866 | 3,091 | +257% | 0 | 0 | — |
case-22 | fail→fail | 5,818 | 8,676 | +49% | 1 | 1 | 0% | 961 | 2,412 | +151% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.