Install any skill in seconds. Free to start, no credit card required.
Get Started Free →ISO 42001 AI Management System (AIMS) audit-prep playbook. Use when an ISO 42001 certification audit is scheduled (Stage 1 or Stage 2), when a surveillance or internal AIMS audit is due, or when preparing an AI Impact Assessment (AIIA).
.claude/skills/borghei-aims-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 67% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-24 | ✗→✓ | ▲ Improved | 26% | 0% |
Operational playbook for ISO 42001:2023 AI Management System (AIMS) audit preparation. Whether targeting initial certification, surveillance audit, or annual internal audit.
When to use this skill vs. iso42001-ai-management:
| Situation | Skill applies | |-----------|---------------| | ISO 42001 Stage 1 audit scheduled | Yes — documentation review prep | | Stage 2 (onsite/operational) audit | Yes — operational evidence sprint | | Annual surveillance audit | Yes — surveillance prep | | Internal AIMS audit | Yes — internal audit playbook | | AI Impact Assessment for new system | Yes — scripts/ai_impact_assessment_checker.py | | Building AIMS from scratch | Use ra-qm-team/iso42001-ai-management |
Week 1: Internal audit + gap analysis
Week 2: Remediation + AI inventory refresh
Week 3: Documentation review + walkthrough rehearsal
Week 4: Auditor onsiteWeeks 1-3: AIMS documentation completion (Annex A controls coverage)
Weeks 4-5: Stage 1 audit + gap closure
Weeks 6-7: Stage 2 operational evidence prep + mock walkthroughs
Week 8: Stage 2 audit| Clause | Topic | |--------|-------| | 4 | Context of the organization | | 5 | Leadership | | 6 | Planning (including AI risk + AI objectives) | | 7 | Support (resources, competence, awareness, communication, documentation) | | 8 | Operation (AI lifecycle, supplier relationships) | | 9 | Performance evaluation (monitoring, internal audit, management review) | | 10 | Improvement (nonconformity, continual improvement) |
| Annex A area | Topics | |--------------|--------| | A.2 | Policies related to AI | | A.3 | Internal organization | | A.4 | Resources for AI systems | | A.5 | Assessing impacts of AI systems | | A.6 | AI system lifecycle | | A.7 | Data for AI systems | | A.8 | Information for interested parties | | A.9 | Use of AI systems | | A.10 | Third-party relationships |
| Item | Evidence | Common gap | |------|----------|------------| | Complete AI system inventory | Inventory document | Shadow AI not captured | | AI Impact Assessment (AIIA) per system | Per-system AIIA | Skipped for "low-risk" systems | | AIIA reviewed periodically | Review records | One-time only | | Risk classification of systems | Per system | Not documented |
| Item | Evidence | Common gap | |------|----------|------------| | AI policy approved + dated | Signed policy | Not signed / stale | | AI ethics principles | Documented principles | Generic; not actionable | | AI governance body | Charter / minutes | Not formalized | | Roles + responsibilities | RACI | Not defined |
| Item | Evidence | Common gap | |------|----------|------------| | AI development lifecycle defined | Process documentation | Not formalized | | Data quality controls | Per system | Generic only | | Model validation procedures | Per system | Validation skipped | | AI system testing | Per system | Inadequate testing | | Deployment controls | Per system | No controls | | Operational monitoring | Per system | Drift not monitored | | Decommissioning procedures | Per system | Not defined |
| Item | Evidence | Common gap | |------|----------|------------| | Data sources documented | Per system | Vague | | Data quality assessed | Quality metrics | Not measured | | Data lineage tracked | Documentation | Untracked | | Sensitive data protection | Controls | Insufficient |
| Item | Evidence | Common gap | |------|----------|------------| | Third-party AI inventory | List | Incomplete | | Vendor due diligence for AI | Per vendor | Generic IT only | | Contract terms for AI vendors | AI-specific clauses | Standard MSA only |
Before running the audit-prep, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the readiness assessment.
python3 scripts/aims_readiness_score.py --config aims-controls.yamlpython3 scripts/ai_impact_assessment_checker.py --aiia system-aiia.yaml| Script | Purpose | |--------|---------| | scripts/aims_readiness_score.py | Score AIMS readiness per clause + Annex A area | | scripts/ai_impact_assessment_checker.py | Validate AI Impact Assessment completeness |
ra-qm-team/iso42001-ai-management — deep ISO 42001 AIMS program managementra-qm-team/eu-ai-act-specialist — EU AI Act regulatory companionra-qm-team/audit-prep/ai-act-readiness — AI Act audit-prep variantra-qm-team/audit-prep/compliance-readiness — multi-framework readiness| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 22,362 | 23,073 | +3% | 1 | 1 | 0% | 3,560 | 5,550 | +56% | 0 | 0 | — |
case-02 | fail→fail | 31,474 | 24,560 | -22% | 1 | 1 | 0% | 4,960 | 5,643 | +14% | 0 | 0 | — |
case-03 | pass→pass | 22,717 | 21,354 | -6% | 1 | 1 | 0% | 3,540 | 5,496 | +55% | 0 | 0 | — |
case-04 | fail→fail | 35,732 | 25,986 | -27% | 1 | 1 | 0% | 5,864 | 6,179 | +5% | 0 | 0 | — |
case-05 | fail→fail | 19,048 | 23,016 | +21% | 1 | 1 | 0% | 2,977 | 5,543 | +86% | 0 | 0 | — |
case-06 | fail→pass | 15,009 | 19,054 | +27% | 1 | 1 | 0% | 2,472 | 4,750 | +92% | 0 | 0 | — |
case-07 | fail→pass | 18,477 | 13,504 | -27% | 1 | 1 | 0% | 2,800 | 4,245 | +52% | 0 | 0 | — |
case-08 | fail→pass | 10,337 | 3,030 | -71% | 1 | 1 | 0% | 1,476 | 2,470 | +67% | 0 | 0 | — |
case-09 | fail→pass | 10,183 | 2,317 | -77% | 1 | 1 | 0% | 1,517 | 2,296 | +51% | 0 | 0 | — |
case-10 | pass→pass | 10,828 | 6,904 | -36% | 1 | 1 | 0% | 1,885 | 3,141 | +67% | 0 | 0 | — |
case-11 | pass→pass | 8,288 | 6,316 | -24% | 1 | 1 | 0% | 1,448 | 2,956 | +104% | 0 | 0 | — |
case-12 | pass→pass | 16,736 | 13,391 | -20% | 1 | 1 | 0% | 2,683 | 3,838 | +43% | 0 | 0 | — |
case-13 | pass→pass | 7,161 | 6,074 | -15% | 1 | 1 | 0% | 1,140 | 2,855 | +150% | 0 | 0 | — |
case-14 | pass→pass | 17,682 | 11,664 | -34% | 1 | 1 | 0% | 2,651 | 3,615 | +36% | 0 | 0 | — |
case-15 | pass→pass | 13,959 | 12,326 | -12% | 1 | 1 | 0% | 2,039 | 3,795 | +86% | 0 | 0 | — |
case-16 | pass→pass | 4,568 | 4,234 | -7% | 1 | 1 | 0% | 777 | 2,628 | +238% | 0 | 0 | — |
case-17 | pass→pass | 3,631 | 4,015 | +11% | 1 | 1 | 0% | 649 | 2,626 | +305% | 0 | 0 | — |
case-18 | pass→pass | 17,380 | 9,336 | -46% | 1 | 1 | 0% | 2,886 | 3,428 | +19% | 0 | 0 | — |
case-19 | pass→pass | 12,756 | 9,387 | -26% | 1 | 1 | 0% | 2,091 | 3,436 | +64% | 0 | 0 | — |
case-20 | pass→pass | 16,540 | 11,175 | -32% | 1 | 1 | 0% | 2,692 | 3,584 | +33% | 0 | 0 | — |
case-21 | pass→pass | 16,543 | 13,269 | -20% | 1 | 1 | 0% | 2,584 | 3,843 | +49% | 0 | 0 | — |
case-22 | pass→pass | 13,617 | 11,888 | -13% | 1 | 1 | 0% | 1,905 | 3,646 | +91% | 0 | 0 | — |
case-23 | pass→pass | 14,650 | 10,453 | -29% | 1 | 1 | 0% | 2,252 | 3,471 | +54% | 0 | 0 | — |
case-24 | fail→pass | 18,129 | 10,336 | -43% | 1 | 1 | 0% | 2,886 | 3,637 | +26% | 0 | 0 | — |
case-25 | pass→pass | 5,383 | 3,525 | -35% | 1 | 1 | 0% | 899 | 2,535 | +182% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +20 percentage points is the difference between those two pass rates over the 25 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.