Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when writing up and releasing the results of a psychological audit of an AI/ML personnel assessment — producing a precise, comprehensive technical report for testing professionals AND a layperson-friendly summary for those the predictions affect, establishing the auditor's standards and credibility in the report, and deciding on public release. Triggers: "write the AI audit report", "release the bias audit results", "dual-audience audit report", "should we publish the audit", "auditor credib
.claude/skills/openmatter-network-ai-audit-reporting/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 41% | 0% |
An audit only creates value if its results are communicated to the right audiences and, where the public interest is at stake, released. Reporting is also where the auditor's own credibility is established or lost.
Present results in multiple formats to meet the needs of all relevant audiences. At minimum:
enough that another competent professional could evaluate and, ideally, reproduce the audit's reasoning. Mirror the structure of a technical-validation-report, but organized around the 12 components and the claims evaluated.
(e.g., candidates) — clear, accurate, non-technical, addressing the fairness concerns in the terms those audiences actually use (justice/transparency — see ai-fairness-lenses).
One format cannot serve both audiences; write both.
validity / utility / lack-of-bias (from ai-audit-planning).
ai-fairness-lenses), statedso conclusions are interpretable across disciplines.
with evidence, gaps, and access limitations encountered.
audits) to improve the model.
intersectional subgroups, etc.).
No auditor or audit is automatically credible. "This system has been audited and is therefore credible" deserves skepticism. So the report must let readers judge the audit itself by disclosing:
in due course across the components;
discretion (the existence and scope of withholding should itself be disclosed).
Access and documentation vary even within auditor type, so transparency about access is essential to interpreting the findings.
Unless there is a compelling, transparently stated reason not to, an audit whose results are in the public interest should be released. Organizations may choose otherwise, but doing so risks the credibility of the audit, the company that built the algorithm, and the auditors. Public-facing, transparent, open audits are especially warranted when a system has outsized societal impact.
Normalize routine auditing. Treat regular internal and external auditing as a public good that raises the probability algorithmic systems in general are valid, valuable, and fair — and that builds public trust. Formative auditing folded into development, with complete documentation, can diminish or even preclude the need for post-hoc audits.
can't evaluate it).
ai-audit-planning (audience & release policy set up front) · ai-fairness-lenses · all model/stakeholder/meta audit skills · technical-validation-report (structure parallel)
Source: Landers & Behrend (2023), "Designing an Effective Psychological Audit" — multiple-format reporting, releasing results in the public interest, normalizing routine auditing, and auditor credibility.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | pass→pass | 12,184 | 9,838 | -19% | 1 | 1 | 0% | 1,842 | 2,669 | +45% | 0 | 0 | — |
case-01 | fail→pass | 36,301 | 24,577 | -32% | 1 | 1 | 0% | 6,235 | 5,209 | -16% | 0 | 0 | — |
case-02 | fail→pass | 34,429 | 33,686 | -2% | 1 | 1 | 0% | 6,230 | 6,043 | -3% | 0 | 0 | — |
case-03 | fail→pass | 36,038 | 25,861 | -28% | 1 | 1 | 0% | 5,889 | 5,403 | -8% | 0 | 0 | — |
case-04 | fail→pass | 24,108 | 14,226 | -41% | 1 | 1 | 0% | 3,684 | 3,397 | -8% | 0 | 0 | — |
case-05 | pass→pass | 9,632 | 5,538 | -43% | 1 | 1 | 0% | 1,603 | 2,130 | +33% | 0 | 0 | — |
case-06 | pass→pass | 12,225 | 10,633 | -13% | 1 | 1 | 0% | 1,587 | 2,686 | +69% | 0 | 0 | — |
case-07 | pass→pass | 10,141 | 10,637 | +5% | 1 | 1 | 0% | 1,436 | 2,745 | +91% | 0 | 0 | — |
case-08 | pass→pass | 12,822 | 9,064 | -29% | 1 | 1 | 0% | 2,248 | 2,353 | +5% | 0 | 0 | — |
case-09 | pass→pass | 15,331 | 9,591 | -37% | 1 | 1 | 0% | 2,223 | 2,623 | +18% | 0 | 0 | — |
case-10 | pass→pass | 11,768 | 8,534 | -27% | 1 | 1 | 0% | 1,894 | 2,305 | +22% | 0 | 0 | — |
case-11 | pass→pass | 15,670 | 6,737 | -57% | 1 | 1 | 0% | 1,982 | 2,086 | +5% | 0 | 0 | — |
case-12 | pass→pass | 11,669 | 8,780 | -25% | 1 | 1 | 0% | 1,637 | 2,397 | +46% | 0 | 0 | — |
case-13 | fail→pass | 11,900 | 11,309 | -5% | 1 | 1 | 0% | 1,809 | 2,549 | +41% | 0 | 0 | — |
case-15 | pass→pass | 10,856 | 6,412 | -41% | 1 | 1 | 0% | 1,685 | 2,051 | +22% | 0 | 0 | — |
case-16 | pass→pass | 14,147 | 9,679 | -32% | 1 | 1 | 0% | 2,201 | 2,650 | +20% | 0 | 0 | — |
case-17 | fail→fail | 10,687 | 13,325 | +25% | 1 | 1 | 0% | 1,992 | 3,123 | +57% | 0 | 0 | — |
case-18 | fail→pass | 11,240 | 7,252 | -35% | 1 | 1 | 0% | 1,841 | 2,297 | +25% | 0 | 0 | — |
case-19 | pass→pass | 11,808 | 8,554 | -28% | 1 | 1 | 0% | 1,840 | 2,344 | +27% | 0 | 0 | — |
case-20 | pass→fail | 14,809 | 13,878 | -6% | 1 | 1 | 0% | 2,364 | 3,340 | +41% | 0 | 0 | — |
case-21 | fail→pass | 15,508 | 11,989 | -23% | 1 | 1 | 0% | 2,396 | 2,954 | +23% | 0 | 0 | — |
case-22 | pass→pass | 11,523 | 13,050 | +13% | 1 | 1 | 0% | 2,486 | 3,735 | +50% | 0 | 0 | — |
case-23 | pass→pass | 16,109 | 17,669 | +10% | 1 | 1 | 0% | 2,721 | 4,164 | +53% | 0 | 0 | — |
case-24 | pass→pass | 16,231 | 17,911 | +10% | 1 | 1 | 0% | 2,981 | 4,209 | +41% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +25 percentage points is the difference between those two pass rates over the 24 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.