Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audit one research Workspace for declared Unit outputs and Pipeline target Artifacts, writing `output/CONTRACT_REPORT.md`; use for mid-Run coverage snapshots or final delivery completeness, not deep provenance integrity.
.claude/skills/willoscar-artifact-contract-auditor/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 34% | 0% |
This Skill measures coverage: whether files promised by Units and the locked Pipeline exist at the point they are required. It does not replace the deeper Attempt, Manifest, Artifact-hash, and ledger checks in pipeline.py audit.
UNITS.csv.PIPELINE.lock.md.pipelines/*.pipeline.md target-Artifact contract.output/CONTRACT_REPORT.md.Read PIPELINE.lock.md, resolve the Pipeline inside this repository, and parse UNITS.csv. Treat an invalid lock or malformed Unit table as a reportable contract failure rather than guessing another Workflow.
Completion criterion: one Pipeline contract and one readable Unit table are bound to the audit.
For every DONE Unit, verify each required output exists. Outputs prefixed with ? are optional. A missing required output is Unit-level contract drift even when the Pipeline is still running.
Completion criterion: every DONE Unit is classified as output-complete or is listed with each missing required path.
Determine whether all Units are terminal (DONE or SKIP). Only then require every non-optional target Artifact declared by the Pipeline. During a partial Run, report missing final targets as expected rather than failures.
Completion criterion: final-target completeness is evaluated against the actual Run phase, not merely file absence.
Always write output/CONTRACT_REPORT.md with one status:
PASS: terminal Run, complete Unit outputs, complete Pipeline targets.OK: partial Run, consistent completed Unit outputs.FAIL: missing output from a DONE Unit, or missing final target from aterminal Run.
Completion criterion: the report names the evaluated Pipeline, Run phase, status, and every blocking missing path.
For missing Unit outputs, reopen or rerun the owning Unit through the Pipeline adapter. For missing final targets, identify the owning Unit or Skill. When file coverage passes, use scripts/pipeline.py audit for provenance integrity.
Completion criterion: every FAIL item has an owner and repair route; a PASS claim is explicitly limited to declared file coverage.
UNITS.csv is the only source for Unit output obligations.Decisions, Failures, and Evaluation consistency.
bashuv run python .codex/skills/artifact-contract-auditor/scripts/run.py \ --workspace workspaces/<name>
--workspace <path>: Workspace to audit.--unit-id <U###>, --inputs, --outputs, --checkpoint: optionalPipeline-runner compatibility arguments.
Inspect the interface:
bashuv run python .codex/skills/artifact-contract-auditor/scripts/run.py --help
Run the deeper provenance audit after coverage passes:
bashuv run python scripts/pipeline.py audit --workspace workspaces/<name>
The audit may create or replace output/CONTRACT_REPORT.md. It must not edit research content, Unit statuses, checkpoint approvals, or source Artifacts.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→fail | 6,340 | 5,456 | -14% | 1 | 1 | 0% | 215 | 1,135 | +428% | 0 | 0 | — |
case-03 | fail→fail | 5,151 | 4,712 | -9% | 1 | 1 | 0% | 246 | 966 | +293% | 0 | 0 | — |
case-01 | fail→fail | 2,648 | 5,161 | +95% | 1 | 1 | 0% | 265 | 1,044 | +294% | 0 | 0 | — |
case-04 | fail→fail | 13,048 | 6,409 | -51% | 1 | 1 | 0% | 2,505 | 1,082 | -57% | 0 | 0 | — |
case-05 | fail→fail | 4,372 | 4,973 | +14% | 1 | 1 | 0% | 153 | 1,079 | +605% | 0 | 0 | — |
case-06 | fail→fail | 5,050 | 4,507 | -11% | 1 | 1 | 0% | 243 | 954 | +293% | 0 | 0 | — |
case-07 | fail→pass | 5,590 | 3,592 | -36% | 1 | 1 | 0% | 876 | 1,454 | +66% | 0 | 0 | — |
case-08 | pass→pass | 9,224 | 2,925 | -68% | 1 | 1 | 0% | 1,457 | 1,302 | -11% | 0 | 0 | — |
case-09 | pass→pass | 8,714 | 3,722 | -57% | 1 | 1 | 0% | 1,308 | 1,443 | +10% | 0 | 0 | — |
case-10 | pass→pass | 6,237 | 2,440 | -61% | 1 | 1 | 0% | 962 | 1,231 | +28% | 0 | 0 | — |
case-11 | pass→pass | 3,551 | 2,809 | -21% | 1 | 1 | 0% | 589 | 1,285 | +118% | 0 | 0 | — |
case-12 | pass→pass | 7,451 | 2,550 | -66% | 1 | 1 | 0% | 1,159 | 1,203 | +4% | 0 | 0 | — |
case-13 | pass→pass | 8,792 | 3,632 | -59% | 1 | 1 | 0% | 932 | 1,429 | +53% | 0 | 0 | — |
case-14 | fail→pass | 21,939 | 2,603 | -88% | 1 | 1 | 0% | 1,613 | 1,185 | -27% | 0 | 0 | — |
case-15 | fail→fail | 7,643 | 4,299 | -44% | 1 | 1 | 0% | 1,140 | 1,545 | +36% | 0 | 0 | — |
case-16 | pass→pass | 11,043 | 1,695 | -85% | 1 | 1 | 0% | 1,667 | 1,028 | -38% | 0 | 0 | — |
case-17 | fail→pass | 8,103 | 2,768 | -66% | 1 | 1 | 0% | 1,215 | 1,254 | +3% | 0 | 0 | — |
case-18 | fail→pass | 7,360 | 1,511 | -79% | 1 | 1 | 0% | 1,193 | 1,013 | -15% | 0 | 0 | — |
case-19 | pass→pass | 4,226 | 3,279 | -22% | 1 | 1 | 0% | 710 | 1,281 | +80% | 0 | 0 | — |
case-20 | pass→pass | 5,898 | 4,195 | -29% | 1 | 1 | 0% | 823 | 1,481 | +80% | 0 | 0 | — |
case-21 | fail→pass | 6,105 | 2,724 | -55% | 1 | 1 | 0% | 894 | 1,197 | +34% | 0 | 0 | — |
case-22 | fail→pass | 12,273 | 3,400 | -72% | 1 | 1 | 0% | 1,749 | 1,301 | -26% | 0 | 0 | — |
case-23 | fail→pass | 13,060 | 1,551 | -88% | 1 | 1 | 0% | 1,901 | 1,006 | -47% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +30 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.