Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Privileged applier that LANDS meta-optimize / corpus-audit patches the user approved — the ONLY skill permitted to mutate the skill corpus from a self-modification proposal, with cross-model jury and human approval at landing. Use when the user says "meta apply", "/meta-apply", "land the staged patches", "应用优化", after a /meta-optimize run.
.claude/skills/wanshuiyin-meta-apply/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 154% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 159% | 0% |
> Codex assurance: a base landing review records > review_independence: same-family and acceptance_status: provisional via > stamp-provisional; the artifact is not auto-curatable. Only an overlay or > deterministic verifier may produce accepted authorization.
This skill exists to enforce a privilege boundary. Producers like /meta-optimize (and, later, corpus-audit) are read-only — no Write/Edit, no apply step; they can only stage candidate patches under .aris/meta/pending/. This skill is the only place a staged patch becomes a real change to the corpus. Splitting "propose" from "land" across two skills with different tool grants is what makes "a loop cannot apply its own patch" structural rather than a sentence the producer is asked to obey.
It is human-invoked only. It runs when the user explicitly types /meta-apply after reading the producer's REPORT. That invocation IS the landing authorization (the locked 2026-05-30 decision: the human stays at the landing gate).
The earlier design trusted a jury_verdict: pass field in the producer's manifest. That is forgeable — the producer writes that field, so a buggy/misaligned producer could stage a killed patch labelled "pass." So this skill does not trust any producer-written verdict. It runs the fresh landing reviewer itself, at landing time, on the actual staged diff (fresh reviewer, paths-only). The verdict is produced where it is consumed, by the privileged human-invoked skill — nothing to forge.
For each staged patch the user asks to land, in order — any failure ⇒ skip & report, never silently apply:
/meta-apply 1,3or all); default to applying nothing.
gpt-6-astrareviewer via spawn_agent (reasoning_effort: ultra, read-only, paths-only per reviewer-independence.md) on the staged .diff + its target. Ask: does this change improve the harness without regressions; PASS or KILL + one-line reason. Include the scope-limits block from review-scope-limits.md: this jury judges ARIS's own mechanism, so an over-defensive KILL permanently blocks a good patch. The block bans proposing new hash binding — it is not a reason to KILL a patch that touches the existing provenance stamp. KILL ⇒ refuse. The human cannot override a KILL — they may only pick among reviewer-PASSED survivors.
lands only with stamp-provisional; it can complete this explicit human-invoked operation but does not authorize future auto-curation. A Claude/Gemini overlay or deterministic verifier uses strict stamp and may record accepted. See skill-governance.md.
bashPENDING=".aris/meta/pending" [ -d "$PENDING" ] || { echo "Nothing staged. Run /meta-optimize first."; exit 0; } echo "Staged:"; cat "$PENDING/manifest.jsonl"
Resolve provenance.py through the Codex manifest:
bashif [ -z "${ARIS_REPO:-}" ] && [ -f .aris/installed-skills-codex.txt ]; then ARIS_REPO=$(awk -F'\t' '$1=="repo_root"{print $2; exit}' .aris/installed-skills-codex.txt 2>/dev/null) || true fi PROVENANCE="" [ -n "${ARIS_REPO:-}" ] && [ -f "$ARIS_REPO/tools/provenance.py" ] && PROVENANCE="$ARIS_REPO/tools/provenance.py" [ -z "$PROVENANCE" ] && [ -f tools/provenance.py ] && PROVENANCE="tools/provenance.py" [ -n "$PROVENANCE" ] || { echo "ERROR: provenance.py unresolved" >&2; exit 1; }
For every patch the user asked to land, read its staged .diff and target, then spawn the fresh reviewer jury (Rule 2) — paths-only, no producer reasoning, no prior-round context. Record {patch, jury_verdict, jury_review_id, one_line_reason}. Print a one-line result per patch (PASS → eligible / KILL → refused: <reason>).
> The producer may have written an advisory pre-screen into the manifest to help the > human read the REPORT — ignore it for the landing decision. Only this fresh verdict > counts.
For each patch that PASSED Step 1 and was named by the user:
.aris/meta/backups/<date>/<target> (use the Write toolto copy contents; corpus paths are not Bash-writable when corpus_write_guard is active — and the applier should use Write/Edit for corpus mutation anyway).
bash python3 "$PROVENANCE" stamp-provisional "$TARGET" --author "$AUTHOR" \ --reviewer "$JURY_MODEL" --verdict-id "$JURY_REVIEW_ID" This records review_independence: same-family and acceptance_status: provisional; is_auto_curatable remains false. If the active overlay produced a cross-family result, use strict stamp instead.
.aris/meta/optimizations.jsonl:{ts, patch, target, author_model, reviewer_model, jury_review_id, applied: true}.
Per patch: LANDED <target> (+ backup path + provenance sidecar) or REFUSED <patch>: <reason>. Remove landed patches from .aris/meta/pending/. Remind the user a landed patch is revertable from its backup, and to test the changed skill next run.
A stamp records that a change passed a process (fresh landing review + human landing), not that it is correct. To prevent "approved-but-wrong with a stamp that vouches for it" (false-authority laundering — worse than no stamp, because a later auto-curator reads it as evidence):
verdict_id (auditable review) + content_hash (a later hand-editinvalidates it).
artifacts, and a behavioral auditor that REVOKES a stamp when a landed skill misbehaves. Track as follow-up; never treat a stamp as permanent truth.
on the staged diff; never trust a producer-written verdict; the human picks among survivors, never resurrects a KILL.
stamp-provisional; only an overlay or deterministic verifier may use strict stamp.
corpus_write_guard hook (if installed) additionally denies Bash corpus writes — it does NOT gate Write/Edit, so it does not by itself stop this skill from editing the corpus; the jury-at-landing + stamp discipline above is what governs Write/Edit mutations (that discipline is procedure, not a hook-enforced mechanism).
.aris/meta/pending/;invents nothing of its own.
Save each landing-jury reviewer call's trace per review-tracing.md to .aris/traces/meta-apply/<date>_run<NN>/ — the acquittal that landed a corpus change must be forensically recoverable.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→pass | 13,001 | 8,169 | -37% | 1 | 1 | 0% | 1,026 | 2,604 | +154% | 0 | 0 | — |
case-06 | fail→pass | 48,842 | 8,480 | -83% | 1 | 1 | 0% | 2,399 | 2,714 | +13% | 0 | 0 | — |
case-01 | fail→fail | 14,739 | 20,342 | +38% | 1 | 1 | 0% | 208 | 2,954 | +1320% | 0 | 0 | — |
case-02 | fail→fail | 14,727 | 16,940 | +15% | 1 | 1 | 0% | 265 | 2,550 | +862% | 0 | 0 | — |
case-03 | fail→fail | 14,888 | 18,536 | +25% | 1 | 1 | 0% | 1,710 | 2,713 | +59% | 0 | 0 | — |
case-04 | fail→fail | 16,100 | 9,484 | -41% | 1 | 1 | 0% | 1,761 | 3,003 | +71% | 0 | 0 | — |
case-07 | fail→fail | 15,617 | 10,273 | -34% | 1 | 1 | 0% | 1,631 | 3,140 | +93% | 0 | 0 | — |
case-08 | fail→pass | 13,211 | 7,563 | -43% | 1 | 1 | 0% | 1,343 | 2,487 | +85% | 0 | 0 | — |
case-09 | pass→pass | 16,328 | 10,537 | -35% | 1 | 1 | 0% | 1,621 | 3,015 | +86% | 0 | 0 | — |
case-10 | fail→pass | 15,060 | 6,886 | -54% | 1 | 1 | 0% | 1,562 | 2,427 | +55% | 0 | 0 | — |
case-11 | fail→fail | 18,513 | 10,640 | -43% | 1 | 1 | 0% | 2,075 | 3,020 | +46% | 0 | 0 | — |
case-12 | fail→pass | 28,961 | 9,484 | -67% | 1 | 1 | 0% | 1,144 | 2,961 | +159% | 0 | 0 | — |
case-13 | fail→pass | 15,163 | 8,789 | -42% | 1 | 1 | 0% | 1,634 | 2,762 | +69% | 0 | 0 | — |
case-14 | fail→pass | 19,484 | 8,150 | -58% | 1 | 1 | 0% | 2,579 | 2,712 | +5% | 0 | 0 | — |
case-15 | pass→pass | 12,020 | 8,223 | -32% | 1 | 1 | 0% | 1,040 | 2,645 | +154% | 0 | 0 | — |
case-16 | fail→pass | 18,579 | 7,410 | -60% | 1 | 1 | 0% | 1,984 | 2,476 | +25% | 0 | 0 | — |
case-17 | pass→pass | 16,209 | 10,255 | -37% | 1 | 1 | 0% | 1,631 | 2,957 | +81% | 0 | 0 | — |
case-18 | pass→pass | 15,097 | 11,646 | -23% | 1 | 1 | 0% | 1,589 | 3,040 | +91% | 0 | 0 | — |
case-19 | pass→pass | 18,600 | 10,210 | -45% | 1 | 1 | 0% | 1,855 | 2,824 | +52% | 0 | 0 | — |
case-20 | fail→fail | 15,009 | 19,214 | +28% | 1 | 1 | 0% | 231 | 2,671 | +1056% | 0 | 0 | — |
case-21 | fail→fail | 9,491 | 16,440 | +73% | 1 | 1 | 0% | 719 | 2,473 | +244% | 0 | 0 | — |
case-22 | fail→fail | 72,284 | 37,143 | -49% | 1 | 1 | 0% | 6,706 | 8,359 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 15 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/17/2026 | +45% |
| gemini-3.6-flash | verified | 8/11/2026 | +52% |
Other measured skills in the registry, with their headline benchmark lift.