Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Start a safe, governed autonomous loop from a plain-language goal. Use when the user says /coco-loop, "run an autonomous loop", "keep my build green for N hours", "work on X autonomously for a while", or wants a bounded self-improving loop that proposes fixes and never commits on its own. Turns a vague goal into a readable charter the user confirms, then arms and runs the loop propose-only under the coco-loops governance framework.
.claude/skills/coco-research-coco-loop/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 491% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 149% | 0% |
Turn a plain-language goal into a governed, propose-only loop. You compile the goal into a charter, the person confirms it, and the loop runs under the coco-loops framework: it proposes changes, a separate verifier judges them, and nothing is committed. The loop stops itself when its lifetime cap of active work is spent.
The framework enforces an immutable constitution underneath every charter. You cannot write a charter that disables the verifier, commits without approval, edits the constitution, sends content to raw public-cloud endpoints, or runs unattended destructive operations. The coco-loops charter command rejects any charter that tries. Do not attempt to work around it. Present the floor to the person as a feature.
coco-loops must be on the PATH. Install it from coco-research/coco-loops:
bashgit clone https://github.com/coco-research/coco-loops pip install -e coco-loops
Check with coco-loops --help. If it is missing, tell the person how to install it and stop.
Read the person's goal. Ask at most one or two clarifying questions, and only when the answer genuinely changes the charter:
so a dedicated worktree is strongly recommended (the loop reverts its own changes each cycle and refuses to run on a dirty tree, to avoid clobbering uncommitted work).
4h. The clock only advances while the loop is actually working, so a closed laptop or an idle wait does not burn it.
Do not over-question. If the goal is clear, proceed.
Get the exact field list with coco-loops charter --schema, then produce a charter as JSON. Derive the principles from the goal, and add sensible safe defaults. Keep principles readable: they are what the person confirms and what we audit the loop against.
Example, for "keep my frontend build green for the next 4 hours":
json{ "loop_id": "keep-build-green", "objective": "Make failing frontend tests pass and keep the build green.", "principles": [ "Only touch code covered by a failing test; keep the diff as small as possible.", "If a test is genuinely stale, propose updating it rather than weakening it silently.", "Prefer fixing the code over changing the test when a real regression is likely." ], "guardrails": [ "No dependency or lockfile changes.", "No CI or build-config edits." ], "context": ["the frontend test suite"], "inputs": ["failing tests detected each cycle"], "lifetime": "4h", "level": "L1" }
Rules for the charter you produce:
level is always L1. New loops are propose-only. Higher maturity is earned throughthe promotion gate, never declared here.
loop_id is a slug: lowercase letters, digits, and hyphens.lifetime from what the person asked, defaulting to 4h.Write the charter JSON to a temporary file, then run:
coco-loops charter --from-json <file>This validates the charter against the constitution floor and prints a readable version. If validation fails, it prints why (with the invariant it violated); fix the charter and try again. Show the rendered charter to the person and ask them to confirm before anything is armed.
On confirmation:
coco-loops charter --from-json <file> --installThis writes <loop_id>.contract.md into the loops directory, arms the lifetime cap, and clears the kill-switch. The loop is now defined and armed, propose-only.
The loop reads tasks from a JSON queue (a list of {"id", "name", "instruction"} objects). Seed the queue with concrete, grounded tasks, for example one task per failing test with the file and the assertion. A loop is only as good as its task list; vague tasks produce no-ops.
Run it in the chosen worktree:
coco-loops run <loop_id> --repo <worktree> --tasks <queue.json>Each cycle it pops a task, asks the maker (cursor-agent by default) to propose a change, has the verifier judge it, logs the proposal, and reverts the tree. Use --once to run a single cycle first and watch what it does. Recommend a dedicated worktree so the loop's revert can never touch the person's uncommitted work.
coco-loops status <loop_id>shows the recent proposals, the verifier verdicts, and how much of the lifetime is spent. Everything is propose-only: the person reviews the batch of proposals and applies what they want. When the lifetime is spent the loop stops itself. To run another session, re-arm with coco-loops start <loop_id> --for <duration> (or run charter --install again).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 27,392 | 9,150 | -67% | 1 | 1 | 0% | 3,514 | 2,016 | -43% | 0 | 0 | — |
case-02 | fail→fail | 19,819 | 3,133 | -84% | 1 | 1 | 0% | 2,435 | 1,649 | -32% | 0 | 0 | — |
case-03 | fail→pass | 11,102 | 25,802 | +132% | 1 | 1 | 0% | 810 | 4,786 | +491% | 0 | 0 | — |
case-04 | fail→pass | 6,455 | 7,448 | +15% | 1 | 1 | 0% | 1,117 | 1,802 | +61% | 0 | 0 | — |
case-05 | fail→pass | 11,638 | 8,470 | -27% | 1 | 1 | 0% | 924 | 1,912 | +107% | 0 | 0 | — |
case-06 | fail→pass | 23,352 | 7,475 | -68% | 1 | 1 | 0% | 1,545 | 1,659 | +7% | 0 | 0 | — |
case-07 | fail→pass | 11,771 | 13,523 | +15% | 1 | 1 | 0% | 1,124 | 2,794 | +149% | 0 | 0 | — |
case-08 | fail→pass | 14,324 | 5,256 | -63% | 1 | 1 | 0% | 1,770 | 2,300 | +30% | 0 | 0 | — |
case-09 | fail→pass | 11,494 | 7,306 | -36% | 1 | 1 | 0% | 2,051 | 2,675 | +30% | 0 | 0 | — |
case-10 | fail→pass | 49,568 | 2,533 | -95% | 1 | 1 | 0% | 8,058 | 1,750 | -78% | 0 | 0 | — |
case-11 | fail→fail | 21,652 | 7,621 | -65% | 1 | 1 | 0% | 3,054 | 1,767 | -42% | 0 | 0 | — |
case-12 | fail→pass | 16,485 | 8,494 | -48% | 1 | 1 | 0% | 2,593 | 2,925 | +13% | 0 | 0 | — |
case-13 | fail→pass | 26,372 | 8,014 | -70% | 1 | 1 | 0% | 4,050 | 1,890 | -53% | 0 | 0 | — |
case-14 | fail→pass | 12,956 | 7,232 | -44% | 1 | 1 | 0% | 1,868 | 1,621 | -13% | 0 | 0 | — |
case-15 | fail→pass | 28,912 | 8,483 | -71% | 1 | 1 | 0% | 3,751 | 1,797 | -52% | 0 | 0 | — |
case-16 | fail→fail | 50,490 | 14,779 | -71% | 1 | 1 | 0% | 1,483 | 2,888 | +95% | 0 | 0 | — |
case-22 | fail→fail | 16,615 | 20,384 | +23% | 1 | 1 | 0% | 1,868 | 2,125 | +14% | 0 | 0 | — |
case-17 | pass→fail | 13,708 | 12,281 | -10% | 1 | 1 | 0% | 2,154 | 2,482 | +15% | 0 | 0 | — |
case-18 | pass→pass | 11,426 | 5,953 | -48% | 1 | 1 | 0% | 1,705 | 2,255 | +32% | 0 | 0 | — |
case-19 | fail→pass | 16,770 | 8,209 | -51% | 1 | 1 | 0% | 1,616 | 1,863 | +15% | 0 | 0 | — |
case-20 | pass→pass | 20,562 | 16,851 | -18% | 1 | 1 | 0% | 2,202 | 3,035 | +38% | 0 | 0 | — |
case-21 | fail→pass | 29,499 | 10,543 | -64% | 1 | 1 | 0% | 1,871 | 3,241 | +73% | 0 | 0 | — |
case-23 | pass→pass | 7,716 | 9,935 | +29% | 1 | 1 | 0% | 1,436 | 2,267 | +58% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +57 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.