Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run the explicitly requested independent review and refinement cycle for OpenDesign HTML through the native open-design-review loop. Use for a new brief or existing docs/design artifacts; ordinary edits and _uiux.md input alone do not trigger it.
.claude/skills/compozy-open-design-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -48% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -14% | 0% |
Use this skill only for an explicit request for the independent review and refinement cycle. The open-design skill owns the shared craft, project authority, and output contract. Ordinary design work stays with open-design-designer.
Inspect open-design-review through compozy__loop_inspect and read the live descriptors for loop operations. Use the active workspace; the open-design profile provides the designer, critic, and linter. If operations are unavailable, report that limitation without claiming an independent cycle ran.
Supply a concise brief and a workspace-relative artifact_path under docs/design/. Preserve existing or explicitly requested paths; otherwise use docs/design/<slug>/index.html. For related HTMLs, name the entry file and list the other boards or _uiux.md paths in the brief. The designer returns every HTML path for lint and review.
Pass the following config_overrides on both calls. CompozyOS runtime defaults override definition values, so the definition alone does not enforce this workflow's limit. full_body also reruns the complete workflow after an action failure. A rejected review always starts a fresh generation.
Dry-run with compozy__loop_run using dry: true, verify effective iteration_cap: 3 and reattempt_strategy: full_body, then execute with the same inputs and overrides without dry:
json{ "name": "open-design-review", "workspace": "<active-workspace-id>", "config_overrides": { "iteration_cap": 3, "reattempt_strategy": "full_body" }, "inputs": { "artifact_path": "docs/design/session-actions/index.html", "brief": "Review and refine the existing session-actions board against the requested selection, confirmation, empty, and failure states. Preserve the product styling and scope." } }
Optional designer and critic inputs select existing agent definitions. Provider and model selection follow CompozyOS's runtime settings. Do not create a new loop definition, specification, or registry to run this workflow.
Each pass runs designer → native lint → independent critic → native lint verification. The designer refines the same files; the linter reads every returned HTML; the critic evaluates craft, purpose and states, brand, accessibility, and copy. The critic runs as a separate agent session with filesystem and browser tools, owns judgment, and returns a schema-validated verdict with evidence for every file. The designer owns corrections and receives the previous critic verdict, blockers, evidence, and exceptions on the next generation.
The linter must run successfully. Its findings are heuristic: inspect each one, require applicable fixes, and document a specific source-backed exception where appropriate. A lint passed value only reports its P0 result; it does not establish visual quality or complete review.
Completion requires an explicit approved verdict, no blocking issues, evidence for every reviewed file, and unchanged paths and lint digests before and after review. Each review identifies its file by lint_index, path, and digest. Its exceptions object must cover exactly the remaining lint IDs, with a specific source and reason for each; a clean file uses {}. IDs are unique within each artifact, and the lint remains the severity authority. The completion condition verifies this coverage and file identity without copying finding prose. The native completion condition fails on evaluation errors; missing or invalid critic output cannot become approval.
Use open-design-browser for agent-browser inspection when rendered evidence is useful and available. Load the current screenshot through an image-capable harness before making visual claims; text snapshots alone are insufficient. Missing screenshots are not an automatic rejection, but required unavailable evidence remains a disclosed limitation.
With these per-run overrides, the native loop permits at most three generations: the initial pass and up to two refinements, stopping early on approval. Generations repeated after action failures consume this limit; individual node retries follow the native retry policy. A designer or critic inside the loop must return its requested schema and never start another run. Report transport or dependency failures as failures.
Follow the returned run ID with compozy__loop_status and inspect its node outputs, especially critique and verify. Use compozy loop why <run-id> -o json when a result needs diagnosis. Keep observing the existing run after a wait timeout rather than starting it again.
Successful completion requires terminal done and the recorded approved critique result with matching lint verification. Link the actual HTML files, report the terminal outcome and remaining findings or accepted exceptions, and state which browser or visual checks ran. Exhaustion or interruption leaves the latest files on disk; there is no automatic rollback or best-version restoration.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,330 | 8,378 | +32% | 1 | 1 | 0% | 319 | 1,394 | +337% | 0 | 0 | — |
case-02 | fail→fail | 5,959 | 6,179 | +4% | 1 | 1 | 0% | 289 | 1,313 | +354% | 0 | 0 | — |
case-03 | fail→fail | 5,824 | 7,209 | +24% | 1 | 1 | 0% | 295 | 1,368 | +364% | 0 | 0 | — |
case-04 | fail→pass | 13,198 | 7,977 | -40% | 1 | 1 | 0% | 2,505 | 2,613 | +4% | 0 | 0 | — |
case-05 | pass→pass | 8,099 | 7,172 | -11% | 1 | 1 | 0% | 1,553 | 2,030 | +31% | 0 | 0 | — |
case-06 | pass→pass | 9,248 | 4,102 | -56% | 1 | 1 | 0% | 1,031 | 1,683 | +63% | 0 | 0 | — |
case-07 | fail→pass | 5,031 | 9,427 | +87% | 1 | 1 | 0% | 730 | 1,539 | +111% | 0 | 0 | — |
case-08 | fail→pass | 15,893 | 4,773 | -70% | 1 | 1 | 0% | 1,677 | 1,737 | +4% | 0 | 0 | — |
case-09 | fail→pass | 35,198 | 3,603 | -90% | 1 | 1 | 0% | 3,039 | 1,573 | -48% | 0 | 0 | — |
case-10 | pass→pass | 7,811 | 6,324 | -19% | 1 | 1 | 0% | 972 | 2,015 | +107% | 0 | 0 | — |
case-11 | fail→pass | 9,779 | 2,780 | -72% | 1 | 1 | 0% | 1,685 | 1,450 | -14% | 0 | 0 | — |
case-12 | pass→pass | 10,133 | 18,315 | +81% | 1 | 1 | 0% | 1,344 | 1,662 | +24% | 0 | 0 | — |
case-13 | pass→pass | 14,013 | 9,468 | -32% | 1 | 1 | 0% | 1,970 | 2,375 | +21% | 0 | 0 | — |
case-14 | fail→pass | 8,278 | 5,981 | -28% | 1 | 1 | 0% | 1,384 | 2,183 | +58% | 0 | 0 | — |
case-15 | pass→pass | 10,774 | 3,936 | -63% | 1 | 1 | 0% | 1,422 | 1,598 | +12% | 0 | 0 | — |
case-16 | pass→pass | 7,581 | 5,219 | -31% | 1 | 1 | 0% | 1,022 | 1,849 | +81% | 0 | 0 | — |
case-17 | pass→pass | 11,409 | 6,072 | -47% | 1 | 1 | 0% | 1,575 | 1,865 | +18% | 0 | 0 | — |
case-18 | fail→pass | 9,507 | 4,691 | -51% | 1 | 1 | 0% | 1,374 | 1,920 | +40% | 0 | 0 | — |
case-19 | fail→pass | 9,099 | 2,564 | -72% | 1 | 1 | 0% | 1,251 | 1,361 | +9% | 0 | 0 | — |
case-20 | pass→pass | 40,172 | 45,329 | +13% | 1 | 1 | 0% | 8,257 | 9,315 | +13% | 0 | 0 | — |
case-21 | fail→pass | 35,239 | 10,933 | -69% | 1 | 1 | 0% | 1,186 | 2,988 | +152% | 0 | 0 | — |
case-22 | fail→pass | 18,490 | 17,376 | -6% | 1 | 1 | 0% | 2,782 | 2,716 | -2% | 0 | 0 | — |
case-23 | pass→pass | 7,545 | 4,949 | -34% | 1 | 1 | 0% | 1,094 | 1,827 | +67% | 0 | 0 | — |
case-24 | fail→pass | 23,680 | 5,072 | -79% | 1 | 1 | 0% | 2,299 | 1,627 | -29% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 19 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +46 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.