Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Five-role Design Jury review that streams scored rounds, persists a replayable transcript, and ships through the daemon's Critique Theater protocol.
.claude/skills/nexu-io-critique-theater/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -31% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -34% | 0% |
Design Jury is the user-facing name for Critique Theater. When the daemon enables it for an artifact-generating run, it appends the active protocol, brand source, thresholds, and weights to the agent prompt. The agent then plays the fixed v1 cast — Designer, Critic, Brand, Accessibility, and Copy — as turns inside the same CLI session.
<CRITIQUE_RUN> envelope containing <ROUND>, <PANELIST>, and <ROUND_END> blocks and, only after a ship decision, one final <SHIP> block. Do not emit prose outside the envelope.
their own scopes and record actionable <MUST_FIX> items. The cast is fixed in v1; a plugin does not replace it with a custom list of axes.
PanelEvent variantsand publishes them on critique.* SSE channels. It recomputes the composite from the scoring panelists, so an agent-supplied composite is advisory.
critique.json in the project cwd or emit akind: "critique-panel" object. Those were the pre-orchestrator shape and are not inputs to the current runtime.
The default review uses a 0–10 scale, an 8.0 ship threshold, and at most three rounds. A round ships only when its daemon-computed composite reaches the configured threshold and no open must-fix items remain. The active values come from OD_CRITIQUE_* configuration, including OD_CRITIQUE_MAX_ROUNDS, OD_CRITIQUE_SCORE_THRESHOLD, and the fallback policy.
OD_MAX_DEVLOOP_ITERATIONS still caps an outer plugin pipeline stage; it is not the Design Jury round limit or ship rule. Likewise, a pipeline's critique.score signal is a scheduler-facing projection, not the wire output the agent should manufacture.
The daemon stores the terminal run and per-round summaries in SQLite and writes the ordered PanelEvent stream as a replayable NDJSON transcript (gzip-compressed when large). Final artifact bytes are stored separately; the ship event carries only an artifactRef.
The web Design Jury surface exposes Interrupt and an Esc shortcut. After the user triggers either, the web app posts to /api/projects/:projectId/critique/:runId/interrupt; the daemon aborts the registered run, persists the partial best-so-far state, and emits critique.interrupted. od ui respond handles GenUI surfaces and does not provide a break-loop action for Critique Theater.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 39,708 | 26,986 | -32% | 1 | 1 | 0% | 8,269 | 5,964 | -28% | 0 | 0 | — |
case-02 | fail→pass | 49,448 | 26,589 | -46% | 1 | 1 | 0% | 8,258 | 5,801 | -30% | 0 | 0 | — |
case-03 | fail→pass | 45,006 | 67,161 | +49% | 1 | 1 | 0% | 8,267 | 6,314 | -24% | 0 | 0 | — |
case-04 | fail→pass | 18,206 | 6,336 | -65% | 1 | 1 | 0% | 2,357 | 1,628 | -31% | 0 | 0 | — |
case-05 | fail→pass | 13,777 | 3,622 | -74% | 1 | 1 | 0% | 1,733 | 1,139 | -34% | 0 | 0 | — |
case-06 | fail→pass | 16,645 | 8,085 | -51% | 1 | 1 | 0% | 2,131 | 1,659 | -22% | 0 | 0 | — |
case-07 | fail→pass | 23,633 | 5,561 | -76% | 1 | 1 | 0% | 2,025 | 1,446 | -29% | 0 | 0 | — |
case-08 | fail→pass | 58,079 | 2,485 | -96% | 1 | 1 | 0% | 2,049 | 879 | -57% | 0 | 0 | — |
case-09 | fail→pass | 17,595 | 5,136 | -71% | 1 | 1 | 0% | 1,381 | 944 | -32% | 0 | 0 | — |
case-10 | pass→pass | 22,915 | 3,831 | -83% | 1 | 1 | 0% | 2,180 | 1,164 | -47% | 0 | 0 | — |
case-11 | pass→pass | 7,083 | 4,193 | -41% | 1 | 1 | 0% | 876 | 1,159 | +32% | 0 | 0 | — |
case-12 | fail→pass | 37,546 | 2,792 | -93% | 1 | 1 | 0% | 2,371 | 1,000 | -58% | 0 | 0 | — |
case-13 | fail→pass | 8,166 | 2,267 | -72% | 1 | 1 | 0% | 1,382 | 1,033 | -25% | 0 | 0 | — |
case-14 | fail→pass | 13,392 | 5,704 | -57% | 1 | 1 | 0% | 2,010 | 1,119 | -44% | 0 | 0 | — |
case-15 | fail→pass | 16,502 | 5,730 | -65% | 1 | 1 | 0% | 2,660 | 1,623 | -39% | 0 | 0 | — |
case-16 | fail→pass | 33,694 | 2,227 | -93% | 1 | 1 | 0% | 1,122 | 1,042 | -7% | 0 | 0 | — |
case-17 | fail→pass | 13,890 | 1,819 | -87% | 1 | 1 | 0% | 1,956 | 967 | -51% | 0 | 0 | — |
case-18 | pass→pass | 14,000 | 2,538 | -82% | 1 | 1 | 0% | 2,214 | 982 | -56% | 0 | 0 | — |
case-24 | fail→pass | 15,727 | 3,530 | -78% | 1 | 1 | 0% | 1,933 | 1,074 | -44% | 0 | 0 | — |
case-19 | fail→pass | 9,633 | 2,107 | -78% | 1 | 1 | 0% | 1,413 | 979 | -31% | 0 | 0 | — |
case-20 | fail→pass | 21,071 | 12,699 | -40% | 1 | 1 | 0% | 3,254 | 2,443 | -25% | 0 | 0 | — |
case-21 | fail→pass | 15,527 | 6,781 | -56% | 1 | 1 | 0% | 2,059 | 1,535 | -25% | 0 | 0 | — |
case-22 | fail→pass | 13,706 | 4,352 | -68% | 1 | 1 | 0% | 1,733 | 1,183 | -32% | 0 | 0 | — |
case-23 | fail→pass | 4,203 | 4,884 | +16% | 1 | 1 | 0% | 669 | 1,274 | +90% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 23 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +88 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.