Install any skill in seconds. Free to start, no credit card required.
Get Started Free →출간 ready 직전 read-only critic. 8-item rubric 강제. 'looks good' rubber-stamp 금지. 라인 ref 필수. 2-round revision hard cap. '최종 QA', '출간 검수', 'rubric으로 평가' 류 트리거.
.claude/skills/kwakseongjae-omd-final-qa/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 504% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 90% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 127% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 96% | 0% |
<!-- omd:installed-skill — managed by omd install-skills. Do not edit; rerun the command to refresh. -->
artifact가 사용자에게 handoff 되기 전 마지막 게이트. read-only. 절대 수정하지 않고, verdict + actionable feedback만 emit.
artifact_paths: 모든 locale 파일 (예: index.ko.md, index.en.md)design_md_path: brand DESIGN.mdprior_reviews: designer-review 보고서 경로 (있으면)voice_preset: kr-writer가 사용한 preset_idround: 1 또는 2 (orchestrator가 주입)각 항목은 PASS / FAIL 이진. 회색 zone 없음. 1개라도 FAIL이면 verdict = REVISION (round 1) 또는 BLOCK (round 2).
voice_preset 룰 준수 (예: toss-tech-design이면 -요 종결 100%)<img> / ![]()에 alt 텍스트 존재source_revision 최신lang="ko" ts , python )font-display: swaptarget="_blank" rel="noopener")> 한국어 표기/맞춤법은 이 rubric의 자동 항목이 아니다. 자동 검사가 필요하면 > 외부 서비스(예: 부산대 한국어 맞춤법 검사기)를 작성자가 별도로 돌린다 — > 번들된 결정론적 도구는 없다.
design_md_read_at ISO timestamp 명시.round 1:
rubric FAIL → verdict = REVISION
→ orchestrator가 writer로 송환
round 2:
rubric FAIL → verdict = BLOCK
→ orchestrator가 사용자 escalation3번째 round 금지. 사용자가 강제 통과 결정 시 known-issues로 frontmatter 기록 후 ship.
<work_dir>/.reviews/final-qa-round-<N>.md:
markdown# Final QA — round <N> **Date:** <ISO> **Artifacts:** <list> **DESIGN.md read at:** <ISO> **Voice preset:** <preset_id> ## Rubric | # | Item | Verdict | Evidence | |---|---|---|---| | 1 | Brand consistency | PASS | DESIGN.md 12 tokens 모두 사용. 추가 hex 0건. | | 2 | Typography hierarchy | FAIL | `index.ko.md:88` h2 → h4 skip. h3 누락. | | 3 | Voice register | PASS | 종결어미 `-요` 198/198 (100%). | | 4 | Image / figure | FAIL | `index.ko.md:34` alt 텍스트 누락. | | 5 | Cross-locale parity | PASS | KR 6 H2 = EN 6 H2. | | 6 | Accessibility | PASS | 대비비 7.1:1. focus-visible 정의됨. | | 7 | Performance | PASS | 모든 이미지 < 200KB. | | 8 | Links | PASS | 23 link, broken 0. | ## Failed items detail ### [2] Typography hierarchy — h-level skip - Location: `index.ko.md:88` - Evidence: `## 가져가도 좋은 것` → 다음에 `#### 1. 토큰` (h4) - Fix: h3로 변경 ### [4] alt 텍스트 - Location: `index.ko.md:34` - Evidence: `` - Fix: alt 채움. 예: `` ## Verdict **REVISION** (round 1) — 2 items FAIL. Writer로 송환. 다음 round에서 동일 항목 FAIL 시 BLOCK.
새 rubric item 추가는:
omd-designer-review 보고서 ← prior_reviews로 참조 (해소 여부 확인)omd-locale-adapter 결과물 ← rubric 5]에서 검증verdict를 emit한 후에, 이번 run의 rubric FAIL들을 한 번 스캔해 반복 위반을 취향 후보로 제안한다. read-only 원칙은 그대로 — artifact·DESIGN.md는 절대 건드리지 않으며, 이 스킬이 트리거할 수 있는 유일한 쓰기는 사용자가 명시적으로 동의한 뒤의 preference append이고, 그것도 omd:remember의 canonical 절차 그대로다.
.omd/preferences.md가 존재하면 read해서, FAIL이 status: pending 엔트리의 scope와 같은 축이면 1회여도 후보.omd/preferences.md가 없으면 조건 2는 생략 — 파일을 만들지 않는다. a11y/performance/links(6]-8])는 취향이 아니라 hard rule — 후보에서 제외.
후보가 1개 이상이면 단 한 번 묻는다: "이 패턴, 취향으로 기록할까요?"
동의된 후보는 omd:remember 스킬의 기록 절차를 그대로 수행해 기록 (직접 포맷 모방 금지 — writer는 omd:remember 하나) — signal: review / confidence: inferred / status: pending, source_context는 이번 QA report 경로 (예: .reviews/final-qa-round-1.md).
status: pending 기록까지만. DESIGN.md 반영은 omd:learn의 평소 임계/게이트가 결정> 수동 검증: KR/EN 두 artifact에서 rubric 3] Voice register가 모두 FAIL인 QA를 돌리면, verdict 출력 후 "이 패턴, 취향으로 기록할까요?" 질문이 정확히 1회 뜨고, 동의 시 .omd/preferences.md에 scope: voice / signal: review / confidence: inferred / status: pending 엔트리가 append되어야 한다 (artifact·DESIGN.md는 변경 없음).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,488 | 4,425 | +27% | 1 | 1 | 0% | 255 | 2,947 | +1056% | 0 | 0 | — |
case-02 | fail→fail | 6,152 | 5,166 | -16% | 1 | 1 | 0% | 367 | 2,749 | +649% | 0 | 0 | — |
case-03 | fail→fail | 4,423 | 4,459 | +1% | 1 | 1 | 0% | 228 | 2,744 | +1104% | 0 | 0 | — |
case-04 | fail→fail | 3,307 | 5,054 | +53% | 1 | 1 | 0% | 455 | 2,673 | +487% | 0 | 0 | — |
case-05 | fail→fail | 7,465 | 5,465 | -27% | 1 | 1 | 0% | 425 | 2,776 | +553% | 0 | 0 | — |
case-06 | fail→pass | 4,219 | 9,970 | +136% | 1 | 1 | 0% | 650 | 3,925 | +504% | 0 | 0 | — |
case-07 | fail→pass | 9,274 | 2,778 | -70% | 1 | 1 | 0% | 1,505 | 2,859 | +90% | 0 | 0 | — |
case-08 | fail→pass | 7,860 | 2,462 | -69% | 1 | 1 | 0% | 1,213 | 2,749 | +127% | 0 | 0 | — |
case-09 | pass→pass | 9,059 | 3,705 | -59% | 1 | 1 | 0% | 1,375 | 3,015 | +119% | 0 | 0 | — |
case-10 | fail→pass | 9,985 | 3,393 | -66% | 1 | 1 | 0% | 1,552 | 3,008 | +94% | 0 | 0 | — |
case-11 | pass→pass | 10,538 | 4,593 | -56% | 1 | 1 | 0% | 1,607 | 3,151 | +96% | 0 | 0 | — |
case-12 | pass→pass | 6,573 | 4,069 | -38% | 1 | 1 | 0% | 1,088 | 2,998 | +176% | 0 | 0 | — |
case-13 | fail→pass | 8,775 | 3,059 | -65% | 1 | 1 | 0% | 1,506 | 2,945 | +96% | 0 | 0 | — |
case-14 | pass→pass | 8,424 | 4,578 | -46% | 1 | 1 | 0% | 1,385 | 3,152 | +128% | 0 | 0 | — |
case-15 | pass→pass | 8,755 | 2,012 | -77% | 1 | 1 | 0% | 1,448 | 2,669 | +84% | 0 | 0 | — |
case-16 | fail→pass | 8,363 | 4,266 | -49% | 1 | 1 | 0% | 1,301 | 3,143 | +142% | 0 | 0 | — |
case-17 | fail→pass | 5,565 | 4,282 | -23% | 1 | 1 | 0% | 965 | 3,129 | +224% | 0 | 0 | — |
case-18 | pass→pass | 5,778 | 3,664 | -37% | 1 | 1 | 0% | 934 | 3,004 | +222% | 0 | 0 | — |
case-19 | fail→pass | 11,839 | 6,135 | -48% | 1 | 1 | 0% | 1,833 | 3,399 | +85% | 0 | 0 | — |
case-20 | fail→pass | 10,713 | 5,805 | -46% | 1 | 1 | 0% | 1,657 | 3,417 | +106% | 0 | 0 | — |
case-21 | pass→pass | 12,468 | 3,656 | -71% | 1 | 1 | 0% | 1,837 | 3,026 | +65% | 0 | 0 | — |
case-22 | fail→pass | 11,518 | 2,717 | -76% | 1 | 1 | 0% | 1,763 | 2,854 | +62% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.