Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a design-challenging Claude Code review of local git changes in this repository. Args: --wait, --background, --base <ref>, --scope <auto|working-tree|branch>, --model <model>, --effort <low|medium|high|xhigh|max>, [focus text]. Defaults to opus with no forced effort. Use only when the user wants stronger scrutiny than a normal review, such as explicit tradeoff challenge, risky-change review, or custom focus text.
.claude/skills/hashgraph-online-adversarial-review/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 89% | 24 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 1236% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 1010% | 0% |
通常のコードレビューは「正しさの確認」に集中する。 敵対的レビューは 「どう壊れるか」「どう攻撃されるか」「どこが論理的に弱いか」 に集中する。
AIをレビューに使う最大の価値は、情報の整理ではなく 思考の死角を映す鏡 としての活用にある。 このスキルは、2系統の敵対的手法を体系化し、レビューの質を根本的に引き上げる。
| 手法 | 対策するバイアス | 核心の問い | | --------------- | -------------------------- | ------------------------------------ | | Pre-mortem | 生存バイアス・楽観バイアス | 「失敗したとして、なぜ?」 | | War Game | 自己中心バイアス | 「敵の立場から、どう攻撃する?」 | | Logic Torturing | 確証バイアス | 「この論理の穴を潰して」 |
| 手法 | 対象とするズレ | 核心の問い | | -------------------- | ------------------------ | --------------------------------------------- | | Self-Contradiction | 宣言と同一ファイルの実装 | 「規則 X を宣言した本人が破っていない?」 | | Refactor-Claim Audit | 完了主張と残骸 | 「『全部やった』を grep で反証できる?」 | | Cross-File Leakage | 構造変更と caller 側 | 「直したのは変更元だけ、参照元は?」 |
入力に応じて、適切な手法へルーティングする。複数手法の併用も可能。
| キーワード | 手法 | スキルID | | --------------------------------------------------- | -------------------- | ---------------------- | | 失敗, リスク, 負債, インシデント, pre-mortem | Pre-mortem | pre-mortem | | 攻撃, セキュリティ, 悪用, 脆弱性, war-game | War Game | war-game | | 論理, 判断, 根拠, なぜ, 代替案, logic | Logic Torturing | logic-torturing | | 自己矛盾, contradiction, 宣言と実装, declared but | Self-Contradiction | self-contradiction | | 削減, 完了, 全て置換, all replaced, -N%, リファクタ | Refactor-Claim Audit | refactor-claim-audit | | caller, 残骸, 参照漏れ, 再採番, leakage | Cross-File Leakage | cross-file-leakage | | 敵対的, adversarial, 全部, フル | 全手法実行 | 上記6つすべて |
text1. 変更内容の分類 ├─ 設計/ADR → Pre-mortem を優先実行 ├─ セキュリティ関連 → War Game を優先実行 ├─ 判断を含む変更 → Logic Torturing を実行 ├─ 宣言的フレーズ → Self-Contradiction を実行 ├─ 完了主張 → Refactor-Claim Audit を実行 └─ 構造変更 → Cross-File Leakage を実行 2. 各手法の実行(並列可能) ├─ Pre-mortem: 失敗シナリオ × 最大5件 ├─ War Game: 攻撃シナリオ × 最大5件 ├─ Logic Torturing: 論理検証 × 最大5件 ├─ Self-Contradiction: 宣言と実装の乖離 × 最大5件 ├─ Refactor-Claim Audit: 完了主張の反証 × 最大5件 └─ Cross-File Leakage: caller 側残骸 × 最大5件 3. 統合サマリの生成 ├─ 重複する指摘の統合 ├─ 重大度による優先順位付け └─ Human Handoff 条件の判定
markdown## 🔍 Adversarial Review Summary ### 検出された盲点: N件 - Pre-mortem: X件 (失敗シナリオ) - War Game: Y件 (攻撃シナリオ) - Logic Torturing: Z件 (論理的な穴) - Self-Contradiction: A件 (宣言と実装の乖離) - Refactor-Claim Audit: B件 (完了主張の反証) - Cross-File Leakage: C件 (caller 側残骸) ### 最も重大な発見 <最も致命的な1件の要約> ### 詳細 (各手法の出力を統合)
| 既存スキル | 関係 | 棲み分け | | ---------------------------- | ---- | ----------------------------------------------------------------------------------------------------------------- | | architecture-risk-register | 補完 | risk-register は「リスクが文書化されているか」を確認。Pre-mortem は「文書化されていないリスクを発見」する | | security-basic | 補完 | security-basic は既知パターン(SQLi, XSS等)を検出。War Game は「既知パターンに当てはまらない攻撃経路」を発見する | | adr-decision-quality | 補完 | adr-decision は ADR の形式品質を確認。Logic Torturing は「記述された判断の論理的強度」を検証する |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 30,645 | 35,836 | +17% | 1 | 1 | 0% | 3,874 | 6,119 | +58% | 0 | 0 | — |
case-02 | fail→pass | 26,113 | 27,416 | +5% | 1 | 1 | 0% | 3,336 | 4,947 | +48% | 0 | 0 | — |
case-03 | pass→pass | 37,198 | 36,280 | -2% | 1 | 1 | 0% | 3,994 | 6,195 | +55% | 0 | 0 | — |
case-04 | fail→pass | 17,525 | 22,481 | +28% | 1 | 1 | 0% | 2,572 | 4,733 | +84% | 0 | 0 | — |
case-05 | fail→fail | 16,808 | 16,153 | -4% | 1 | 1 | 0% | 2,242 | 3,760 | +68% | 0 | 0 | — |
case-06 | pass→pass | 16,930 | 16,374 | -3% | 1 | 1 | 0% | 2,571 | 3,951 | +54% | 0 | 0 | — |
case-07 | pass→pass | 19,303 | 28,098 | +46% | 1 | 1 | 0% | 2,539 | 4,755 | +87% | 0 | 0 | — |
case-08 | pass→pass | 18,226 | 20,832 | +14% | 1 | 1 | 0% | 2,402 | 4,076 | +70% | 0 | 0 | — |
case-09 | fail→pass | 9,272 | 19,747 | +113% | 1 | 1 | 0% | 293 | 3,914 | +1236% | 0 | 0 | — |
case-10 | pass→pass | 22,435 | 34,358 | +53% | 1 | 1 | 0% | 3,444 | 6,152 | +79% | 0 | 0 | — |
case-11 | fail→pass | 19,227 | 31,114 | +62% | 1 | 1 | 0% | 492 | 5,460 | +1010% | 0 | 0 | — |
case-12 | fail→pass | 30,214 | 32,619 | +8% | 1 | 1 | 0% | 1,327 | 5,057 | +281% | 0 | 0 | — |
case-13 | fail→pass | 16,623 | 34,188 | +106% | 1 | 1 | 0% | 1,310 | 6,564 | +401% | 0 | 0 | — |
case-14 | fail→pass | 15,319 | 25,413 | +66% | 1 | 1 | 0% | 1,737 | 5,126 | +195% | 0 | 0 | — |
case-15 | pass→pass | 19,016 | 26,906 | +41% | 1 | 1 | 0% | 2,853 | 4,749 | +66% | 0 | 0 | — |
case-16 | pass→pass | 13,104 | 20,452 | +56% | 1 | 1 | 0% | 2,126 | 4,655 | +119% | 0 | 0 | — |
case-22 | pass→fail | 14,205 | 27,079 | +91% | 1 | 1 | 0% | 2,491 | 4,535 | +82% | 0 | 0 | — |
case-17 | pass→pass | 17,505 | 28,410 | +62% | 1 | 1 | 0% | 1,867 | 4,607 | +147% | 0 | 0 | — |
case-18 | fail→pass | 22,125 | 15,235 | -31% | 1 | 1 | 0% | 2,383 | 3,715 | +56% | 0 | 0 | — |
case-19 | fail→pass | 7,737 | 29,930 | +287% | 1 | 1 | 0% | 1,172 | 5,455 | +365% | 0 | 0 | — |
case-20 | fail→fail | 7,349 | 12,901 | +76% | 1 | 1 | 0% | 351 | 2,714 | +673% | 0 | 0 | — |
case-21 | pass→fail | 18,491 | 30,222 | +63% | 1 | 1 | 0% | 2,093 | 5,067 | +142% | 0 | 0 | — |
case-23 | pass→pass | 10,566 | 5,145 | -51% | 1 | 1 | 0% | 989 | 2,311 | +134% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +35 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.