Install any skill in seconds. Free to start, no credit card required.
Get Started Free →/cs:cross-eval <memo> — Multi-model consensus on a board memo or strategy brief. Claude + Codex + Gemini cross-review with graceful degradation. Use when a high-stakes memo needs an independent sanity check before the boardroom — e.g. a bet-the-company pivot or fundraise terms.
.claude/skills/alirezarezvani-cross-eval/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -37% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 42% | 0% |
Command: /cs:cross-eval <memo-or-brief>
Runs the same memo through multiple model providers and reconciles divergences. Use for high-stakes, irreversible decisions where single-model bias is too costly: M&A, major fundraises, layoffs, strategic pivots, regulatory commitments.
Adapted from gstack's /codex cross-review pattern, generalized to business memos instead of code PRs.
The command tries to invoke each available model in order:
OPENAI_API_KEY or codex CLI available)GEMINI_API_KEY or gemini CLI available)If only Claude is available, the command runs Claude-only with adversarial mode — same model, different prompt seeds — and clearly labels the output as single-model.
> "You are an independent C-suite reviewer. The following is a board memo from another company's boardroom. Identify the top 3 concerns, the top 3 supports, and your vote (APPROVE / REJECT / DEFER). Do not deferentially agree — assume the memo's reasoning is flawed until proven otherwise."
Saved to ~/.claude/cross-eval/YYYY-MM-DD-<slug>.md:
markdown# Cross-Eval: <memo title> **Date:** YYYY-MM-DD **Memo reviewed:** <link> **Models invoked:** Claude / Codex / Gemini (or noted fallbacks) ## Vote Tally | Model | Vote | Confidence | |---|---|---| | Claude | APPROVE | High | | Codex | DEFER | Med | | Gemini | APPROVE | Low | ## Consensus Concerns (≥2 models flagged) 1. <concern> — flagged by Claude + Codex 2. <concern> — flagged by all 3 ## Divergent Concerns (1 model flagged) - <Codex only:> <concern> — worth a second look - <Gemini only:> <concern> — likely noise, but check ## Consensus Supports (≥2 models endorsed) 1. <support> 2. <support> ## Recommendation - 🟢 GO if 2+ models APPROVE and no CRITICAL concerns from any model - 🟡 PAUSE if any model is DEFER or any concern is CRITICAL - 🔴 STOP if 2+ models REJECT ## Open Questions for Founder 1. <question raised by divergence> 2. <question raised by divergence>
Single-model recommendations have systematic biases. Claude trends helpful and may under-weight risk. Codex (OpenAI) trends more cautious on emerging-market and regulatory topics. Gemini trends more cautious on technical scale claims. Disagreement is signal, not noise.
This is the safety net before irreversibility — not a replacement for outside counsel or a real board.
If only Claude is available:
markdown**Models available:** Claude only **Mode:** ADVERSARIAL — running 3 independent Claude passes with different system prompts: 1. Standard reviewer 2. Devil's advocate (must find 3 critical concerns) 3. Steelman (must find 3 strongest reasons to approve) This is weaker than true multi-model. Treat the result as suggestive, not conclusive.
/cs:decide — if consensus is GO/cs:freeze — if consensus is PAUSE/cs:boardroom (re-run) — if consensus is STOPboard-meeting, executive-mentor/codex cross-review pattern (adapted to business memos)Version: 1.0.0
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | pass→pass | 17,237 | 9,553 | -45% | 1 | 1 | 0% | 1,923 | 2,601 | +35% | 0 | 0 | — |
case-06 | pass→pass | 13,503 | 3,013 | -78% | 1 | 1 | 0% | 1,455 | 1,616 | +11% | 0 | 0 | — |
case-07 | pass→pass | 13,708 | 9,371 | -32% | 1 | 1 | 0% | 1,594 | 1,818 | +14% | 0 | 0 | — |
case-13 | pass→pass | 11,320 | 12,244 | +8% | 1 | 1 | 0% | 1,677 | 2,133 | +27% | 0 | 0 | — |
case-01 | pass→fail | 28,677 | 29,875 | +4% | 1 | 1 | 0% | 4,135 | 4,196 | +1% | 0 | 0 | — |
case-02 | fail→pass | 31,381 | 20,801 | -34% | 1 | 1 | 0% | 3,563 | 4,465 | +25% | 0 | 0 | — |
case-03 | fail→fail | 32,555 | 21,772 | -33% | 1 | 1 | 0% | 4,482 | 3,420 | -24% | 0 | 0 | — |
case-04 | fail→pass | 9,782 | 11,170 | +14% | 1 | 1 | 0% | 1,308 | 2,172 | +66% | 0 | 0 | — |
case-05 | fail→pass | 21,997 | 7,995 | -64% | 1 | 1 | 0% | 1,043 | 1,541 | +48% | 0 | 0 | — |
case-08 | fail→pass | 15,930 | 8,009 | -50% | 1 | 1 | 0% | 2,616 | 1,656 | -37% | 0 | 0 | — |
case-09 | fail→pass | 8,882 | 9,917 | +12% | 1 | 1 | 0% | 1,299 | 1,848 | +42% | 0 | 0 | — |
case-10 | pass→pass | 18,267 | 8,486 | -54% | 1 | 1 | 0% | 1,974 | 2,303 | +17% | 0 | 0 | — |
case-11 | pass→pass | 10,257 | 15,233 | +49% | 1 | 1 | 0% | 1,626 | 2,401 | +48% | 0 | 0 | — |
case-12 | pass→pass | 10,073 | 17,340 | +72% | 1 | 1 | 0% | 1,627 | 2,489 | +53% | 0 | 0 | — |
case-14 | fail→pass | 9,462 | 14,048 | +48% | 1 | 1 | 0% | 1,490 | 2,458 | +65% | 0 | 0 | — |
case-15 | pass→pass | 7,645 | 8,877 | +16% | 1 | 1 | 0% | 1,148 | 1,697 | +48% | 0 | 0 | — |
case-16 | fail→pass | 14,478 | 13,554 | -6% | 1 | 1 | 0% | 2,241 | 3,142 | +40% | 0 | 0 | — |
case-17 | pass→pass | 15,474 | 7,070 | -54% | 1 | 1 | 0% | 1,672 | 2,263 | +35% | 0 | 0 | — |
case-19 | pass→pass | 10,561 | 9,508 | -10% | 1 | 1 | 0% | 1,576 | 2,552 | +62% | 0 | 0 | — |
case-20 | pass→pass | 21,736 | 9,276 | -57% | 1 | 1 | 0% | 1,404 | 2,197 | +56% | 0 | 0 | — |
case-21 | pass→pass | 14,753 | 20,704 | +40% | 1 | 1 | 0% | 2,397 | 3,738 | +56% | 0 | 0 | — |
case-22 | pass→pass | 12,089 | 11,494 | -5% | 1 | 1 | 0% | 1,845 | 2,866 | +55% | 0 | 0 | — |
case-23 | pass→pass | 19,752 | 10,627 | -46% | 1 | 1 | 0% | 1,377 | 1,895 | +38% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +26 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/4/2026 | +55% |
Other measured skills in the registry, with their headline benchmark lift.