Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a Virtual Think Tank — a structured multi-persona debate — before planning or making architectural/design/strategic decisions. Use this skill whenever the user is about to plan a system, make a technology choice, evaluate trade-offs, decide on an approach, or faces any decision where multiple perspectives would sharpen the outcome. Also trigger when the user says "think tank", "debate this", "perspectives on", "trade-offs", "should I use X or Y", "help me decide", "before we plan", or asks f
.claude/skills/davila7-think-tank/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | 112% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-04 | ✓→✗ | ▼ Worse | 47% | 0% |
| case-05 | ✓→✗ | ▼ Worse | 16% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 235% | 0% |
A pre-planning skill that simulates a moderated expert debate to surface trade-offs, blind spots, and perspectives before committing to a plan. Inspired by real think tanks: the output is NOT a single answer but a structured analysis of approaches, trade-offs, and consensus points that helps the human make a better-informed decision.
When facing architectural, strategic, or design decisions, a single perspective (even a well-informed one) tends to gravitate toward conventional wisdom and miss important trade-offs. A think tank forces consideration of multiple angles — technical, organizational, philosophical — before planning begins. The result is plans that account for more of reality.
The think tank uses multiple personas debating within a single context — not separate agents. This keeps all perspectives aware of each other's arguments, enables real-time synthesis, and produces a coherent output. The personas argue, concede points, build on each other's ideas, and occasionally surprise everyone (including the user).
Before assembling the panel, clearly understand what's being decided. Ask the user (if not already clear):
Restate the problem back to the user in a crisp problem statement before proceeding. This ensures the think tank debates the right question.
Build a panel of 4–6 personas. The composition matters — diversity of perspective is the whole point.
Panel structure:
Important persona guidelines:
Structure the debate as a moderated discussion, not a series of independent monologues. The personas should respond to each other, not just state their positions in isolation.
The user is a participant, not a spectator. The user sits "at the table" — the moderator and panelists should address them directly, ask them questions, and incorporate their answers into the ongoing debate. The user is the decision-maker; the panel is there to serve them. Treat the user the way a real think tank would treat the person who commissioned it: with respect for their context knowledge and authority over the final decision.
Debate structure:
This is a real pause — wait for the user's response and feed it back into the debate. If the user reveals something (e.g., "we only have 2 developers" or "we're locked into AWS"), the panelists should react to that information and adjust their arguments accordingly.
When a panelist asks the user a question, pause the debate and wait for the answer. Then resume with the panelists reacting to the new information. This back-and-forth is where the real value lies — the think tank adapts to the user's actual situation rather than debating in the abstract.
This gives the user a chance to redirect before the summary phase.
Tone guidelines:
After the debate, produce a structured summary. This is what feeds into the planning phase.
Output format:
## Think Tank Summary: [Problem Statement]
### Panel
[List panelists and their roles/perspectives]
### Key Debate Highlights
[2-3 of the most illuminating exchanges or insights from the debate — the moments where something shifted or crystallized. Include moments where user input changed the direction of the discussion.]
### User-Revealed Context
[Key constraints, preferences, or realities the user shared during the debate that shaped the panel's thinking. This section ensures nothing the user said gets lost.]
### Consensus Points
[Things all or most panelists agreed on — these are high-confidence inputs to planning]
### Core Trade-offs
[The real axes of disagreement, stated as trade-offs rather than as one side being right]
- Trade-off 1: [X] vs [Y] — choosing X gives you [...] but costs you [...]
- Trade-off 2: ...
### Conditional Recommendations
[Recommendations framed as "if-then" rather than absolutes]
- If [condition], then [approach] because [reasoning]
- If [condition], then [approach] because [reasoning]
### Risks & Blind Spots
[Things the panel identified as under-discussed or easy to overlook]
### Open Questions
[Questions that couldn't be resolved in the debate and need more information or experimentation to answer]
### Suggested Next Steps
[Concrete actions: things to research, prototype, test, or decide before planning]After presenting the summary, ask the user:
The think tank output should be treated as input to the plan — not as the plan itself. The human makes the decision; the think tank provides the analysis.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 28,330 | 36,319 | +28% | 1 | 1 | 0% | 4,437 | 7,926 | +79% | 0 | 0 | — |
case-02 | fail→fail | 13,872 | 13,682 | -1% | 1 | 1 | 0% | 2,042 | 4,436 | +117% | 0 | 0 | — |
case-03 | fail→fail | 28,537 | 38,007 | +33% | 1 | 1 | 0% | 4,366 | 8,366 | +92% | 0 | 0 | — |
case-04 | pass→fail | 18,451 | 11,103 | -40% | 1 | 1 | 0% | 2,843 | 4,165 | +47% | 0 | 0 | — |
case-05 | pass→fail | 24,464 | 15,580 | -36% | 1 | 1 | 0% | 4,264 | 4,941 | +16% | 0 | 0 | — |
case-06 | pass→fail | 9,242 | 25,573 | +177% | 1 | 1 | 0% | 2,127 | 7,118 | +235% | 0 | 0 | — |
case-07 | pass→pass | 15,115 | 13,415 | -11% | 1 | 1 | 0% | 2,184 | 4,554 | +109% | 0 | 0 | — |
case-08 | fail→fail | 10,062 | 4,545 | -55% | 1 | 1 | 0% | 1,582 | 3,159 | +100% | 0 | 0 | — |
case-09 | pass→pass | 10,039 | 13,776 | +37% | 1 | 1 | 0% | 1,472 | 4,445 | +202% | 0 | 0 | — |
case-10 | pass→pass | 13,667 | 15,907 | +16% | 1 | 1 | 0% | 2,116 | 4,847 | +129% | 0 | 0 | — |
case-11 | pass→pass | 7,925 | 16,292 | +106% | 1 | 1 | 0% | 1,213 | 4,965 | +309% | 0 | 0 | — |
case-12 | pass→pass | 17,065 | 12,100 | -29% | 1 | 1 | 0% | 2,493 | 4,277 | +72% | 0 | 0 | — |
case-13 | pass→pass | 17,187 | 12,285 | -29% | 1 | 1 | 0% | 2,683 | 4,306 | +60% | 0 | 0 | — |
case-14 | fail→pass | 10,618 | 6,314 | -41% | 1 | 1 | 0% | 1,581 | 3,347 | +112% | 0 | 0 | — |
case-15 | pass→pass | 24,821 | 15,799 | -36% | 1 | 1 | 0% | 2,952 | 4,712 | +60% | 0 | 0 | — |
case-16 | pass→pass | 13,981 | 12,983 | -7% | 1 | 1 | 0% | 2,205 | 4,454 | +102% | 0 | 0 | — |
case-17 | fail→fail | 8,716 | 3,980 | -54% | 1 | 1 | 0% | 1,270 | 3,060 | +141% | 0 | 0 | — |
case-18 | pass→pass | 12,802 | 13,454 | +5% | 1 | 1 | 0% | 1,976 | 4,427 | +124% | 0 | 0 | — |
case-19 | pass→pass | 12,378 | 11,083 | -10% | 1 | 1 | 0% | 1,975 | 4,030 | +104% | 0 | 0 | — |
case-20 | pass→pass | 15,840 | 10,655 | -33% | 1 | 1 | 0% | 2,138 | 4,094 | +91% | 0 | 0 | — |
case-21 | pass→pass | 13,987 | 8,508 | -39% | 1 | 1 | 0% | 1,931 | 3,743 | +94% | 0 | 0 | — |
case-22 | fail→pass | 15,673 | 15,870 | +1% | 1 | 1 | 0% | 2,461 | 4,852 | +97% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of -5 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.