Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Consult Claude Code for second opinions, brainstorming, or difficult debugging from Codex. Use when the user asks to "consult claude", "ask claude", "get claude's opinion", "brainstorm with claude", or "discuss with claude".
.claude/skills/tobihagemann-consult-claude/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-19 | ✓→✗ | ▼ Worse | 27% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 29% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 66% | 0% |
Use Claude Code as an external collaborator from Codex. Unlike $peer-review, this skill is conversational and exploratory.
State:
When a recommendation is wanted, bar answers that appeal to scope: state that "out of scope" or "leave it alone" is not an acceptable argument on its own, and that recommending no change must be justified on technical merit. Demand one pick per decision, the reasoning, and the strongest counterargument to that pick, with hedging across options ruled out.
When the consultation is a prose rewrite bound by a house style, name the shapes that style forbids in the first prompt, so they do not have to be corrected across follow-up turns. Common ones: prefixing a summary with a grammatical subject the convention omits, expanding a pronoun to its full noun phrase at every occurrence, and splitting a sentence so a condition is restated in both halves.
$claude-print SkillRun the $claude-print skill with the assembled question. Default to read-only permissions.
For follow-up questions, include Claude's previous answer and the new evidence gathered since then.
When the recommendation would violate a documented constraint, follow up rather than discarding or adopting it. Quote the constraint back and ask Claude to argue it out: whether the constraint is sound or was set without the problem Claude identified in view, whether that problem is reachable given code Claude may not have accounted for, and what the best fix that respects the constraint is. Ask it to quantify the exposure rather than assert it, and say that reversing its prior recommendation is acceptable.
Summarize the useful parts of Claude's response. Cross-reference suggestions with the repository before acting.
When the consultation rewrote prose rather than answering a question, check the rewrite against the source yourself before adopting it. Treat its own report that the rewrite is faithful as a claim awaiting verification. Verify the source's own factual claims against what they describe, since a rewrite can be faithful to a source that was itself wrong. Read for these drift shapes in the rewrite:
Take the plainer sentences and keep the load-bearing why.
When the consultation was opened from a pending question, resolve that question with the answer in hand, re-asking the user when the choice stays theirs. Then call update_plan to mark this step completed and continue with the next step of the active workflow.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | fail→fail | 9,200 | 4,856 | -47% | 1 | 1 | 0% | 1,357 | 1,005 | -26% | 0 | 0 | — |
case-13 | fail→fail | 3,444 | 3,974 | +15% | 1 | 1 | 0% | 248 | 962 | +288% | 0 | 0 | — |
case-01 | fail→fail | 19,913 | 78,078 | +292% | 1 | 1 | 0% | 3,641 | 3,403 | -7% | 0 | 0 | — |
case-02 | fail→fail | 20,230 | 34,692 | +71% | 1 | 1 | 0% | 3,820 | 1,038 | -73% | 0 | 0 | — |
case-03 | fail→fail | 9,580 | 48,365 | +405% | 1 | 1 | 0% | 1,137 | 3,644 | +220% | 0 | 0 | — |
case-14 | fail→fail | 6,643 | 5,921 | -11% | 1 | 1 | 0% | 1,192 | 1,609 | +35% | 0 | 0 | — |
case-04 | fail→fail | 12,657 | 21,321 | +68% | 1 | 1 | 0% | 1,428 | 2,950 | +107% | 0 | 0 | — |
case-05 | pass→pass | 6,734 | 35,912 | +433% | 1 | 1 | 0% | 887 | 1,476 | +66% | 0 | 0 | — |
case-06 | fail→fail | 9,139 | 15,290 | +67% | 1 | 1 | 0% | 1,285 | 3,206 | +149% | 0 | 0 | — |
case-07 | fail→fail | 9,858 | 5,581 | -43% | 1 | 1 | 0% | 1,620 | 1,105 | -32% | 0 | 0 | — |
case-08 | pass→pass | 13,205 | 19,375 | +47% | 1 | 1 | 0% | 2,104 | 1,837 | -13% | 0 | 0 | — |
case-09 | fail→fail | 8,724 | 6,554 | -25% | 1 | 1 | 0% | 1,423 | 1,118 | -21% | 0 | 0 | — |
case-10 | fail→pass | 9,320 | 6,652 | -29% | 1 | 1 | 0% | 1,515 | 1,702 | +12% | 0 | 0 | — |
case-11 | fail→pass | 8,178 | 7,176 | -12% | 1 | 1 | 0% | 1,376 | 1,822 | +32% | 0 | 0 | — |
case-15 | fail→fail | 10,817 | 5,523 | -49% | 1 | 1 | 0% | 1,857 | 1,098 | -41% | 0 | 0 | — |
case-16 | pass→pass | 10,193 | 3,799 | -63% | 1 | 1 | 0% | 1,767 | 1,491 | -16% | 0 | 0 | — |
case-17 | pass→pass | 6,380 | 5,002 | -22% | 1 | 1 | 0% | 1,068 | 1,695 | +59% | 0 | 0 | — |
case-18 | pass→pass | 7,851 | 6,424 | -18% | 1 | 1 | 0% | 1,370 | 1,676 | +22% | 0 | 0 | — |
case-19 | pass→fail | 7,843 | 6,675 | -15% | 1 | 1 | 0% | 1,021 | 1,293 | +27% | 0 | 0 | — |
case-20 | pass→pass | 7,105 | 4,485 | -37% | 1 | 1 | 0% | 1,203 | 1,425 | +18% | 0 | 0 | — |
case-21 | pass→pass | 6,002 | 4,472 | -25% | 1 | 1 | 0% | 1,053 | 1,553 | +47% | 0 | 0 | — |
case-22 | pass→fail | 8,246 | 8,998 | +9% | 1 | 1 | 0% | 1,197 | 1,549 | +29% | 0 | 0 | — |
case-23 | fail→fail | 4,809 | 16,594 | +245% | 1 | 1 | 0% | 268 | 3,045 | +1036% | 0 | 0 | — |
case-24 | fail→fail | 11,653 | 6,027 | -48% | 1 | 1 | 0% | 2,061 | 1,155 | -44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 13 counted toward the lift figure. The other 11 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 13 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/21/2026 | +27% |
Other measured skills in the registry, with their headline benchmark lift.