Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Parallel multi-AI cross-validation research workflow (大版本). Dispatch N internal sub-agents + grok + gemini in parallel, automatically cross-validate findings, tier by confidence (strong consensus / partial / conflict / insufficient), generate tiered action items with arbitration. Use when user says "多 AI 调研", "交叉验证", "独立共识", "三脑调研", "multi-ai research", "parallel research", "cross-validate", or needs deep research that benefits from internal data + external 2026 consensus. NOT for quick factual
.claude/skills/majiayu000-multi-ai-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 248% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 127% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 137% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 202% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 117% | 0% |
<!-- v1 | 2026-04-09 | 从 2026-04-08 X reply deboost 调研实战沉淀;集成自动置信度分级 + 仲裁 + 改动清单 -->
核心价值:能力乘法,不是加法。Claude(主脑)+ grok(X 社区/实时)+ gemini(Google 生态/结构化)+ N 个内部 sub-agent = N+3 个 agent 并行处理同一个研究问题。
关键洞察(2026-04-08 实战验证):两个独立外部 AI 的共识信号 强于 任何单个 AI 的深度。"更深度" < "更少错"。
和 ask-opencli 的关系:
ask-opencli = 单次 grok 或 gemini 调用(日常 second opinion)multi-ai-research = 完整调研工作流,并行多 AI + 内部数据 + 交叉验证 + 自动仲裁如果用户只是想"问 grok 一个问题",用 ask-opencli。如果用户要做"深度调研"或"交叉验证多个维度",用这个 skill。
从用户的一个研究问题,自动拆分成:
默认分解策略:
用户可覆盖:用户明确说"只问 grok 和 gemini"或"只派内部 agent"时按用户指令。
对每个并行任务,自动生成具体 prompt:
你的任务是**只读数据分析**,不要修改任何文件。
## 背景
{{研究问题的 2-3 句背景描述}}
## 数据源
{{数据库路径或文件列表}}
## 任务
{{具体要查的维度,1-5 个 task}}
## 输出格式
- 结构化 markdown 报告
- 每个结论标注 n(样本数)和 置信度
- 3 屏幕内
- 纯文本返回,不要尝试写文件{{研究问题的精简描述,≤300 字}}
具体问:
(1) {{子问题 1}}
(2) {{子问题 2}}
...
请基于 2026 年上半年真实情况/最新数据回答,要具体可引用。关键要求:问 grok 和 gemini 的 prompt 必须一致(独立对比的前提)。
Tool 1: Agent (general-purpose) run_in_background=true [内部数据 agent A]
Tool 2: Agent (general-purpose) run_in_background=true [内部数据 agent B]
Tool 3: Agent (general-purpose) run_in_background=true [内部数据 agent C]
Tool 4: Bash run_in_background=true [OPENCLI_BROWSER_COMMAND_TIMEOUT=300 opencli grok ask "..." --timeout 300 -f json]
Tool 5: Bash run_in_background=true [opencli gemini ask "..." --format plain]必须在 Claude 的同一条 assistant 消息里一次性调用多个工具,才能真正并行。
Background agent / Bash 任务完成会自动通知。在等的时候,Claude 可以:
明确禁止:
所有结果回来后,按 4 层分类规则自动仲裁每个发现:
| Tier | 判定条件 | 置信度 | 处理 | |---|---|---|---| | 🟢 Strong consensus | 内部数据(n≥20) + grok + gemini 全部支持 | 高 | 可直接写进最终结论,作为硬依据 | | 🟡 Partial consensus | 内部数据 + 1 家外部 AI 支持 | 中 | 小幅建议,保留怀疑 | | 🔴 Conflict | 内部数据 vs 外部 AI 矛盾 | 需仲裁 | 按仲裁原则决定 | | ⚪ Insufficient data | 内部数据 n<10 且没有强外部支持 | 低 | 标为假设,需 A/B 测试 |
原则 1:样本量门槛
原则 2:数据 vs 理论冲突时
原则 3:外部 AI 可疑概念识别
原则 4:反直觉发现的特别处理
外部 AI 有时会给出具体的"案例"(如 @morsyxbt 100→1600/月),这类"单家独有案例"需要二次验证:
验证方式(按优先级):
twitter -c user <handle> 查是否真实存在2026-04-09 实战案例:
@morsyxbt 100→1600/月twitter -c user morsyxbt 验证:真实账号,10.2k followers, verified潜在幻觉的识别:
hallucination,从 report 删除partial verified,保留但降低权重输出结构:
markdown## Action Items (auto-generated, tiered) ### 🔴 极高置信度必做(Strong consensus,可直接落地) 1. [action] — 依据:{数据来源 + AI 共识} — 预期效果:{具体可衡量} - Rollback:{怎么撤销} - Verify:{24-48h 后如何验证} 2. ... ### 🟡 高置信度建议做(Partial consensus) 1. [action] — 依据:{单侧支持} — 前提假设:{什么条件下才成立} 2. ... ### ⚪ 待验证假设(Insufficient data) 1. [hypothesis] — 需要:{什么数据才能确认} — 建议:{A/B 测试设计} 2. ... ### 🚫 不做(Conflict / 反对证据强) 1. [originally planned action] — 反对依据:{为什么不做}
保存到 .omx/artifacts/multi-ai-research-<slug>-<YYYYMMDD-HHMMSS>.md
必须包含:
grok 和 gemini 的有效子命令只有这些,其他都是错的:
bash# ✅ Grok —— 只有一个子命令 ask OPENCLI_BROWSER_COMMAND_TIMEOUT=300 opencli grok ask "问题" --timeout 300 -f json # ✅ Gemini —— 5 个子命令,最常用是 ask opencli gemini ask "问题" --format plain # 最常用(单次问答) opencli gemini new # 开新对话 opencli gemini deep-research "问题" # Deep Research opencli gemini deep-research-result # 取 Deep Research 结果 opencli gemini image "画一个..." # 生图
❌ 常见错误命令(运行会直接报 unknown command):
| ❌ 错误 | ✅ 正确 | 备注 | |---|---|---| | opencli gemini chat "..." | opencli gemini ask "..." | 没有 chat 子命令(常见错误,别习惯性用) | | opencli gemini query "..." | opencli gemini ask "..." | 没有 query | | opencli grok chat "..." | opencli grok ask "..." | grok 只有 ask | | opencli grok new | opencli grok ask "..." --new true | grok 的"新对话"是 ask 的参数 |
记忆锚点:ask 是两家唯一的"问一次"命令。不是 chat,不是 query,不是 prompt。
如果不确定,跑 opencli grok --help 或 opencli gemini --help 看完整子命令列表。
> 官方仓库:https://github.com/jackwener/opencli > npm 包:@jackwener/opencli (npmjs) > 作者:jackwener > License:Apache-2.0
bashnpm install -g @jackwener/opencli
验证:
bashopencli --version
opencli 通过一个轻量的 Chrome 扩展 + 本地 daemon 复用你浏览器已登录的 session。首次运行会自动引导安装:
bashopencli doctor
按提示把扩展加载到 Chrome(通常是 chrome://extensions → 开发者模式 → 加载已解压的扩展,路径 doctor 会告诉你)。
用装了扩展的那个 Chrome profile 打开并登录:
登录一次就行,session 会被 opencli 长期复用。
bash# 加到 ~/.zshrc 或 ~/.bashrc export OPENCLI_BROWSER_COMMAND_TIMEOUT=300
为什么必须:opencli 的默认 browser command timeout 是 60 秒(runtime.js:25),对 grok 复杂问题不够。不设这个 grok 会报 timed out after 60s。这是血泪教训。
opencli 自己提供了几个 AI skill,也可以装:
bashnpx skills add jackwener/opencli
(这些 skill 和 multi-ai-research 不冲突,是互补的。)
bashOPENCLI_BROWSER_COMMAND_TIMEOUT=300 opencli grok ask "请只回复:OK" --timeout 300 -f json opencli gemini ask "请只回复:OK" --format plain
两个都返回 OK 就代表全链路通了。
Detailed material starting at ## Phase 0: Pre-flight Prerequisites(强制检查) has been moved to reference/extended.md to keep this skill concise. Load that reference when the task requires the moved examples, command catalogs, checklists, platform details, or implementation templates.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 37,362 | 33,646 | -10% | 1 | 1 | 0% | 6,231 | 4,780 | -23% | 0 | 0 | — |
case-02 | fail→fail | 27,171 | 48,827 | +80% | 1 | 1 | 0% | 4,026 | 3,858 | -4% | 0 | 0 | — |
case-03 | fail→fail | 30,496 | 45,334 | +49% | 1 | 1 | 0% | 4,607 | 4,034 | -12% | 0 | 0 | — |
case-04 | pass→pass | 7,981 | 3,873 | -51% | 1 | 1 | 0% | 1,360 | 4,094 | +201% | 0 | 0 | — |
case-05 | pass→pass | 9,226 | 5,232 | -43% | 1 | 1 | 0% | 1,612 | 4,420 | +174% | 0 | 0 | — |
case-06 | pass→pass | 4,392 | 3,627 | -17% | 1 | 1 | 0% | 633 | 4,125 | +552% | 0 | 0 | — |
case-07 | pass→pass | 6,863 | 3,984 | -42% | 1 | 1 | 0% | 1,119 | 4,203 | +276% | 0 | 0 | — |
case-08 | fail→pass | 15,926 | 2,672 | -83% | 1 | 1 | 0% | 1,135 | 3,946 | +248% | 0 | 0 | — |
case-09 | fail→pass | 9,708 | 2,486 | -74% | 1 | 1 | 0% | 1,741 | 3,954 | +127% | 0 | 0 | — |
case-10 | fail→pass | 14,271 | 9,085 | -36% | 1 | 1 | 0% | 2,103 | 4,975 | +137% | 0 | 0 | — |
case-11 | pass→pass | 13,105 | 9,265 | -29% | 1 | 1 | 0% | 2,017 | 5,005 | +148% | 0 | 0 | — |
case-12 | fail→pass | 9,426 | 5,826 | -38% | 1 | 1 | 0% | 1,453 | 4,387 | +202% | 0 | 0 | — |
case-13 | pass→pass | 12,988 | 8,285 | -36% | 1 | 1 | 0% | 1,900 | 4,840 | +155% | 0 | 0 | — |
case-14 | fail→fail | 20,648 | 6,815 | -67% | 1 | 1 | 0% | 2,200 | 4,589 | +109% | 0 | 0 | — |
case-15 | fail→pass | 14,526 | 7,904 | -46% | 1 | 1 | 0% | 2,141 | 4,651 | +117% | 0 | 0 | — |
case-16 | fail→pass | 14,478 | 6,707 | -54% | 1 | 1 | 0% | 2,411 | 4,736 | +96% | 0 | 0 | — |
case-17 | fail→pass | 10,021 | 3,440 | -66% | 1 | 1 | 0% | 1,453 | 4,021 | +177% | 0 | 0 | — |
case-18 | fail→pass | 9,852 | 2,737 | -72% | 1 | 1 | 0% | 1,490 | 3,889 | +161% | 0 | 0 | — |
case-19 | fail→pass | 6,253 | 5,240 | -16% | 1 | 1 | 0% | 869 | 4,222 | +386% | 0 | 0 | — |
case-20 | fail→pass | 6,507 | 3,745 | -42% | 1 | 1 | 0% | 943 | 4,117 | +337% | 0 | 0 | — |
case-21 | pass→pass | 13,831 | 8,047 | -42% | 1 | 1 | 0% | 1,952 | 4,828 | +147% | 0 | 0 | — |
case-22 | fail→pass | 12,871 | 5,317 | -59% | 1 | 1 | 0% | 2,069 | 4,382 | +112% | 0 | 0 | — |
case-23 | fail→pass | 7,595 | 3,607 | -53% | 1 | 1 | 0% | 1,181 | 4,083 | +246% | 0 | 0 | — |
case-24 | fail→pass | 11,556 | 3,084 | -73% | 1 | 1 | 0% | 1,687 | 4,027 | +139% | 0 | 0 | — |
case-25 | fail→fail | 7,359 | 1,841 | -75% | 1 | 1 | 0% | 1,121 | 3,768 | +236% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 22 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +52 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.