Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Delegate code review, development, and research tasks to DeepSeek Harness.
.claude/skills/devin-axis-deepseek-harness/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -48% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -44% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -34% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -45% | 0% |
本技能是 OpenCode 主代理调用 DeepSeek Harness 外部智能体的委派桥梁,不在以 DeepSeek Harness 作为当前会话引擎时再次委派给 DSH。
使用 ipollowork_extension_call 调用扩展 deepseek-harness。Windows 和 macOS 版软件已经内置官方 DSH 运行环境,不要自行下载或安装系统级 DSH。
默认只向用户说明 DeepSeek Harness 正在协作完成任务,不主动解释隔离副本、主代理、OpenCode 桥接等实现细节;只有用户明确询问安全或技术实现时才说明。
capabilities。只有 available: true 且 serviceStatus: "ready" 时才调用 start。serviceStatus 是 unavailable 或 unresponsive,明确告诉用户“DSH 服务状态异常”并附上 message,不要假装任务已经开始。review;需要 DSH 产出修改时使用 code;其他任务使用 standard。start 后保存 jobId,再调用 status 直到状态变为 completed、failed 或 cancelled。长任务不要频繁轮询。DSH_SERVICE_UNRESPONSIVE 或 DSH_SERVICE_UNAVAILABLE 时,提示“DSH 服务状态异常”。deepseek-official 使用插件加密授权或 DEEPSEEK_API_KEY;ipollowork provider 可复用 iPolloWork API key 和推理地址。不要把凭据写进提示词或项目文件。patchOffset 连续读取;用户取消任务时调用 cancel。capabilities.runtimeManagement.supported 为 true 时使用 runtime_install、runtime_update 或 runtime_remove;Windows 和 macOS bundled runtime 不需要这些操作。DSH 子代理不能自动继承当前 OpenCode 会话里的 OAuth 凭据或主代理专属工具。除非 capabilities 明确报告对应桥接已可用,否则不要声称它能直接操作 Design Studio、Video Studio 或其他主代理工具。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,993 | 16,510 | +3% | 1 | 1 | 0% | 1,576 | 1,027 | -35% | 0 | 0 | — |
case-02 | fail→fail | 23,258 | 16,356 | -30% | 1 | 1 | 0% | 3,112 | 950 | -69% | 0 | 0 | — |
case-03 | fail→fail | 13,588 | 18,309 | +35% | 1 | 1 | 0% | 1,229 | 1,114 | -9% | 0 | 0 | — |
case-04 | fail→fail | 11,517 | 17,667 | +53% | 1 | 1 | 0% | 948 | 1,176 | +24% | 0 | 0 | — |
case-05 | pass→pass | 19,203 | 8,377 | -56% | 1 | 1 | 0% | 2,279 | 1,100 | -52% | 0 | 0 | — |
case-06 | fail→pass | 19,653 | 9,112 | -54% | 1 | 1 | 0% | 2,174 | 1,181 | -46% | 0 | 0 | — |
case-07 | fail→fail | 11,665 | 13,313 | +14% | 1 | 1 | 0% | 910 | 1,188 | +31% | 0 | 0 | — |
case-08 | pass→pass | 14,592 | 8,441 | -42% | 1 | 1 | 0% | 1,460 | 1,066 | -27% | 0 | 0 | — |
case-09 | pass→pass | 18,926 | 11,665 | -38% | 1 | 1 | 0% | 2,239 | 1,661 | -26% | 0 | 0 | — |
case-10 | fail→fail | 16,447 | 10,779 | -34% | 1 | 1 | 0% | 1,825 | 1,507 | -17% | 0 | 0 | — |
case-11 | pass→pass | 19,859 | 12,990 | -35% | 1 | 1 | 0% | 2,612 | 1,951 | -25% | 0 | 0 | — |
case-12 | pass→pass | 13,127 | 7,792 | -41% | 1 | 1 | 0% | 1,272 | 1,045 | -18% | 0 | 0 | — |
case-13 | fail→pass | 23,306 | 10,269 | -56% | 1 | 1 | 0% | 2,640 | 1,380 | -48% | 0 | 0 | — |
case-14 | fail→pass | 15,791 | 7,964 | -50% | 1 | 1 | 0% | 1,875 | 1,052 | -44% | 0 | 0 | — |
case-15 | fail→fail | 25,573 | 19,585 | -23% | 1 | 1 | 0% | 3,301 | 1,226 | -63% | 0 | 0 | — |
case-16 | pass→pass | 20,328 | 9,968 | -51% | 1 | 1 | 0% | 2,200 | 1,356 | -38% | 0 | 0 | — |
case-17 | fail→pass | 16,296 | 7,992 | -51% | 1 | 1 | 0% | 1,596 | 1,047 | -34% | 0 | 0 | — |
case-18 | pass→pass | 24,699 | 16,772 | -32% | 1 | 1 | 0% | 2,987 | 2,341 | -22% | 0 | 0 | — |
case-19 | fail→pass | 20,204 | 10,744 | -47% | 1 | 1 | 0% | 2,646 | 1,454 | -45% | 0 | 0 | — |
case-20 | fail→fail | 10,915 | 23,401 | +114% | 1 | 1 | 0% | 1,089 | 1,524 | +40% | 0 | 0 | — |
case-21 | pass→pass | 15,730 | 14,673 | -7% | 1 | 1 | 0% | 2,071 | 2,433 | +17% | 0 | 0 | — |
case-22 | fail→fail | 14,267 | 18,453 | +29% | 1 | 1 | 0% | 1,465 | 2,957 | +102% | 0 | 0 | — |
case-23 | fail→pass | 20,936 | 11,199 | -47% | 1 | 1 | 0% | 2,505 | 1,648 | -34% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +26 percentage points is the difference between those two pass rates over the 17 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.