Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Worker-Agent integration for intelligent task dispatch and performance tracking
.claude/skills/ruvnet-worker-integration/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -37% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -20% | 0% |
Intelligent coordination between background workers and specialized agents.
bash# View agent recommendations for a trigger npx agentic-flow workers agents ultralearn npx agentic-flow workers agents optimize # View performance metrics npx agentic-flow workers metrics # View integration stats npx agentic-flow workers stats --integration
Workers automatically dispatch to optimal agents based on trigger type:
| Trigger | Primary Agents | Fallback | Pipeline Phases | |---------|---------------|----------|-----------------| | ultralearn | researcher, coder | planner | discovery → patterns → vectorization → summary | | optimize | performance-analyzer, coder | researcher | static-analysis → performance → patterns | | audit | security-analyst, tester | reviewer | security → secrets → vulnerability-scan | | benchmark | performance-analyzer | coder, tester | performance → metrics → report | | testgaps | tester | coder | discovery → coverage → gaps | | document | documenter, researcher | coder | api-discovery → patterns → indexing | | deepdive | researcher, security-analyst | coder | call-graph → deps → trace | | refactor | coder, reviewer | researcher | complexity → smells → patterns |
The system learns from execution history to improve agent selection:
typescript// Agent selection considers: // 1. Quality score (0-1) // 2. Success rate // 3. Average latency // 4. Execution count const { agent, confidence, reasoning } = selectBestAgent('optimize'); // agent: "performance-analyzer" // confidence: 0.87 // reasoning: "Selected based on 45 executions with 94.2% success"
Workers store results using consistent patterns:
{trigger}/{topic}/{phase}
Examples:
- ultralearn$auth-module$analysis
- optimize$database$performance
- audit$payment$vulnerabilities
- benchmark$api$metricsAgents are monitored against performance thresholds:
json{ "researcher": { "p95_latency": "<500ms", "memory_mb": "<256MB" }, "coder": { "p95_latency": "<300ms", "quality_score": ">0.85" }, "security-analyst": { "scan_coverage": ">95%", "p95_latency": "<1000ms" } }
Workers provide feedback for continuous improvement:
typescriptimport { workerAgentIntegration } from 'agentic-flow$workers$worker-agent-integration'; // Record execution feedback workerAgentIntegration.recordFeedback( 'optimize', // trigger 'coder', // agent true, // success 245, // latency ms 0.92 // quality score ); // Check compliance const { compliant, violations } = workerAgentIntegration.checkBenchmarkCompliance('coder');
bash$ npx agentic-flow workers stats --integration Worker-Agent Integration Stats ══════════════════════════════ Total Agents: 6 Tracked Agents: 4 Total Feedback: 156 Avg Quality Score: 0.89 Model Cache Stats ───────────────── Hits: 1,234 Misses: 45 Hit Rate: 96.5%
Enable integration features in .claude$settings.json:
json{ "workers": { "enabled": true, "parallel": true, "memoryDepositEnabled": true, "agentMappings": { "ultralearn": ["researcher", "coder"], "optimize": ["performance-analyzer", "coder"] } } }
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 10,776 | 3,326 | -69% | 1 | 1 | 0% | 1,935 | 1,615 | -17% | 0 | 0 | — |
case-02 | fail→pass | 7,965 | 2,707 | -66% | 1 | 1 | 0% | 1,470 | 1,401 | -5% | 0 | 0 | — |
case-03 | fail→pass | 11,299 | 5,965 | -47% | 1 | 1 | 0% | 2,142 | 2,204 | +3% | 0 | 0 | — |
case-04 | fail→pass | 12,507 | 1,659 | -87% | 1 | 1 | 0% | 1,994 | 1,262 | -37% | 0 | 0 | — |
case-05 | fail→pass | 9,585 | 1,801 | -81% | 1 | 1 | 0% | 1,553 | 1,250 | -20% | 0 | 0 | — |
case-06 | fail→pass | 13,420 | 2,898 | -78% | 1 | 1 | 0% | 2,253 | 1,162 | -48% | 0 | 0 | — |
case-07 | fail→pass | 11,941 | 2,121 | -82% | 1 | 1 | 0% | 2,095 | 1,238 | -41% | 0 | 0 | — |
case-08 | fail→pass | 7,430 | 1,975 | -73% | 1 | 1 | 0% | 1,305 | 1,178 | -10% | 0 | 0 | — |
case-09 | fail→pass | 14,059 | 1,638 | -88% | 1 | 1 | 0% | 2,476 | 1,153 | -53% | 0 | 0 | — |
case-10 | fail→pass | 10,633 | 1,667 | -84% | 1 | 1 | 0% | 1,678 | 1,187 | -29% | 0 | 0 | — |
case-11 | fail→pass | 14,604 | 1,394 | -90% | 1 | 1 | 0% | 2,523 | 1,064 | -58% | 0 | 0 | — |
case-12 | fail→pass | 8,639 | 1,891 | -78% | 1 | 1 | 0% | 1,598 | 1,221 | -24% | 0 | 0 | — |
case-13 | fail→pass | 7,230 | 2,459 | -66% | 1 | 1 | 0% | 1,432 | 1,403 | -2% | 0 | 0 | — |
case-14 | fail→pass | 8,510 | 2,582 | -70% | 1 | 1 | 0% | 1,533 | 1,325 | -14% | 0 | 0 | — |
case-15 | fail→pass | 7,462 | 1,494 | -80% | 1 | 1 | 0% | 1,170 | 1,143 | -2% | 0 | 0 | — |
case-16 | fail→pass | 9,832 | 2,101 | -79% | 1 | 1 | 0% | 1,631 | 1,228 | -25% | 0 | 0 | — |
case-17 | fail→pass | 7,648 | 3,290 | -57% | 1 | 1 | 0% | 1,119 | 1,126 | +1% | 0 | 0 | — |
case-18 | fail→pass | 9,963 | 3,633 | -64% | 1 | 1 | 0% | 2,021 | 1,646 | -19% | 0 | 0 | — |
case-19 | fail→pass | 12,759 | 1,683 | -87% | 1 | 1 | 0% | 2,104 | 1,210 | -42% | 0 | 0 | — |
case-20 | pass→pass | 10,255 | 12,031 | +17% | 1 | 1 | 0% | 2,051 | 3,429 | +67% | 0 | 0 | — |
case-21 | pass→pass | 12,129 | 7,358 | -39% | 1 | 1 | 0% | 2,601 | 2,397 | -8% | 0 | 0 | — |
case-22 | pass→pass | 9,909 | 19,552 | +97% | 1 | 1 | 0% | 1,978 | 2,921 | +48% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +86 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.