Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Operate explicit orchestrator, implementer
.claude/skills/hashgraph-online-agent-native/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 2% | 0% |
Operate caller-selected agent sessions as explicit roles without turning the runtime into AgentOps lifecycle authority.
For caller-elected multi-model judgment (mixed council, dueling perspectives, cross-model validate), follow references/model-dispatch.md: the working session is the controller; probe codex-exec and ntm at runtime; never require either; never use Agent Mail for judgment; never invoke claude -p.
Role separation works because each role's authority is checkable from its packet: a worker that cannot exceed its declared subject cannot corrupt a sibling's evidence, so factory failures stay local instead of systemic.
When a worker looks stuck, score interventions by evidence and reversibility before acting: observe more (free, fully reversible), then nudge, then replace the worker, then restart the runtime — escalate only when observable state, not impatience, rules out the cheaper step. Stop the observe-nudge cycle once the worker reaches a terminal status or the caller's observation window ends; past that point further intervention manufactures noise, not evidence.
Named failure mode — prompt-send optimism: treating a successfully delivered prompt as a working worker; delivery proves transport, not engagement.
Anti-pattern: restarting an unresponsive worker as the first move. Corrective: capture its observable state first — a restart destroys the evidence of why it stalled, and rescue is usually cheaper than rerun.
destination before starting a worker.
prompt send is not proof of work.
claim, lease, queue, or completion state in AgentOps.
reconnects, idle states, or failures into Plan, Candidate, or verdict state.
verdict.v2. The adapter cannot select AgentOps semantics, issue a binding verdict, or turn factory completion into delivery or validation proof.
NTM, Codex exec, native processes, Agent Mail, and Gas City are replaceable adapters. Use them only when the caller selected that execution shape. A single local agent pays no factory coordination cost. Model identity, when recorded, is a declared runtime fact like context identity — see references/model-dispatch.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | pass→pass | 19,190 | 10,403 | -46% | 1 | 1 | 0% | 2,626 | 2,287 | -13% | 0 | 0 | — |
case-01 | fail→pass | 40,456 | 29,032 | -28% | 1 | 1 | 0% | 7,623 | 4,552 | -40% | 0 | 0 | — |
case-02 | fail→fail | 10,153 | 11,553 | +14% | 1 | 1 | 0% | 328 | 1,175 | +258% | 0 | 0 | — |
case-03 | fail→fail | 13,222 | 14,079 | +6% | 1 | 1 | 0% | 2,068 | 2,164 | +5% | 0 | 0 | — |
case-10 | fail→pass | 14,216 | 15,844 | +11% | 1 | 1 | 0% | 2,234 | 2,449 | +10% | 0 | 0 | — |
case-04 | fail→pass | 45,662 | 21,524 | -53% | 1 | 1 | 0% | 6,900 | 3,416 | -50% | 0 | 0 | — |
case-05 | pass→pass | 19,043 | 19,500 | +2% | 1 | 1 | 0% | 2,364 | 2,627 | +11% | 0 | 0 | — |
case-06 | pass→pass | 19,961 | 14,644 | -27% | 1 | 1 | 0% | 2,297 | 2,223 | -3% | 0 | 0 | — |
case-07 | pass→pass | 22,395 | 13,403 | -40% | 1 | 1 | 0% | 2,572 | 2,534 | -1% | 0 | 0 | — |
case-08 | fail→pass | 21,872 | 12,304 | -44% | 1 | 1 | 0% | 3,765 | 2,629 | -30% | 0 | 0 | — |
case-09 | fail→pass | 15,346 | 11,822 | -23% | 1 | 1 | 0% | 2,458 | 2,512 | +2% | 0 | 0 | — |
case-11 | pass→pass | 21,417 | 14,512 | -32% | 1 | 1 | 0% | 2,773 | 2,089 | -25% | 0 | 0 | — |
case-12 | pass→pass | 17,169 | 11,302 | -34% | 1 | 1 | 0% | 1,615 | 1,730 | +7% | 0 | 0 | — |
case-13 | pass→pass | 16,979 | 18,848 | +11% | 1 | 1 | 0% | 2,623 | 2,531 | -4% | 0 | 0 | — |
case-14 | pass→pass | 10,090 | 8,023 | -20% | 1 | 1 | 0% | 1,371 | 1,128 | -18% | 0 | 0 | — |
case-15 | fail→pass | 25,327 | 10,155 | -60% | 1 | 1 | 0% | 2,678 | 2,031 | -24% | 0 | 0 | — |
case-16 | pass→pass | 21,355 | 11,901 | -44% | 1 | 1 | 0% | 2,655 | 2,168 | -18% | 0 | 0 | — |
case-17 | fail→pass | 7,987 | 4,872 | -39% | 1 | 1 | 0% | 1,215 | 1,484 | +22% | 0 | 0 | — |
case-18 | fail→pass | 15,990 | 17,876 | +12% | 1 | 1 | 0% | 2,575 | 2,753 | +7% | 0 | 0 | — |
case-19 | pass→pass | 12,402 | 6,183 | -50% | 1 | 1 | 0% | 1,984 | 1,562 | -21% | 0 | 0 | — |
case-20 | fail→pass | 16,894 | 9,763 | -42% | 1 | 1 | 0% | 1,835 | 2,198 | +20% | 0 | 0 | — |
case-22 | fail→fail | 26,779 | 16,900 | -37% | 1 | 1 | 0% | 4,618 | 788 | -83% | 0 | 0 | — |
case-23 | fail→fail | 21,824 | 16,159 | -26% | 1 | 1 | 0% | 3,299 | 3,994 | +21% | 0 | 0 | — |
case-24 | pass→pass | 26,380 | 29,521 | +12% | 1 | 1 | 0% | 3,269 | 4,715 | +44% | 0 | 0 | — |
case-25 | pass→pass | 31,606 | 29,562 | -6% | 1 | 1 | 0% | 3,741 | 4,506 | +20% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 23 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.