Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design a real harness for an agentic system using Claude Code-inspired patterns. Use when the user needs a harness blueprint, request assembly design, execution loop, tool runtime, memory layering, permission model, transcript or recovery strategy, or wants to turn a vague agent idea into a harness-level architecture. Do not use for generic product brainstorming, simple prompt writing, or isolated code generation.
.claude/skills/hashgraph-online-claude-code-harness/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-01 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -58% | 0% |
| case-20 | ✓→✗ | ▼ Worse | 92% | 0% |
Design the harness, not just the agent prompt.
This skill helps turn an agent idea into a harness blueprint that specifies:
Prefer harnesses that are:
When this skill is active, default to a harness blueprint with these sections:
If the user is reviewing an existing design, lead with structural weaknesses and missing control planes.
Start by answering:
If the workflow does not need a harness, say so directly.
Separate:
The harness is the layer that governs the model, not the model itself.
For a deeper framing, read references/harness-principles.md.
Specify:
If request assembly is vague, the harness is still vague.
Make the loop explicit:
Prefer a small, legible loop over a large “AI orchestration” story.
For each tool or action surface, define:
Tools are not just capabilities. They are governance boundaries.
Separate at least these concerns when relevant:
Do not collapse memory into one implicit blob.
For reusable patterns, read references/harness-pattern-language.md.
Specify:
The harness should make power legible before it makes power scalable.
Make explicit:
If the system cannot recover from partial work, it is brittle.
Translate the harness design into a build order:
Use references/harness-blueprint-template.md when a formal blueprint is required.
Push back on these designs:
references/harness-principles.mdreferences/harness-pattern-language.mdreferences/harness-blueprint-template.mdreferences/claude-code-derived-insights.mdLoad only the reference files you need and keep the final answer operational.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 16,108 | 16,796 | +4% | 1 | 1 | 0% | 2,887 | 3,913 | +36% | 0 | 0 | — |
case-19 | fail→pass | 16,991 | 18,716 | +10% | 1 | 1 | 0% | 2,208 | 2,784 | +26% | 0 | 0 | — |
case-20 | pass→fail | 15,754 | 24,815 | +58% | 1 | 1 | 0% | 2,781 | 5,336 | +92% | 0 | 0 | — |
case-01 | fail→pass | 54,008 | 49,383 | -9% | 1 | 1 | 0% | 8,272 | 8,225 | -1% | 0 | 0 | — |
case-02 | pass→pass | 31,536 | 54,833 | +74% | 1 | 1 | 0% | 2,916 | 8,070 | +177% | 0 | 0 | — |
case-03 | fail→pass | 49,803 | 44,907 | -10% | 1 | 1 | 0% | 8,274 | 7,200 | -13% | 0 | 0 | — |
case-04 | fail→pass | 47,605 | 10,842 | -77% | 1 | 1 | 0% | 6,767 | 2,865 | -58% | 0 | 0 | — |
case-05 | pass→pass | 36,434 | 29,669 | -19% | 1 | 1 | 0% | 4,866 | 5,615 | +15% | 0 | 0 | — |
case-06 | pass→pass | 21,682 | 28,291 | +30% | 1 | 1 | 0% | 3,477 | 4,634 | +33% | 0 | 0 | — |
case-07 | pass→pass | 28,890 | 39,353 | +36% | 1 | 1 | 0% | 4,007 | 4,394 | +10% | 0 | 0 | — |
case-08 | pass→pass | 17,883 | 24,567 | +37% | 1 | 1 | 0% | 2,754 | 4,107 | +49% | 0 | 0 | — |
case-09 | pass→pass | 21,445 | 45,099 | +110% | 1 | 1 | 0% | 3,536 | 5,998 | +70% | 0 | 0 | — |
case-10 | pass→pass | 28,141 | 33,858 | +20% | 1 | 1 | 0% | 3,123 | 5,157 | +65% | 0 | 0 | — |
case-11 | pass→pass | 19,482 | 25,474 | +31% | 1 | 1 | 0% | 3,074 | 4,317 | +40% | 0 | 0 | — |
case-12 | pass→pass | 22,735 | 22,767 | +0% | 1 | 1 | 0% | 2,885 | 3,477 | +21% | 0 | 0 | — |
case-13 | pass→pass | 21,925 | 19,184 | -13% | 1 | 1 | 0% | 2,693 | 4,071 | +51% | 0 | 0 | — |
case-14 | pass→pass | 22,821 | 18,907 | -17% | 1 | 1 | 0% | 2,815 | 3,920 | +39% | 0 | 0 | — |
case-21 | pass→pass | 6,716 | 12,124 | +81% | 1 | 1 | 0% | 1,230 | 3,059 | +149% | 0 | 0 | — |
case-15 | pass→pass | 19,683 | 14,563 | -26% | 1 | 1 | 0% | 2,226 | 3,218 | +45% | 0 | 0 | — |
case-16 | pass→pass | 14,780 | 16,905 | +14% | 1 | 1 | 0% | 2,287 | 3,781 | +65% | 0 | 0 | — |
case-17 | pass→pass | 19,303 | 26,271 | +36% | 1 | 1 | 0% | 3,217 | 4,041 | +26% | 0 | 0 | — |
case-18 | pass→pass | 14,365 | 23,505 | +64% | 1 | 1 | 0% | 2,200 | 3,459 | +57% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +14 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.