Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when /better-harness reviews the outer coding-agent Harness for lifecycle controls, repeated work, project feedback, agent assets, session outcomes, repair planning, durable reports, finding-bound fixes, or manual direct fixes. Invoke only via slash command.
.claude/skills/qoderai-better-harness/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 2119% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 133% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 61% | 0% |
Review the coding-agent system: context, execution, control, feedback, and learning; keep Sessions, project, and Agent assets independent.
Route:
<better-harness-fix-output>: Finding-bound Fix.fix, repair, or \u4fee\u590d: Manual Direct Fix.Resolve the Skill path, <better-harness-root> as ../.., a supported <node>, and <cli> as <node> <better-harness-root>/scripts/better-harness.mjs. Stop if any owner is missing; never select another cache or runtime by search order.
Resolve absolute target, decision, acceptance boundary, risks, locale (request language by default), output mode, provider, and depth. Quick uses three items and 7 days; normal uses five and 30 days. Default Qoder/Cursor to durable Canvas and other rendering hosts to HTML. Providers without REPORT_RENDERING proceed only inline or no-files and must not create HTML, Markdown, or Canvas output. Keep providers separate. Use the current one unless project-wide review explicitly authorizes multiple supported providers. Qoder project Memory title metadata is part of the selected workspace baseline. Memory bodies, Codex Memory, Qoder global Memory, user-home, raw Session, installed-plugin, marketplace, and historical-insight access require explicit scope.
Before delegation, collect one versioned evidence bundle per authorized provider:
text<cli> harness evidence-bundle --platform <provider> --workspace <target> --cwd <effective-cwd> --language <locale> --depth <quick|normal> --since <window-start> --until <window-end> --format json [--include-memories] [--include-user-home] [--canvas-out <run-dir>/canvas.json]
Use --canvas-out only for Qoder/Cursor durable reports. For Qoder, keep the default project Memory-title scan; --include-user-home widens it to authorized global Memory/config and other user assets. For Codex, Memory metadata requires --include-memories; user/global or installed-Plugin metadata requires --include-user-home. Apply both when both scopes are authorized. Neither flag authorizes Memory bodies.
It freezes topology, provider, window, depth, limit, and authority. Before delegation, read bundle.context.topology.target; report kind, route, and packageRoute (memberRoute or null). Providers must agree. It returns sessionEvidence, projectHarness, agentCustomize, and the lead envelope. Agent Customize holds bounded lint, inventory, and integrity envelopes from one shared asset snapshot. Keep lane/stage status and providers distinct. Use the individual session-analysis facts, core-change-watch evidence-pack, coding-agent-practices asset-baseline, or harness analyze command only to diagnose a named unavailable or evidence-loss stage; do not substitute diagnostic output into the bundle or rerun all owners. Counts for Rules, Skills, MCP, Memory, Agents, Hooks, Commands, Workflows, and Plugins only route inspection. Zero or high counts never create findings or scores. A normal Qoder report with project Memories blocks when the integrity stage is unavailable; do not replace the missing review with an unobserved disposition.
If the provider discovers or the user supplies a historical insight source, the lead may inspect only a few authorized architecture/history notes. Never assume or search a conventional path; notes cannot prove current behavior, configured capability, or effectiveness.
Launch exactly three fresh, read-only agents in parallel. In Codex use spawn_agent with fork_turns: "none"; otherwise run the same briefs locally and independently. No evidence agent may delegate.
The lead takes the provider-labelled facts envelopes from bundle.lanes.sessionEvidence.data, whose production collector is routed by Sessions Diagnostics, using only the production facts route. Do not pass the complete bundle, collection reference, debug output, or raw sessions to Agent 1.
Give Agent 1 only the provider-labelled facts envelopes, the compact Step 1 asset counts needed to notice zero Skills, and the resolved scope. Require it to read Session Evidence and conditionally read Repeated Workflow Discovery when repeated procedure demand is in scope. It must not inspect the project, configured assets, raw sessions, or another brief.
Give Agent 2 only the target, scoped history/current-change boundary, bundle.lanes.projectHarness.data, decision, risks, and owner limit. Require it to read Project Harness Evidence. It must not receive Session or Agent Customize conclusions.
Give Agent 3 only bundle.lanes.agentCustomize.data with its provider-labelled lint, inventory, and integrity envelopes; asset authority; decision; risks; and owner limit. Require it to read Agent Customize Evidence. It consumes the deterministic envelopes and must not rerun their commands or receive Session/Project conclusions.
Each agent follows its reference-local free-form return contract: normally three to five candidates, up to three in quick mode, and fewer when evidence is sparse. Specialists never assign final severity or scores.
While they run, use only bundle.lead.data as the lead analyzer result. The bundle maps --include-user-home to the analyzer's global-capability boundary; this preserves authorized MCP, Plugin, Skill, Hook, and Memory counts without authorizing content reads or proving use.
Stop if the bundle is failed, the lead lane is unavailable, or its data omits evidence or summaryFacts. In quick mode a partial bundle lowers confidence and every unavailable specialist remains explicit; in normal mode any unavailable or partial specialist lane blocks the report. This evidence pass has a hard cap of three delegated agents.
Read Harness Findings Input for field roles and Agent Work Loop for the five dimensions, checks, evidence states, scoring, and Learning Capture rules. Replace all example content. Never derive the contract from prior reports, Memory, recommend files, or validators.
Perform one reconciliation. Start by retaining every specialist candidate. Merge only candidates with the same target, observed consequence, owner, and repair route; preserve independent consequences even when they share a broader theme. Keep a working reason for every unsupported or deferred candidate. Never drop an eligible finding to reach five rows, shorten the report, simplify a score, or match the three priority moves. Then the lead alone:
confidence, and verifier;
dimension scores before shaping priority moves, repair prompts, or reader copy.
Before drafting, read Findings Quality Gates and apply its eligibility, consistency, privacy, asset, candidate-promotion, and repair-prompt checks directly. For repeated procedure or knowledge demand, also read Asset Demand Reconciliation.
Do not author summary.suggestions in a new report. Promote a suggestion candidate to an ordinary Low finding only when it passes the same consequence, owner, evidence, output, verifier, and repair-prompt gates as every finding; otherwise keep it deferred in the working reconciliation.
After findings and dimension scores are frozen, select exactly one support track from the evidence and requested outcome. The parenthetical ranges are user-journey labels, never score thresholds:
findings establish a missing foundational navigation, validation, or risk route.
show they are not wired into ordinary work or exercised through an outcome.
least two distinct comparable Task Episodes for the repeated goal or friction.
Read only the selected track: Bootstrap Support, Operationalize Support, or Optimize Support. A track may shape at most three priority moves, repair prompts, and reader copy for already-supported findings. It must not add a finding, change severity, rescore a dimension, add a report field, or expand evidence and mutation authority.
For a durable report, draft findings.json only after the three evidence agents finish. Do not launch a fourth review agent. The lead applies the quality gates once, preserves all eligible findings, and fixes any machine-validation failure before rendering.
Inline analysis writes nothing. After lead checks pass, treat the draft as the one final findings.json, then render and validate it once:
textQoder/Cursor: <mode>=<provider>-canvas; <host-root>=<target>/.<provider>/better-harness Other providers: <mode>=html; <host-root>=<target>/.<provider>/better-harness <cli> harness render --findings <run-dir>/findings.json --mode <mode> --out <host-root> --run-dir <run-dir> --target <target> --validate --json
Qoder/Cursor analysis owns adjacent canvas.json; do not copy its summaryFacts into findings. HTML keeps analyzer summaryFacts verbatim. Succeed only on status: pass and return the exact paths reported by render. Never hand-write Canvas, Markdown, or HTML.
Finish with one compact sentence: <count> findings. [Open the report](<renderer-path>). Link the renderer-reported primary report; never return inline-code paths, a bare directory, or an output-file inventory.
a separate independent post-fix agent may update verified finding state and Repair Progress; Loop Effectiveness waits for comparable later Task Episodes.
session-analysis usage-summary once.Loop Discovery.
Core Change Watch, Report Routing, Source Review.
The durable route authorizes only renderer-owned artifacts in its host root. Other creation, activation, mutation, cleanup, scheduling, external writes, and high-risk access require task-local authority. If an owner or value is unresolved, stop with the condition to resume; do not invent a substitute artifact or inspect internal validators.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→fail | 38,557 | 21,427 | -44% | 1 | 1 | 0% | 196 | 3,017 | +1439% | 0 | 0 | — |
case-04 | fail→pass | 20,468 | 42,729 | +109% | 1 | 1 | 0% | 140 | 3,106 | +2119% | 0 | 0 | — |
case-01 | fail→fail | 30,792 | 41,946 | +36% | 1 | 1 | 0% | 4,318 | 3,038 | -30% | 0 | 0 | — |
case-02 | fail→fail | 13,160 | 17,843 | +36% | 1 | 1 | 0% | 1,340 | 3,088 | +130% | 0 | 0 | — |
case-03 | fail→fail | 13,866 | 17,042 | +23% | 1 | 1 | 0% | 207 | 3,013 | +1356% | 0 | 0 | — |
case-06 | fail→fail | 11,814 | 19,193 | +62% | 1 | 1 | 0% | 795 | 3,137 | +295% | 0 | 0 | — |
case-07 | pass→fail | 24,930 | 21,201 | -15% | 1 | 1 | 0% | 2,916 | 3,304 | +13% | 0 | 0 | — |
case-08 | fail→pass | 17,851 | 13,933 | -22% | 1 | 1 | 0% | 1,716 | 3,990 | +133% | 0 | 0 | — |
case-09 | fail→pass | 17,152 | 9,990 | -42% | 1 | 1 | 0% | 1,811 | 3,410 | +88% | 0 | 0 | — |
case-14 | pass→pass | 29,524 | 8,452 | -71% | 1 | 1 | 0% | 1,699 | 3,196 | +88% | 0 | 0 | — |
case-10 | fail→pass | 29,919 | 31,255 | +4% | 1 | 1 | 0% | 2,221 | 3,368 | +52% | 0 | 0 | — |
case-11 | fail→pass | 18,808 | 7,517 | -60% | 1 | 1 | 0% | 1,848 | 2,974 | +61% | 0 | 0 | — |
case-12 | pass→pass | 16,013 | 25,021 | +56% | 1 | 1 | 0% | 1,374 | 3,146 | +129% | 0 | 0 | — |
case-13 | fail→pass | 24,611 | 7,533 | -69% | 1 | 1 | 0% | 1,549 | 3,087 | +99% | 0 | 0 | — |
case-15 | fail→pass | 17,314 | 12,418 | -28% | 1 | 1 | 0% | 1,824 | 3,975 | +118% | 0 | 0 | — |
case-16 | pass→pass | 14,368 | 9,722 | -32% | 1 | 1 | 0% | 1,412 | 3,222 | +128% | 0 | 0 | — |
case-17 | fail→pass | 61,247 | 9,772 | -84% | 1 | 1 | 0% | 2,570 | 3,469 | +35% | 0 | 0 | — |
case-18 | fail→pass | 15,113 | 23,489 | +55% | 1 | 1 | 0% | 1,411 | 3,289 | +133% | 0 | 0 | — |
case-19 | fail→pass | 15,339 | 9,008 | -41% | 1 | 1 | 0% | 1,523 | 3,177 | +109% | 0 | 0 | — |
case-20 | pass→pass | 16,145 | 7,922 | -51% | 1 | 1 | 0% | 1,620 | 3,136 | +94% | 0 | 0 | — |
case-21 | pass→pass | 17,387 | 8,132 | -53% | 1 | 1 | 0% | 1,875 | 3,151 | +68% | 0 | 0 | — |
case-22 | pass→pass | 15,010 | 6,822 | -55% | 1 | 1 | 0% | 1,522 | 2,913 | +91% | 0 | 0 | — |
case-23 | fail→pass | 28,411 | 9,881 | -65% | 1 | 1 | 0% | 971 | 3,203 | +230% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 15 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +43 percentage points is the difference between those two pass rates over the 15 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/27/2026 | +52% |
| gemini-3.6-flash | verified | 8/13/2026 | +27% |
Other measured skills in the registry, with their headline benchmark lift.