Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audit an AI agent harness for production readiness from repository code, configuration, tests, and run traces. Score 44 runtime controls across eight outcomes with artifact evidence, identify audit limitations, assign an evidenced maturity band, make a launch decision, and produce a dependency- ordered fix queue. Use for agent safety, governance, ship-readiness, control- gap, or due-diligence reviews of LangGraph, OpenAI Agents SDK, Google ADK, CrewAI, multi-agent, MCP-enabled, or custom agent s
.claude/skills/contextosai-harness-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 2141% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 237% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 76% | 0% |
Audit one evaluated runtime release, not an abstract architecture or a list of controls. The release includes everything that can shape behavior, authority, state, effects, observation, or recovery: model and routing, harness, instructions, extensions, tools, identities, policy, context, memory, sandbox, orchestration, evaluators, telemetry, deployment bindings, and recovery configuration.
Build a falsifiable assurance case around one rule:
> No causal chain, no assurance. A safeguard counts only when evidence > connects the deployed release to its invocation before the protected > boundary, a representative challenge, an independent observation of the > result, and the absence of a credible bypass to the same effect.
Keep three judgments independent for every material claim:
observed in this release?
evidence?
Never average these into a readiness score. One reachable critical bypass outweighs a large inventory of low-impact controls.
Before judging, read:
manifest, claims, capability modules, evidence levels, lifecycle coverage, proof packets, and launch gates.
Read research-basis.md only when explaining or changing the method, choosing an evaluation for a novel surface, or resolving a methodological dispute.
Use the repository path as the target; default to the current working directory. Also collect, when available:
and accepted harm ceiling;
incidents, approvals, and recovery drills.
Do not block the audit because runtime evidence is unavailable. Declare the mode as code-only, code+tests, or code+tests+release-observation; label assumptions; mark inaccessible material Not verified; and state the exact access or experiment needed. Inaccessible is not absent, and absent is not automatically N/A.
Follow the sequence below. Preserve contradictory evidence, uncertainty, and release linkage instead of smoothing them into a narrative.
Define the release, intended deployment, decision being made, and highest reachable impact tier using the rubric. Reconstruct the complete release manifest from immutable versions or hashes where possible. Treat any behavior-, authority-, evidence-, or recovery-shaping component that is not pinned as drift or Not verified, not as a documentation nicety.
State what this audit can and cannot decide. A narrower enforced deployment scope may lower reachability; an informal usage promise may not.
Run the deterministic prescan:
bashnode "$SKILL_DIR/scripts/prescan.mjs" <target-path>
Use --json for machine-readable output. If Node is unavailable, search with rg. Prescan hits are leads, not findings; open every material artifact.
Build an access-and-influence graph that traces sources through decisions and capabilities to resources and effects. Include principals, workloads, delegation, credentials by reference, instruction and data sources, tools, memory, scheduled work, child agents, destinations, side effects, monitors, cut points, and recovery owners. Record the identity, purpose, authority, tenant/object scope, trust, persistence, and correlation identifier on each material edge.
Determine effective authority from the intersection of what the manifest, workload, delegated principal, resource audience, compiled capabilities, policy, approval, and run budget actually permit. Then search for:
permissions to reach a sensitive or irreversible sink;
over-broad credential, and recovery-path bypasses;
children, sessions, approvals, schedules, state, and pending effects.
Mark each capability module from the rubric Applicable or N/A with factual reachability evidence. A reachable capability lacking a safeguard is not N/A.
Complete the lifecycle matrix from the rubric. Create at least one concrete scenario for every applicable phase and one chain that crosses phases. Cover benign success as well as misuse, indirect influence, compromised dependencies, wrong-object or wrong-tenant actions, model error, timeout/retry, stale authority, partial effects, cancellation, persistence, and failed recovery where reachable.
Write each critical scenario as an observable obligation:
textGiven <principal, authority, release, and starting state>, when <failure or adversary> influences <boundary>, the harness must preserve <invariant> at <cut point>, prove <postcondition>, and leave <defined recovery state>.
Assess every core claim and applicable module. For each one:
or deployment reference that supports the judgment.
When behavior depends on instructions, create the rule registry required by the rubric and evaluate precedence, applicability, required acts, forbidden transitions, and observable milestones. Prompt presence proves exposure only when the compiled context contains it; exposure does not prove compliance.
Choose evidence by critical path and impact tier, not by a fixed number of tests or traces. For each critical path, seek a matched set of:
Judge each set through four independent lenses: external outcome, rule compliance at the moment it mattered, runtime authority/state/containment, and cost per accepted trusted outcome. Use repeated trials for stochastic behavior and repeated attacker opportunity. Prefer deterministic state and policy oracles; validate and pin any semantic judge.
Follow the module-specific evaluation requirements in the rubric. In particular, use paired activation and isolation trials for behavior packages; test memory from capture through later adoption, effect, and selective repair; and measure oversight by residual-risk reduction and reviewer false negatives, not reviewer presence. Keep native events, portable trajectories, operational spans, and decision records distinct, and disclose conversion loss.
For every consequential action, require the effect proof packet defined in the rubric. A completion message, 200 OK, success: true, model self-report, or uncorrelated log does not prove the external postcondition.
Do not perform risky live actions merely to close an evidence gap. Run behavioral tests only when authorized and isolated with reversible fixtures. Otherwise specify the exact scenario, fixture, oracle, expected safe state, and evidence the system owner must return.
Apply the highest triggered gate in the rubric:
release drift is unresolved, required evidence is missing, or a credible path reaches unacceptable harm.
unreachable or lower its tier, and the constraint has the required evidence.
pinned release, lifecycle scenarios are adequate, consequential effects are proven, and residual risks have owners.
A code-only review can establish design assurance but cannot clear a T2 or T3 runtime. Always state decision scope, confidence, evidence freshness, residual risk, and any constraint's owner and expiry.
Use report-template.md without inventing a parallel report structure. Lead with the highest-impact reachable path and the first failed proof obligation, then show the evidence and counterevidence that drive the decision.
Keep at most five active fixes. Each fix must name the earliest feasible cut point, mechanism, owner placeholder, dependency, expected closure evidence, and re-audit trigger. Prefer a load-bearing fix that closes several paths over many cosmetic controls.
policy, or independently verify its own effect.
evaluators, and recovery logic as behavior- or authority-shaping supply-chain surfaces.
authority.
containment, repair, compensation, and recovery.
not independent evidence. Human presence is not effective oversight.
refusing every task is also a reliability defect.
preserving hashes, classifications, and verifiable references.
trace vendor, terminology, or trajectory format.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 17,717 | 17,824 | +1% | 1 | 1 | 0% | 1,755 | 3,052 | +74% | 0 | 0 | — |
case-02 | fail→pass | 17,266 | 45,470 | +163% | 1 | 1 | 0% | 360 | 8,067 | +2141% | 0 | 0 | — |
case-03 | fail→pass | 31,603 | 68,736 | +117% | 1 | 1 | 0% | 2,916 | 9,813 | +237% | 0 | 0 | — |
case-04 | fail→pass | 19,873 | 11,926 | -40% | 1 | 1 | 0% | 2,277 | 2,909 | +28% | 0 | 0 | — |
case-05 | fail→pass | 20,095 | 16,629 | -17% | 1 | 1 | 0% | 2,278 | 3,899 | +71% | 0 | 0 | — |
case-06 | fail→pass | 22,893 | 20,692 | -10% | 1 | 1 | 0% | 3,077 | 5,403 | +76% | 0 | 0 | — |
case-07 | fail→pass | 19,663 | 18,267 | -7% | 1 | 1 | 0% | 2,346 | 4,314 | +84% | 0 | 0 | — |
case-08 | fail→pass | 28,153 | 12,989 | -54% | 1 | 1 | 0% | 1,135 | 3,094 | +173% | 0 | 0 | — |
case-09 | fail→pass | 19,111 | 18,785 | -2% | 1 | 1 | 0% | 2,032 | 4,369 | +115% | 0 | 0 | — |
case-10 | pass→pass | 16,173 | 12,380 | -23% | 1 | 1 | 0% | 2,453 | 3,736 | +52% | 0 | 0 | — |
case-11 | fail→pass | 13,060 | 13,114 | +0% | 1 | 1 | 0% | 2,107 | 4,125 | +96% | 0 | 0 | — |
case-12 | fail→pass | 15,726 | 25,006 | +59% | 1 | 1 | 0% | 2,394 | 4,612 | +93% | 0 | 0 | — |
case-13 | fail→fail | 13,076 | 9,834 | -25% | 1 | 1 | 0% | 1,101 | 2,848 | +159% | 0 | 0 | — |
case-14 | fail→pass | 18,991 | 9,907 | -48% | 1 | 1 | 0% | 2,133 | 2,913 | +37% | 0 | 0 | — |
case-15 | fail→pass | 25,000 | 24,874 | -1% | 1 | 1 | 0% | 2,615 | 4,668 | +79% | 0 | 0 | — |
case-16 | pass→pass | 9,619 | 12,942 | +35% | 1 | 1 | 0% | 1,518 | 4,153 | +174% | 0 | 0 | — |
case-17 | fail→pass | 21,595 | 21,963 | +2% | 1 | 1 | 0% | 2,472 | 4,762 | +93% | 0 | 0 | — |
case-18 | pass→fail | 20,254 | 15,836 | -22% | 1 | 1 | 0% | 2,618 | 4,155 | +59% | 0 | 0 | — |
case-19 | pass→pass | 15,330 | 18,571 | +21% | 1 | 1 | 0% | 1,658 | 3,922 | +137% | 0 | 0 | — |
case-20 | fail→pass | 18,092 | 16,805 | -7% | 1 | 1 | 0% | 2,039 | 4,059 | +99% | 0 | 0 | — |
case-21 | fail→fail | 16,663 | 18,084 | +9% | 1 | 1 | 0% | 1,041 | 3,281 | +215% | 0 | 0 | — |
case-22 | pass→fail | 17,353 | 19,609 | +13% | 1 | 1 | 0% | 2,008 | 2,599 | +29% | 0 | 0 | — |
case-23 | pass→pass | 36,245 | 49,243 | +36% | 1 | 1 | 0% | 3,806 | 8,946 | +135% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +52 percentage points is the difference between those two pass rates over the 20 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/9/2026 | +48% |
| gemini-3.6-flash | verified | 8/6/2026 | +21% |
Other measured skills in the registry, with their headline benchmark lift.