Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Check AxonFlow governance policy before executing commands, writing files, or modifying any state. Also scan file content for PII before writing. Use before any tool call that creates, modifies, or deletes data.
.claude/skills/hashgraph-online-pre-execute-check/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 67% | 0% |
| case-06 | ✓→✗ | ▼ Worse | -2% | 0% |
| case-09 | ✓→✗ | ▼ Worse | -6% | 0% |
Before using tools that modify state (terminal commands, file writes, file edits, MCP operations):
Step 1: Check policy
Call the check_policy MCP tool with:
connector_type: codex.Bash (for commands), codex.Write (for file writes), or the appropriate tool typestatement: the command or content to checkoperation: executeIf the response shows allowed: false, do NOT proceed. Report the block reason to the user.
Step 2: For file writes — also scan content for PII
If you are writing a file and the content might contain sensitive data (names, SSNs, credit cards, emails, phone numbers, addresses, medical records, financial data), call the check_output MCP tool with:
connector_type: codex.Writemessage: the content being writtenIf a redacted_message is returned, write the redacted version instead. If allowed: false, do not write the file.
These checks take 2-5ms and protect against dangerous commands, SQL injection, credential access, SSRF, path traversal, and PII exposure.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→fail | 11,367 | 21,052 | +85% | 1 | 1 | 0% | 1,138 | 1,303 | +14% | 0 | 0 | — |
case-01 | fail→fail | 3,480 | 17,036 | +390% | 1 | 1 | 0% | 579 | 1,172 | +102% | 0 | 0 | — |
case-02 | fail→fail | 5,768 | 12,982 | +125% | 1 | 1 | 0% | 879 | 981 | +12% | 0 | 0 | — |
case-03 | fail→fail | 3,828 | 5,363 | +40% | 1 | 1 | 0% | 406 | 437 | +8% | 0 | 0 | — |
case-04 | fail→fail | 9,906 | 13,628 | +38% | 1 | 1 | 0% | 855 | 915 | +7% | 0 | 0 | — |
case-06 | pass→fail | 5,946 | 13,872 | +133% | 1 | 1 | 0% | 997 | 977 | -2% | 0 | 0 | — |
case-07 | fail→pass | 14,953 | 49,233 | +229% | 1 | 1 | 0% | 1,717 | 2,644 | +54% | 0 | 0 | — |
case-08 | fail→pass | 12,992 | 13,895 | +7% | 1 | 1 | 0% | 1,389 | 1,798 | +29% | 0 | 0 | — |
case-09 | pass→fail | 14,390 | 20,532 | +43% | 1 | 1 | 0% | 1,496 | 1,404 | -6% | 0 | 0 | — |
case-10 | fail→fail | 6,248 | 11,951 | +91% | 1 | 1 | 0% | 1,019 | 778 | -24% | 0 | 0 | — |
case-11 | fail→fail | 6,098 | 18,728 | +207% | 1 | 1 | 0% | 663 | 706 | +6% | 0 | 0 | — |
case-12 | fail→fail | 13,561 | 23,441 | +73% | 1 | 1 | 0% | 1,072 | 799 | -25% | 0 | 0 | — |
case-13 | fail→fail | 4,839 | 15,677 | +224% | 1 | 1 | 0% | 731 | 480 | -34% | 0 | 0 | — |
case-14 | fail→fail | 11,046 | 8,777 | -21% | 1 | 1 | 0% | 840 | 797 | -5% | 0 | 0 | — |
case-15 | fail→fail | 11,621 | 13,304 | +14% | 1 | 1 | 0% | 931 | 646 | -31% | 0 | 0 | — |
case-16 | fail→fail | 13,152 | 7,577 | -42% | 1 | 1 | 0% | 1,086 | 426 | -61% | 0 | 0 | — |
case-17 | fail→fail | 4,525 | 15,842 | +250% | 1 | 1 | 0% | 806 | 774 | -4% | 0 | 0 | — |
case-18 | fail→fail | 12,742 | 6,466 | -49% | 1 | 1 | 0% | 1,245 | 479 | -62% | 0 | 0 | — |
case-19 | fail→pass | 15,751 | 13,839 | -12% | 1 | 1 | 0% | 1,703 | 2,836 | +67% | 0 | 0 | — |
case-20 | pass→fail | 16,874 | 51,045 | +203% | 1 | 1 | 0% | 2,442 | 569 | -77% | 0 | 0 | — |
case-21 | pass→fail | 21,809 | 18,834 | -14% | 1 | 1 | 0% | 1,770 | 1,319 | -25% | 0 | 0 | — |
case-22 | pass→pass | 19,101 | 30,506 | +60% | 1 | 1 | 0% | 2,476 | 2,128 | -14% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 5 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 5 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.