Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Runs repeatable AI work as checked, budgeted workflow files. The agent captures a repeated task as a .nika.yaml DAG, audits cost/permits/schema before a single token is spent, and runs it with tamper-evident traces.
.claude/skills/davepoon-nika/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 65% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 22% | 0% |
Nika turns repeatable AI work into files: one .nika.yaml, four verbs, audited before it runs. You author the file; nika check is the oracle; the human runs it.
nika examples list · nika examples show <slug> · nika new --from <template> <file>.nika.yaml
nika: v1 +workflow: <kebab-id> + tasks:. Pick models and builtins from the embedded catalogs — nika catalog (providers · models · capabilities · which env var each needs) and nika tools (the nika:* builtins an invoke reaches without MCP); before a run, nika inspect <file> shows the anatomy: tasks · waves · the cost floor.
nika check <file> (exit 0 = clean · 2 = findings),then nika check --native-strict <file> — it fails on any native-first hint (an exec: a builtin covers).
and fix. Unknown code? nika explain NIKA-XXXX.
not pass nika check — and pass --native-strict too, unless every remaining exec: is in the exec ledger (below).
nika run <file>. Preview offline with--model mock/echo; run locally with --model ollama/<model>. Inputs ride --var key=value (repeatable · unknown keys refused); a run paused on a nika:prompt resumes with nika run <file> --resume <trace> --answer <task>=<value> (confirm gates take booleans: --answer approve=true).
nika test <file> --update writes<file>.golden.json from an offline mock run; nika test <file> replays and compares — deterministic, zero keys.
journal to .nika/traces/ — nika trace verify <trace> checks the chain (tamper-evidence), nika trace show <trace> reads the card. Cite the trace, never a memory of the run.
nika check prints the cost ceiling BEFORE any token: ≤ $X is aceiling · ≥ $X FLOOR means at least one task is unbounded — name the reason (a missing max_tokens, an uncataloged model, an expression fan-out), never round it to $0.
ollama/…) is unpriced compute, not « free » —say "unpriced", never "$0" or "free".
nika run <file> --max-cost-usd <n>blocks BEFORE the call that would cross the cap.
nika explain <file> narrates all of this (waves · cost · touches ·how to run) — use it before handing a workflow to a human.
infer: — an LLM call (prompt, schema? for typed output,max_tokens?)
exec: — a shell command (command, capture: text|structured) ·last resort: run the native-first interrogation first (below)
invoke: — a builtin or MCP tool (tool, args) · HTTP fetch istool: "nika:fetch", a tool, not a verb
agent: — a bounded multi-turn loop (prompt, tools allowlist,max_turns, max_tokens_total)
The order is invoke: nika:* → invoke: mcp:<server>/<tool> → exec:. Before writing ANY exec:, answer in your head:
nika tools --json is the catalog.HTTP (curl/wget/helper fetch) → nika:fetch · uploads → multipart: · site crawls → traverse: · file plumbing (cat/tee/cp/mkdir) → nika:read/nika:write (create_dirs: true) · JSON shaping (jq/sed) → nika:jq (or an output: binding) · in-place edits → nika:edit · image/speech provider calls → nika:image_generate/nika:tts_generate · image styling (ImageMagick convert / PIL filters / dither scripts) → nika:image_fx (deterministic — same input+args = same bytes, the artifact sha256 joins the trace chain).
server, never a helper script.
exec: is legitimate(build tools · git · a product CLI with no MCP surface yet) and goes in the ledger.
Never write a helper script (node bin/helper.mjs …) that wraps HTTP/files/JSON — that is native-first/005, the exact failure class this law exists for.
Every surviving exec: gets a row in the workflow's header comment:
# EXEC LEDGER ·
# | task | command | why no native path | unlock that removes it |--native-strict + a complete ledger = a reviewable workflow.
${{ tasks.<id>.output }} · ${{ vars.x }} ·${{ env.KEY }} · ${{ secrets.X }} (never inline a credential).
depends_on: [<id>].
provider/name (ollama/llama3.2:3b local-first ·mock/echo offline preview).
timeout: "7m") — give localproviders ≥300s: thinking models routinely think past 30s.
nika check --infer-permits <file> printsthe tightest permits: block — paste it in (default-deny from then on).
infer: a schema:; addadditionalProperties: false for a deterministic shape.
headers: { x-api-key: "${{ secrets.KEY }}" } (masked ·declared in secrets: with its egress: sink) — never exec: curl for the sake of a header.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 18,369 | 9,124 | -50% | 1 | 1 | 0% | 2,339 | 3,856 | +65% | 0 | 0 | — |
case-02 | fail→pass | 15,091 | 9,263 | -39% | 1 | 1 | 0% | 3,333 | 3,505 | +5% | 0 | 0 | — |
case-03 | fail→pass | 13,024 | 9,699 | -26% | 1 | 1 | 0% | 2,344 | 3,549 | +51% | 0 | 0 | — |
case-04 | pass→pass | 7,564 | 5,854 | -23% | 1 | 1 | 0% | 1,608 | 2,674 | +66% | 0 | 0 | — |
case-05 | pass→pass | 11,879 | 10,804 | -9% | 1 | 1 | 0% | 2,386 | 3,438 | +44% | 0 | 0 | — |
case-06 | pass→pass | 9,097 | 10,080 | +11% | 1 | 1 | 0% | 1,778 | 3,221 | +81% | 0 | 0 | — |
case-07 | pass→pass | 2,919 | 2,771 | -5% | 1 | 1 | 0% | 533 | 2,110 | +296% | 0 | 0 | — |
case-08 | fail→pass | 11,615 | 4,025 | -65% | 1 | 1 | 0% | 2,079 | 2,302 | +11% | 0 | 0 | — |
case-09 | fail→pass | 9,709 | 3,974 | -59% | 1 | 1 | 0% | 1,878 | 2,286 | +22% | 0 | 0 | — |
case-10 | fail→pass | 14,771 | 4,164 | -72% | 1 | 1 | 0% | 2,452 | 2,191 | -11% | 0 | 0 | — |
case-11 | pass→pass | 13,190 | 3,864 | -71% | 1 | 1 | 0% | 2,449 | 2,305 | -6% | 0 | 0 | — |
case-12 | fail→pass | 16,277 | 2,990 | -82% | 1 | 1 | 0% | 2,603 | 2,083 | -20% | 0 | 0 | — |
case-13 | pass→pass | 12,163 | 2,884 | -76% | 1 | 1 | 0% | 2,270 | 2,146 | -5% | 0 | 0 | — |
case-14 | fail→pass | 17,733 | 3,057 | -83% | 1 | 1 | 0% | 3,129 | 2,192 | -30% | 0 | 0 | — |
case-15 | pass→pass | 8,911 | 1,826 | -80% | 1 | 1 | 0% | 1,520 | 1,879 | +24% | 0 | 0 | — |
case-16 | fail→pass | 11,790 | 1,524 | -87% | 1 | 1 | 0% | 2,119 | 1,858 | -12% | 0 | 0 | — |
case-17 | fail→pass | 14,589 | 1,962 | -87% | 1 | 1 | 0% | 2,629 | 2,024 | -23% | 0 | 0 | — |
case-18 | fail→pass | 8,300 | 2,215 | -73% | 1 | 1 | 0% | 1,414 | 1,996 | +41% | 0 | 0 | — |
case-19 | fail→pass | 28,441 | 1,959 | -93% | 1 | 1 | 0% | 5,064 | 1,951 | -61% | 0 | 0 | — |
case-20 | fail→pass | 12,249 | 2,281 | -81% | 1 | 1 | 0% | 2,177 | 1,985 | -9% | 0 | 0 | — |
case-21 | fail→pass | 6,619 | 2,127 | -68% | 1 | 1 | 0% | 1,123 | 1,990 | +77% | 0 | 0 | — |
case-22 | pass→pass | 5,853 | 1,507 | -74% | 1 | 1 | 0% | 922 | 1,840 | +100% | 0 | 0 | — |
case-23 | fail→pass | 9,063 | 3,376 | -63% | 1 | 1 | 0% | 1,476 | 2,346 | +59% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +65 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.