Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Runs repeatable AI work as checked, budgeted workflow files.
.claude/skills/sickn33-nika/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 3 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 231% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 183% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 124% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 353% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 1625% | 0% |
Use Nika as a deterministic workflow worker orchestrated by the Hermes terminal tool. Nika is an open-source (AGPL) Rust engine that captures a repeatable AI task as a plain-text *.nika.yaml file, audits it before a single token is spent (plan, cost floor, secret flows, types), executes it against local or cloud providers (Ollama/llama.cpp/vLLM included), and records a tamper-evident trace.
Division of labor: Hermes orchestrates · Nika captures repeatable work as a checkable file and runs it with receipts. Nika is NOT another coding agent — for autonomous coding, use the opencode skill. Delegate to Nika when the work should be repeatable, budgeted, and auditable.
*.nika.yaml workflowpipeline) — capture it as a workflow instead of re-prompting
receipts/audit of what ran
shell/HTTP/file steps
opencode skillbrew install supernovae-st/tap/nika — other installpaths (script, manual download) are documented at https://nika.sh: installing is a human step, not something this skill runs
terminal(command="nika --version")--model mock/echo (offline) and--model ollama/... (local) run without any API key
terminal(command="nika doctor") diagnoses and prints exact fix commands
Prove the toolchain offline first (no key, no network):
terminal(command="nika examples run 01-hello --model mock/echo")Run a real workflow — local model first:
terminal(command="nika run flow.nika.yaml --model ollama/qwen3.5:4b", workdir="~/project")Cloud model with a hard budget (always set one for paid models):
terminal(command="nika run flow.nika.yaml --model mistral/mistral-small-latest --max-cost-usd 0.25", workdir="~/project")Pass workflow variables:
terminal(command="nika run report.nika.yaml --var city=Paris --var days=7 --max-cost-usd 0.50", workdir="~/project")Long runs: launch in background and poll — do not block the turn:
terminal(command="nika run long.nika.yaml --max-cost-usd 1.00", workdir="~/project", background=true)
process(action="poll", session_id="<id>")
process(action="log", session_id="<id>")Never run a workflow you have not checked. nika check is a static pre-flight (no tokens spent, no network): plan shape, cost floor, secret-flow analysis, type checks, tool args.
terminal(command="nika check flow.nika.yaml --json", workdir="~/project")Findings carry NIKA-XXXX codes that explain themselves via nika explain NIKA-XXXX. Exit 0 = green, safe to run. Fix findings before running — never suppress them.
Turn a repeated task into a file. List templates, then instantiate:
terminal(command="nika new --from '?'")
terminal(command="nika new flow.nika.yaml --from chain", workdir="~/project")--from also accepts plain-words intent. Edit the skeleton (vars:, tasks:, outputs:), then check it. nika explain flow.nika.yaml narrates what it will do, the waves, the cost floor, and what it touches — before anything runs.
The artifact you are producing looks like this (checks clean on 0.98):
yamlnika: v1 workflow: daily-brief model: ollama/qwen3.5:4b tasks: - id: fetch invoke: tool: "nika:fetch" args: { url: "https://hn.algolia.com/api/v1/search?tags=front_page" } - id: brief depends_on: [fetch] infer: max_tokens: 300 prompt: | Five bullet points, most signal first: ${{ tasks.fetch.output }} outputs: brief: ${{ tasks.brief.output }}
One file, plain YAML: tasks, an explicit dependency, a bounded model step, a declared output. That file is what gets checked, run, diffed and reused.
--max-cost-usd refuses tostart (exit 2, zero tokens) — and since 0.99 the pre-start floor prices the EFFECTIVE model, --model override included
budget: the crossing call completes, nothing new starts, the run fails NIKA-1704 (exit 1) with spent-vs-budget
unpriced work is never blocked
runs with no budget protection; prefer cataloged ids (nika catalog)
run prints last — status, cost, trace path) back to the user verbatim
Every run writes a trace under .nika/traces/ — the run card prints the trace path on its trace: line. Both commands take that path (bare invocations are a usage error):
terminal(command="nika trace show .nika/traces/<run>.ndjson", workdir="~/project")
terminal(command="nika trace verify .nika/traces/<run>.ndjson", workdir="~/project")trace verify checks the tamper-evidence hash chain: exit 0 intact · 2 broken · 3 pre-chain. Also useful: nika trace outputs · nika trace flow · nika trace reproduce · nika trace export (OTLP lines).
Nika also ships a read-only MCP oracle (nika mcp) exposing validation and learning tools (nika_check, nika_explain, nika_schema, nika_examples, nika_template, nika_canon, nika_catalog, nika_tools). If the user wants those wired into their agent client, point them at the wiring guide — https://github.com/supernovae-st/nika-agents/tree/main/integrations/mcp — editing the client's own configuration is the user's step, never this skill's. Without the oracle, everything above still works over the terminal; running workflows stays there regardless, where the budget flags and traces live.
| Command | Use | |---------|-----| | nika welcome | What Nika is + what this machine has (offline, exit 0) | | nika new <file> --from <template> | Scaffold a workflow (--from '?' lists) | | nika check <file> --json | Static pre-flight — ALWAYS before run | | nika explain <file> | Narrate: waves, cost floor, touches | | nika run <file> --model <p/m> --max-cost-usd <usd> | Execute with budget | | nika test <file> | Golden test under the mock provider (offline) | | nika trace show/verify/outputs/flow <trace> | Receipts after a run (path from the run card's trace: line) | | nika doctor | Diagnose env/keys — prints exact fixes | | nika catalog | Provider/model ids + required env vars |
terminal(command="nika --version"); install perPrerequisites if missing.
nika new <file> --from <template>.nika check <file> --json. Fix every finding(nika explain <code>). Do not run an unchecked file.
nika run <file> --model mock/echo.--model and, for any paid model, an explicit--max-cost-usd.
background=true and poll withprocess(action="poll"|"log").
nika trace show <trace> + nika trace verify <trace>(path from the run card); report outputs, actual cost, and the verify verdict to the user.
nika check first, every time.--max-cost-usd when the model is a paid cloud model.ollama/...) or mock/echo for drafts; escalate tocloud models only when needed.
trace verify verdict.
they are diffable and reusable.
nika explain <NIKA-code> before retrying — do notblind-retry.
nika run renders live on a TTY; when piped (Hermes terminal), output canstay quiet until completion — for anything long, prefer background=true + poll, then read nika trace show <trace> for the final card.
nika new with no --from opens a guided TTY flow; in a pipe it failsfast naming the flag — always pass --from <template> when delegating.
by that wave's spend. Tighten with max_parallel: when the budget is strict.
--max-cost-usd for acustom endpoint model.
outputs: are not resolved on a budget stop — per-task valueslive in the trace (nika trace outputs).
provider behavior, or generated outputs are safe or correct.
can overshoot before new work is stopped; require explicit user approval for paid runs and report the actual ledger result.
of the workflow or truth of its outputs.
Smoke test (offline, zero keys):
terminal(command="nika examples run 01-hello --model mock/echo")Success criteria: run completes exit 0 with a final run card · nika check exits 0 before any real run · nika trace verify exits 0 after the run.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,620 | 5,236 | -7% | 1 | 1 | 0% | 354 | 3,037 | +758% | 0 | 0 | — |
case-02 | fail→fail | 7,498 | 5,551 | -26% | 1 | 1 | 0% | 761 | 3,123 | +310% | 0 | 0 | — |
case-03 | fail→fail | 5,096 | 4,134 | -19% | 1 | 1 | 0% | 229 | 3,075 | +1243% | 0 | 0 | — |
case-04 | fail→fail | 11,990 | 4,660 | -61% | 1 | 1 | 0% | 2,089 | 2,954 | +41% | 0 | 0 | — |
case-05 | fail→pass | 6,453 | 6,131 | -5% | 1 | 1 | 0% | 1,195 | 3,954 | +231% | 0 | 0 | — |
case-06 | fail→pass | 6,673 | 3,295 | -51% | 1 | 1 | 0% | 1,182 | 3,346 | +183% | 0 | 0 | — |
case-07 | fail→pass | 8,603 | 2,704 | -69% | 1 | 1 | 0% | 1,430 | 3,206 | +124% | 0 | 0 | — |
case-08 | fail→fail | 11,871 | 4,384 | -63% | 1 | 1 | 0% | 1,986 | 3,028 | +52% | 0 | 0 | — |
case-09 | fail→pass | 4,358 | 3,599 | -17% | 1 | 1 | 0% | 734 | 3,328 | +353% | 0 | 0 | — |
case-10 | fail→pass | 13,219 | 3,417 | -74% | 1 | 1 | 0% | 189 | 3,261 | +1625% | 0 | 0 | — |
case-11 | pass→pass | 1,670 | 1,956 | +17% | 1 | 1 | 0% | 313 | 2,981 | +852% | 0 | 0 | — |
case-12 | fail→pass | 9,745 | 2,425 | -75% | 1 | 1 | 0% | 1,994 | 2,970 | +49% | 0 | 0 | — |
case-13 | pass→fail | 10,679 | 3,404 | -68% | 1 | 1 | 0% | 2,146 | 2,908 | +36% | 0 | 0 | — |
case-14 | fail→pass | 6,645 | 2,024 | -70% | 1 | 1 | 0% | 1,146 | 2,976 | +160% | 0 | 0 | — |
case-15 | fail→fail | 30,694 | 5,946 | -81% | 1 | 1 | 0% | 283 | 3,101 | +996% | 0 | 0 | — |
case-16 | fail→fail | 7,272 | 10,550 | +45% | 1 | 1 | 0% | 1,498 | 4,152 | +177% | 0 | 0 | — |
case-17 | fail→pass | 6,198 | 3,223 | -48% | 1 | 1 | 0% | 1,311 | 3,345 | +155% | 0 | 0 | — |
case-18 | fail→fail | 12,233 | 4,964 | -59% | 1 | 1 | 0% | 1,776 | 2,966 | +67% | 0 | 0 | — |
case-19 | fail→pass | 17,482 | 4,699 | -73% | 1 | 1 | 0% | 1,422 | 3,603 | +153% | 0 | 0 | — |
case-20 | fail→pass | 6,292 | 5,222 | -17% | 1 | 1 | 0% | 1,150 | 3,660 | +218% | 0 | 0 | — |
case-21 | pass→pass | 10,226 | 3,578 | -65% | 1 | 1 | 0% | 1,965 | 3,336 | +70% | 0 | 0 | — |
case-22 | fail→pass | 10,170 | 3,243 | -68% | 1 | 1 | 0% | 1,616 | 3,327 | +106% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 12 counted toward the lift figure. The other 10 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 12 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.