Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Read every docs/benchmarks/runs/*.json and surface drift in win rate, latency, escalation rate, and LLM-baseline cost over time
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -55% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -84% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -35% | 0% |
The smoke gate is binary (winRate ≥ 0.80 → pass/fail). The corpus benchmarks captured over time form a curve — and curves catch regressions the gate misses (win rate slowly creeping from 100% to 85% is "still passing" by smoke but a real degradation).
This skill reads every persisted run in docs/benchmarks/runs/*.json and reports first→last deltas plus a per-run series, flagging regressions in win rate or latency.
agent-booster — surface latency / strategy changes.bash node plugins/ruflo-cost-tracker/scripts/trend.mjs
Optional env:
TREND_FORMAT=json — emit JSON instead of markdownTREND_LIMIT=10 — consider only the most recent N runsBENCH_ANTHROPIC=1 at run time).> ⚠ Regression callouts when:cost-benchmark — the producer of the run JSONs this skill consumesbench/booster-corpus.json — the corpus version is recorded in each run, so trends across corpus versions remain interpretabledocs/benchmarks/runs/latest.json — the most-recent run; smoke step 23 gates on winRate ≥ 0.80 from this fileOther measured skills in the registry, with their headline benchmark lift.