Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Interrogate a demand forecast before the business commits supply and inventory to it. Use when asked to review a demand plan, challenge a forecast, check forecast accuracy, decompose baseline vs uplift, or find hockey sticks in the numbers. Produces a forecast credibility review with baseline/uplift decomposition, MAPE and bias history, hockey-stick flags, an assumption register, and consensus-vs-statistical divergence analysis.
.claude/skills/mohitagw15856-demand-forecast-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-09 | ✓→✗ | ▼ Worse | 82% | 0% |
| case-12 | ✓→✓ | = Same ✓ | 90% | 0% |
| case-13 | ✓→✓ | = Same ✓ | 70% | 0% |
Every unit of forecast becomes purchase orders, capacity commitments, and inventory. This skill interrogates a forecast the way a supply planner must: separate the defensible baseline from hopeful uplift, confront the forecast with its own accuracy history, hunt for hockey sticks, and register every assumption so that when the number misses, you know which belief broke.
Ask for these if not provided:
With no accuracy history, review structure and assumptions and state plainly: [accuracy unknown — treat forecast as unvalidated]. Never present conclusions as if history existed.
1. Decompose baseline vs. uplift. Baseline = what history alone supports (trend + seasonality). Everything above it is an uplift layer that must be named: which promotion, which customer, which launch. Compute uplift share of total — above ~30% uplift, the forecast is a sales plan wearing a forecast's clothes, and each layer needs its own evidence.
2. Confront accuracy history.
| Metric | Read it as | Action threshold | |---|---|---| | MAPE (lag matched to decision lead time) | Noise level | >30% at family level: forecast can't carry item-level commitments | | Bias (signed error, running) | Systematic lean | Same sign 3+ consecutive periods: correct the input, don't buffer around it |
Persistent over-forecast bias means excess inventory is being manufactured upstream; persistent under-forecast means service failures are planned in. Name which one this forecast has.
3. Hunt hockey sticks. Flag: quarter-end/year-end spikes with no order-book support; growth rates that jump beyond trailing actuals precisely when the plan needs them to; a ramp that has slipped right by one quarter in each successive cycle (the sliding hockey stick — the strongest sell-back signal there is).
4. Register assumptions. Every uplift and step-change gets a row: assumption, owner, evidence (order book / customer commitment / pipeline / hope), the period when reality will confirm or kill it, and the volume at stake if it fails.
5. Flag consensus vs. statistical divergence. Where consensus overrides the statistical line by >10%, the override carries the burden of proof. Check the track record: have past overrides beaten the stat model? If overrides historically added error, recommend planning supply to the statistical line and treating the delta as upside to option, not to stock.
1. Verdict — plan to it / plan with stated buffers / send back for rework, and the one-paragraph why.
2. Decomposition — table: Period | Baseline | Uplift layer(s) | Total | Uplift %.
3. Accuracy history — MAPE and bias at the decision lag, trend, and the buffering implication.
4. Flags — hockey sticks, sliding ramps, anomalies vs. history, each with the volume at stake.
5. Assumption register — Assumption | Owner | Evidence strength (committed / probable / speculative) | Confirms by | Units at stake.
6. Divergence analysis — consensus vs. statistical by family; where overrides exceed 10%, the recommendation on which line supply should plan to.
7. Questions for the demand owner — the 3–5 questions that must be answered before commitment.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | pass→pass | 11,494 | 14,187 | +23% | 1 | 1 | 0% | 1,820 | 3,453 | +90% | 0 | 0 | — |
case-13 | pass→pass | 12,788 | 13,492 | +6% | 1 | 1 | 0% | 1,998 | 3,391 | +70% | 0 | 0 | — |
case-01 | fail→pass | 38,614 | 17,807 | -54% | 1 | 1 | 0% | 6,401 | 4,605 | -28% | 0 | 0 | — |
case-02 | pass→pass | 32,322 | 82,154 | +154% | 1 | 1 | 0% | 5,346 | 5,366 | +0% | 0 | 0 | — |
case-03 | pass→pass | 9,900 | 17,184 | +74% | 1 | 1 | 0% | 2,456 | 4,468 | +82% | 0 | 0 | — |
case-04 | pass→pass | 24,113 | 33,468 | +39% | 1 | 1 | 0% | 4,283 | 6,598 | +54% | 0 | 0 | — |
case-05 | pass→pass | 22,569 | 22,109 | -2% | 1 | 1 | 0% | 3,040 | 5,282 | +74% | 0 | 0 | — |
case-06 | pass→pass | 23,803 | 18,714 | -21% | 1 | 1 | 0% | 3,737 | 4,456 | +19% | 0 | 0 | — |
case-07 | pass→pass | 18,019 | 14,768 | -18% | 1 | 1 | 0% | 2,722 | 3,592 | +32% | 0 | 0 | — |
case-08 | pass→pass | 10,390 | 11,586 | +12% | 1 | 1 | 0% | 1,781 | 2,831 | +59% | 0 | 0 | — |
case-09 | pass→fail | 12,102 | 17,393 | +44% | 1 | 1 | 0% | 2,014 | 3,657 | +82% | 0 | 0 | — |
case-10 | pass→pass | 12,689 | 15,941 | +26% | 1 | 1 | 0% | 2,102 | 3,958 | +88% | 0 | 0 | — |
case-11 | pass→pass | 13,651 | 13,539 | -1% | 1 | 1 | 0% | 2,230 | 3,532 | +58% | 0 | 0 | — |
case-14 | pass→pass | 13,156 | 8,975 | -32% | 1 | 1 | 0% | 1,959 | 2,644 | +35% | 0 | 0 | — |
case-15 | pass→pass | 16,982 | 16,402 | -3% | 1 | 1 | 0% | 2,507 | 4,059 | +62% | 0 | 0 | — |
case-16 | pass→pass | 11,309 | 9,558 | -15% | 1 | 1 | 0% | 1,949 | 2,683 | +38% | 0 | 0 | — |
case-17 | pass→pass | 7,588 | 6,445 | -15% | 1 | 1 | 0% | 1,209 | 2,248 | +86% | 0 | 0 | — |
case-18 | pass→pass | 15,589 | 13,381 | -14% | 1 | 1 | 0% | 2,286 | 3,147 | +38% | 0 | 0 | — |
case-19 | pass→pass | 16,872 | 17,018 | +1% | 1 | 1 | 0% | 2,482 | 3,742 | +51% | 0 | 0 | — |
case-20 | pass→pass | 16,181 | 18,317 | +13% | 1 | 1 | 0% | 2,488 | 4,020 | +62% | 0 | 0 | — |
case-21 | pass→pass | 16,636 | 17,942 | +8% | 1 | 1 | 0% | 2,455 | 4,022 | +64% | 0 | 0 | — |
case-22 | fail→pass | 15,645 | 16,573 | +6% | 1 | 1 | 0% | 2,454 | 3,811 | +55% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.